<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>PLOS
ONE 16 (2021) e0252408. URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.
0252408. doi:10.1371/journal.pone.0252408</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.48550/arXiv.2103.00020</article-id>
      <title-group>
        <article-title>Opportunities in the Evolution of Multimodal Generative Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matteo Spanio</string-name>
          <email>spanio@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Multimodal AI, Emotion-aware AI, Generative models, Human-computer interaction, Multisensory integration</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro di Sonologia Computazionale (CSC), Department of Information Engineering, University of Padova</institution>
          ,
          <addr-line>Via Giovanni</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Gardenigo</institution>
          ,
          <addr-line>6b, 35131 Padova (PD)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>1</volume>
      <fpage>25</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>The evolution of generative models in artificial intelligence (AI) has significantly expanded the capacity of machines to process and generate complex multimodal data such as text, images, audio, and video. Despite these advancements, the integration of emotional awareness remains an underexplored dimension. This paper examines the state of the art in multimodal generative AI, with a focus on existing models developed by major technology companies. It then proposes an approach to incorporate emotional awareness into AI models, which would enhance human-machine interaction by improving the interpretability and explainability of AIgenerated decisions. The paper also addresses the challenges associated with building emotion-aware models, including the need for comprehensive multimodal datasets and the computational complexity of incorporating less-explored sensory modalities like olfaction and gustation. Finally, potential solutions are discussed, including the normalization of existing research data and the application of transfer learning to reduce resource demands. These steps are essential for advancing the field and unlocking the potential of emotion-aware multimodal AI in applications such as healthcare, robotics, and virtual assistants.</p>
      </abstract>
      <kwd-group>
        <kwd>Multimodal Generative</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Generative Artificial Intelligence (AI) has experienced rapid advancements, fundamentally transforming
how machines interact with data and create new content. Generative AI models, particularly those based
on deep neural networks, have revolutionized traditional data processing by autonomously learning and
generating complex patterns from raw data. This shift is especially significant in unsupervised learning,
where machines produce coherent and meaningful outputs without explicit guidance. These models can
now generate text, images, audio, and video, opening vast possibilities across industries such as creative
design, healthcare, and robotics [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Among these advancements, the rise of multimodal generative AI
has been particularly impactful. Multimodality refers to AI systems’ ability to process and integrate
various types of data, such as text, images, audio, and video, to perform tasks involving cross-modal
generation or understanding. By bridging diferent sensory inputs and outputs, multimodal AI models
can generate content spanning multiple domains, mimicking a more human-like understanding of the
world. The significance of multimodal AI lies in its capacity to address the limitations of traditional AI
models confined to single modalities. These systems enhance machine perception and understanding,
thereby increasing their applicability in real-world scenarios. However, the rise of multimodal systems
introduces unique challenges. As AI models expand to include more diverse data forms and sensory
inputs, the need for scalable and interpretable models becomes more pressing. Integrating emotional
understanding into generative AI could enrich human-computer interaction and create systems that
better grasp the subtleties of human experience. Despite significant progress in multimodal AI, the
https://matteospanio.github.io/ (M. Spanio)
      </p>
      <p>CEUR</p>
      <p>
        ceur-ws.org
emotional dimension remains largely underexplored [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The author argue that developing
emotionaware models is crucial. Emotions represent a fundamental intersection of perceptual modalities in the
human brain. Key regions such as the amygdala and the hypothalamus play pivotal roles in sensory
perception and emotional regulation, underscoring the importance of emotions and sensations in our
cognitive architecture. Consequently, an emotionally aware AI system should model data in a manner
more aligned with human cognition. The author firmly believes that emotion-aware models could
revolutionize human-machine interaction by improving the interpretability and explainability of model
decisions, addressing a key limitation of current AI systems [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ].
      </p>
      <p>The subsequent sections will explore various aspects of this domain in the following order:
1. Background and Related Work: Outlining the current state of multimodal models, both with and
without emotional awareness, highlighting widely used models from major tech companies like
Google, Meta, and OpenAI.
2. Benefits of Emotionally Aware Models: Examining existing research on emotion-aware models,
focusing on how emotions mediate other perceptual modalities and the potential advantages of
integrating emotional understanding into AI systems.
3. Challenges and Limitations: Providing a critical analysis of the challenges and potential objections
to emotion-aware models, including the complexities in creating multimodal datasets and the
computational demands of representing olfactory and gustatory modalities.
4. Conclusions: Summarizing the key points, reiterating the importance of developing
emotionaware multimodal generative models to enhance human-computer interaction and some proposals
on how to achieve such results.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Work</title>
      <p>The landscape of multimodal generative artificial intelligence (AI) is currently shaped by significant
investments from major technology companies such as Google, Meta, OpenAI, and Microsoft. These
companies are at the forefront of developing cutting-edge models capable of processing and generating
diverse types of data, including images, audio, and video. Their eforts encompass not only the creation
of expansive datasets but also the release of models that are either open source or proprietary. This
trend underscores the immense resources being allocated globally to develop autonomous systems
proficient in generating rich multimedia content.</p>
      <sec id="sec-2-1">
        <title>2.1. Key Models and Developments</title>
        <p>
          Several notable multimodal generative models have been introduced, showcasing the field’s breadth
of capabilities. Text-to-image generation models, such as DALL·E [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and Stable Difusion [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], create
detailed images from textual descriptions. Text-to-audio models, like MusicLM [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], generate music or
soundscapes from text prompts, with promising applications in entertainment and virtual environments.
Although in its early stages, text-to-video generation shows potential in media production and simulation
environments [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. In the other direction, models such as image-to-text [9, 10, 11] translate visual inputs
into descriptive narratives, providing enhanced capabilities for tasks like automated captioning and
assisting individuals with visual impairments. Audio-to-text models, commonly seen in
speech-totext systems, have long been applied in areas such as transcription and virtual assistants, but recent
advances in generative models enable more nuanced and context-aware interpretations of spoken
language. However, even the simplest models involving two non-textual modalities, like the one
discussed in [12] are essentially concatenations of multiple models exchanging textual information.
Recently, multimodal models such as Mirasol, Chameleon, and others (including gtp4o) [13, 14, 15]
have adopted a diferent approach called early fusion [ 16, 17], where the modalities converge into a
single latent space that mixes tokens from diferent domains. Although this approach has yielded better
results than previously described models, it remains dificult to interpret.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Integration of Emotional Awareness</title>
        <p>
          A relatively less explored but emerging area within multimodal AI is the integration of emotional
awareness. While extensive eforts have been dedicated to recognizing emotions within a single
modality [18], there has been an increasing interest in fusing information from multiple modalities
[18, 19]. This multimodal approach is advantageous because the combined information from diferent
modalities provides a complementary capability for emotion recognition. However, relatively few
eforts have been made to understand the emotion-centric correlation between diferent modalities.
Recent approaches, such as those presented in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], have shown concrete possibilities for connecting
images and sounds through an emotional valence-arousal latent space leveraging supervised contrastive
learning techniques. This contribution allows for a more nuanced and dynamic representation of
emotional states compared to the older theory of discrete emotional states. By capturing the subtleties
and complexities of human emotions, these newer models ofer a more sophisticated understanding of
how emotions interplay across diferent sensory inputs.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Benefits of Emotionally Aware Models</title>
      <p>
        While the experiments conducted by [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] yielded promising results, the potential of contrastive learning in
emotional contexts remains largely underexplored. Utilizing supervised contrastive learning allows for
the alignment of various encoders corresponding to diferent modalities within a shared emotional latent
space. This methodology supports the development of models that align with established psychological
research linking modalities to emotions, as evidenced by studies examining the relationships between
audio and emotions [20], odors and emotions [21], and temperature and emotions [22]. By leveraging
these insights, it becomes feasible to include often-overlooked modalities such as touch, taste, and smell,
which also have emotional correlations. Integrating this knowledge could propel advancements toward
Artificial General Intelligence (AGI). Current research in these areas is still nascent, primarily focusing
on textual descriptions, as seen in [23], which utilize transformer-based models. Emotion-aware models
represent a significant advancement in AI research, as mapping human emotions to the latent space of
generative models can foster more natural interactions in applications like virtual assistants, therapeutic
tools, and social robots. By embedding emotional understanding, these systems can engage users in
a more human-like manner, leading to smoother and more relatable interactions. For instance, an
emotion-aware virtual assistant could adapt its tone and suggestions based on the user’s emotional
state, enhancing the user experience. Additionally, imposing constraints on the latent space allows for
the application of known psychological models to clarify AI behavior and decision-making, enhancing
transparency and trustworthiness. Emotions play a crucial role in human perception, influencing how
we interpret sensory inputs. Emotion-aware models can bridge multiple modalities—vision, hearing,
smell, and taste—creating a richer understanding of the environment and leading to more immersive
applications in virtual reality, gaming, and interactive media.
      </p>
      <p>To realize the benefits of emotionally aware models, we propose a framework utilizing pretrained
encoder/decoder architectures tailored for each modality. This approach eficiently encodes emotional
information from sensory data while minimizing computational demands. A pretrained encoder
ifrst processes the input data, such as a computational description of food, transforming it into a
highdimensional embedding. This embedding is then input into a specialized middle encoder model designed
to capture the emotional essence, producing an emotional embedding represented as a vector of valence
and arousal values. Following this, a pretrained decoder model converts the emotional vector into an
audio token, which is processed by the pretrained audio model to generate the corresponding output.
This architecture efectively manages the computational load through pretrained models, requiring
only the middle encoder/decoder model to be trained. This design streamlines the training process and
enhances the model’s capacity to translate emotional information across modalities, fostering a more
integrated and emotionally aware AI system.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Challenges and Limitations</title>
      <p>The development of multimodal AI models faces several significant challenges and limitations,
particularly concerning the availability and quality of datasets. One of the primary obstacles is finding
comprehensive datasets that integrate multiple modalities. While substantial progress has been made
in creating large-scale datasets for individual modalities, the integration of diverse sensory data, such
as combining visual, auditory, and textual information, remains a challenge. This paucity of integrated
datasets hampers the ability to train and evaluate multimodal models efectively. Moreover, the
collection of multimodal data is both costly and labor-intensive, requiring expert evaluations to ensure
data quality and alignment across modalities. Web scraping, a common method for gathering large
amounts of data, proves insuficient for creating high-quality multimodal datasets. For instance, aligning
diferent types of data (e.g., synchronizing audio with visual inputs) necessitates precise and controlled
conditions, often achievable only in well-equipped laboratories. Existing datasets predominantly rely on
massive, web-scraped data, which, while voluminous, often lack the quality needed for advanced deep
learning applications. The release of models like Microsoft’s Phi [ 24] has underscored the importance
of data quality, demonstrating how high-quality datasets can enhance model eficiency and reduce
computational resource requirements. In addition to these general challenges, specific modalities
such as olfaction and gustation present unique dificulties. In these research communities, there is
no established practice of sharing data in formats suitable for use as training datasets. Furthermore,
there are no widely adopted computational representations for olfactory and gustatory information,
making it challenging to integrate these modalities into multimodal models. Before end-to-end models
that encompass these senses can be developed, significant research is needed to establish standardized
computational frameworks and methodologies for these less-explored sensory domains.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>The integration of emotional awareness into multimodal generative AI represents a pivotal next step
for advancing human-computer interaction. Emotion-aware AI models can interpret and respond more
contextually to human emotions, which is particularly relevant in applications like virtual assistants
and healthcare, where nuanced emotional understanding is critical. Furthermore, these models enhance
interpretability and trustworthiness in AI decisions. However, key challenges remain, especially in
dataset availability and computational complexity. Comprehensive multimodal datasets that incorporate
emotions are scarce, and the dificulty in representing sensory modalities such as olfaction and gustation
hinders progress. Addressing these issues is crucial to fully realize the potential of emotion-aware AI.</p>
      <sec id="sec-5-1">
        <title>5.1. Future Research Directions</title>
        <p>To advance research in emotion-aware multimodal AI, two key strategies are proposed:
1. Dataset aggregation and normalization: a considerable body of research, particularly in
psychology and neuroscience, has already explored the correlation between emotions and various
sensory modalities. Although these data are currently dispersed and non-standardized, they often
come from high-quality studies. A concerted efort to systematically aggregate and normalize
these existing datasets could form the basis for a comprehensive multimodal dataset. Such a
resource would support AI models in learning emotional correlations across multiple sensory
inputs, creating a foundation for more robust emotion-aware systems.
2. Leveraging transfer learning to reduce computational complexity: the computational demands of
building models from scratch, as seen in large tech companies, are a significant barrier for many
research initiatives. However, existing deep learning-based encoders have reached a high level of
performance. By adopting a contrastive learning framework and leveraging transfer learning or
ifne-tuning techniques, researchers can utilize pre-existing models to enhance emotion-aware
capabilities without requiring massive computational resources. This approach not only shortens
the time needed to achieve results but also reduces energy consumption, making research more
sustainable and accessible.</p>
        <p>In summary, advancing emotion-aware multimodal AI requires addressing the current challenges
of dataset availability and computational demands. By capitalizing on existing research data and
leveraging transfer learning, these obstacles can be overcome, enabling the development of AI systems
that are more aligned with human emotions. Such systems will significantly enhance human-computer
interaction and broaden the scope of AI applications in various fields.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>I would like to express my sincere gratitude to Professor Antonio Rodà and Professor Massimiliano
Zampini for their invaluable discussions and advice throughout the development of this paper. Professor
Rodà, from the Centro di Sonologia Computazionale (CSC), Department of Information Engineering,
University of Padova, provided essential guidance and support. Professor Zampini, from the Center for
Mind/Brain Sciences (CIMeC), University of Trento, ofered key insights that significantly shaped this
research. Their expertise and encouragement have been crucial to this work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ramesh</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hesse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , E. Sigler,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Litwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chess</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
          </string-name>
          , Language Models are
          <string-name>
            <surname>Few-Shot Learners</surname>
          </string-name>
          ,
          <year>2020</year>
          . URL: http://arxiv.org/abs/
          <year>2005</year>
          .14165. doi:
          <volume>10</volume>
          .48550/arXiv.
          <year>2005</year>
          .
          <volume>14165</volume>
          , arXiv:
          <year>2005</year>
          .14165 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Keutzer</surname>
          </string-name>
          ,
          <article-title>Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal Space</article-title>
          ,
          <source>in: Proceedings of the 28th ACM International Conference on Multimedia, MM '20</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2020</year>
          , pp.
          <fpage>2945</fpage>
          -
          <lpage>2954</lpage>
          . URL: https://dl.acm.org/doi/10.1145/3394171.3413776. doi:
          <volume>10</volume>
          .1145/3394171. 3413776.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Calvo</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>D'Mello, Afect Detection: An Interdisciplinary Review of Models, Methods, and Their Applications</article-title>
          ,
          <source>IEEE Transactions on Afective Computing</source>
          <volume>1</volume>
          (
          <year>2010</year>
          )
          <fpage>18</fpage>
          -
          <lpage>37</lpage>
          . URL: https: //ieeexplore.ieee.org/document/5520655. doi:
          <volume>10</volume>
          .1109/T- AFFC.
          <year>2010</year>
          .1,
          <string-name>
            <surname>conference</surname>
            <given-names>Name</given-names>
          </string-name>
          :
          <source>IEEE Transactions on Afective Computing.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rudin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Semenova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <source>Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges</source>
          ,
          <year>2021</year>
          . URL: http://arxiv.org/abs/2103.11251. doi:
          <volume>10</volume>
          .48550/arXiv.2103.11251, arXiv:
          <fpage>2103</fpage>
          .11251 [cs, stat].
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pavlov</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Voss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Sutskever</given-names>
            ,
            <surname>Zero-Shot Text-toImage Generation</surname>
          </string-name>
          ,
          <year>2021</year>
          . URL: http://arxiv.org/abs/2102.12092. doi:
          <volume>10</volume>
          .48550/arXiv.2102.12092, arXiv:
          <fpage>2102</fpage>
          .12092 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rombach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Blattmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lorenz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Esser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ommer</surname>
          </string-name>
          ,
          <string-name>
            <surname>High-Resolution Image Synthesis with Latent Difusion Models</surname>
          </string-name>
          ,
          <year>2022</year>
          . URL: http://arxiv.org/abs/2112.10752. doi:
          <volume>10</volume>
          .48550/arXiv.2112. 10752, arXiv:
          <fpage>2112</fpage>
          .10752 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Agostinelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. I.</given-names>
            <surname>Denk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Borsos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Engel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Verzetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caillon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tagliasacchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sharifi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zeghidour</surname>
          </string-name>
          , C. Frank, MusicLM: Generating Music From Text,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2301.11325. doi:
          <volume>10</volume>
          .48550/arXiv.2301.11325, arXiv:
          <fpage>2301</fpage>
          .11325 [cs, eess].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>U.</given-names>
            <surname>Singer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polyak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hayes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Ashual</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Gafni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Taigman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Make-A-Video:</surname>
          </string-name>
          Text-to-Video Generation without
          <string-name>
            <surname>Text-Video Data</surname>
          </string-name>
          ,
          <year>2022</year>
          . URL: http://arxiv.org/abs/2209.14792. doi:
          <volume>10</volume>
          .48550/arXiv.2209.14792, arXiv:
          <fpage>2209</fpage>
          .14792 [cs].
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>