<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>One Picture and a Thousand Words</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Generative Language</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>images Models</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>How to Train Them</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Zamparelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CIMEC, University of Trento / Corso Bettini 31</institution>
          ,
          <addr-line>Rovereto TN, 38068</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Thanks to independent advances in language and image generation we could soon be in the position to have systems that communicate with humans by combining language and images in their output, a skill that humans do not possess (we receive, but do not produce images at high speed). The paper explores some of the implications of this idea: which kinds of data sets need to be developed to train such systems, in which cases language and images could be most usefully integrated and which issues could arise on the image generation and language+image integration side.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Language generation</kwd>
        <kwd>image generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>how (maybe, eventually) Starships will land on Mars, but showing this graphically requires
professional video makers and special efects.</p>
      <p>The goal of this piece is to point out that AI systems from the near future will not have
these limitations: they could in principle be able to produce language and still images (soon,
moving images and ‘stage noises’), mixing them in any way we can imagine and in many we
cannot (yet). Moreover, this mix could be fast enough to take place in dialogue: a chatbot could
answer my questions with an image accompanied by verbal explanations, adjust the image if it
is unclear, answer my questions by expanding details, link referring expressions to points in the
picture, render focus by unblurring parts, making sections transparent or adjusting colors, etc.
None of these capabilities exist in humans: gestures are too generic, photos are restricted in
style and possibly too detailed, drawing is too slow to be usable without interrupting the flow
of the conversation. Put otherwise, combining language and images in on-line generation is a
task in which AI systems not from the realm of science-fiction (due to independent advances in
language generation and image generation, see Sec. 2) could very soon achieve superhuman
abilities and produce actually useful output.1</p>
      <p>This unique status comes, of course, with a number of problems. The aspects where NLP
or artificial vision have historically progressed the fastest are those where they could leverage
huge datasets, annotated via crowd-sourcing &amp; bootstrapping. Humans do create combinations
of images and texts in various contexts and with various purposes, but these contexts and
purposes appear to be a small subset of the cases where it would be useful for an AI to generate
what I am going to call I+L (Images+Language): none of these contexts and cases is both fast
and interactive.</p>
      <p>
        The problem of creating systems that can output both language and images has actually been
around since the 80ies (see [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) often under the name Multimodal Output Generation (MOG,
see [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]). In the previous literature, however, it was seen as a procedural problem: finding the
bast way to combine textual explanations with either pre-stored materials (e.g. product pictures,
see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], medical images, etc.) or data-generated graphs [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The perceived central issue was
planning what to serve where, and in which modality. Machine learning, when present, was
used to find the best combination of the two modalities [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], not as a way to generate either of
them. To take advantage of the additional possibilities given by modern generation models we
need to turn our attention to three aspects: I+L Training, I+L Theory and I+L Input. This paper
gives an overview of some of the issues that arise within these domains, mostly seen from the
vantage point of language and linguistics.
      </p>
      <p>The paper is organized as follows. Sec. 2 gives a landscape of the current AI models the I+L
idea builds on, and discusses the (lack of) current datasets. In Sec. 3 I highlight some questions
raised by the development of I+L models. Sec. 4 discusses possible ways to develop training
materials, while a few linguistic issues arising from the combination of images and language
are discussed in Sec. 5.
1As a side note, no deepfakes are possible at this level: a system that is seen answering questions by producing
intertwined text and images could not be mistaken for a human production simply because no human could
accomplish this feat. Of course, the final output could raise the same issues or authorship and attribution as other
artificial creations, and give raise to the security concerns raised by its individual constituents.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and middle-hanging fruits</title>
      <p>
        The possibility of generating I+L builds on recent advances in a number of neighbouring tasks:
language generation (OpenAI’s GPT3, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Google’s PaLM, [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], META’s OPT, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], etc.),
textto-image (T2I) models (GLIDE, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], DALL-E, [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]), text&amp;image-to-image generation [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and
multiple-round image editing with language [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ]. Many of these advances have been
made possible by a move from annotated images datasets such as ImageNet [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] or MS-COCO
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] to image embeddings like CLIP [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], trained on image-text matching collected on a huge
scale. CLIP does not have a limited vocabulary and has been shown to improve VQA tasks [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ],
including image captioning [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], another task quite relevant for I+L. Other advances come from
applying to image generation the textual embeddings developed by very recent, large-scale
language models (e.g. T5 in IMAGEN, [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]), combined with image upscaling. Given the strict
connection between language and images required by a I+L system, another relevant strand of
research is the one that aims to go from images to comprehensive descriptions, far longer than
typical user-generated captions, either by associating captions to the individual bounding boxes
(Dense Captioning, [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]) or by creating extensive human-readable descriptions from images
(Image Paragraph, [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]). The aim of some of this work is assistive technology for the visually
impaired [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], but they could be very useful to map linguistic descriptions of entities back to
the images, or as part of a circular flow that goes from language to images and back to language
(see Fig. 4).
      </p>
      <p>
        Overall, the field has seen an extremely fast growth, with new open-source data made available
monthly (e.g. the CLIP-based LAION-5B2 dataset), ongoing eforts to make models trainable
at lower costs [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], and the release of open-source T2I software for low-end hardware (Stable
Difusion 3), which is likely to create a huge pool of testers for future models.
      </p>
      <p>
        In all the cases considered, however, the goal is to produce images (from text or text+images),
or text which either describes the image (automatic captioning) or uses the image as ground
truth (e.g. in VQA, [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]). These tasks are generalizations drawn from innumerable examples of
text-image associations (in the form of human-generated captions, textual contexts or image
metadata). The goal of the sort of I+L system envisoned here is diferent: the visual and the
linguistic side of the output should provide contents that are both connected and complementary.
There have been attempts at creating datasets where the text and images are more
complementary (see BD2BB, [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]), but there are many more ways to be complementary than to coincide,
and more systematic ways to generate training materials must be found. At the time of writing,
the Task section of the huge Paperswithcode Dataset4 repository does not contain any task
resembling image and text generation.
      </p>
      <p>One domain where humans routinely combine voice and images are presentations with
slides (Powerpoint©, etc.). Large public slide repositories do exist (e.g. the Slideshare-1M5),
but no dataset that combines the speaker’s voice (or its transcription) with the slide presented.
TED Talks, another possible model of a certain way of combining speech and images, have
been compiled in various machine-learning-friendly forms, but not in one that systematically
2https://laion.ai/blog/laion-5b/
3https://stability.ai/blog/stable-difusion-public-release
4https://paperswithcode.com/datasets
5https://purl.stanford.edu/mv327tb8364
associates the slides used with the speech (TED Talks6 has voice and videos of the speaker’s
body; TED LIUM 37 only audio and transcripts; the TED Gesture Dataset8 could be used to
teach systems to accompany speech with gestures, but obviously the potential of I+L generation
goes much beyond this).</p>
      <p>
        Next, there are instructional videos (on products, procedures, DIY projects, etc.). These could
provide valuable training material, as long as they do no not (mostly) show the speaker. Turning
them into I+L-training material would involve reducing the video to maximally informative still
frames of the type/resolution that could be matched by current T2I systems, then linking them
to the corresponding voice passages. Note that multiple video frames could be used to generate
a 3D representation of the objects shown (e.g. using NeRF-like techniques, [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]). In this case,
the final dataset would contain a paired combination of text/voice and the 3D embeddings of
the entities mentioned at every point. An I+L system could learn to accompany the text passage
with a single 3D image, giving the listener the possibility to explore it interactively (asking for
rotation, zooming, etc.).
      </p>
      <p>
        Finally, there are commercial movies, and there are comics. When the former are combined
with storyline datasets like MPST [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] it may be possible to extract still frames that best highlight
specific storyline passages and that could be used for I+L training (see Phenaki 9 [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] for an
attempt along this line, with videos, minus the textual output). Unlike with videos, it remains to
be seen if a movie’s storyline is not too coarse to be meaningfully mapped to the still frames. The
case of comics is in a way the literary genre that comes closest to the possible output of a type
of I+L system. The problem, in this cases, is in the very limited range of topics and themes that
the comics literature covers (especially if one wants to exclude copyrighted material). Comics
could be useful to train a I+L system that produces just that — comics, but their utility at large
remains to be proved.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Open questions</title>
      <p>
        As we have need, current AI research has a large variety of datasets available for training models,
yet none seems ready and set to be used for I+L development. The questions to address are:
• Which already existing datasets could be repurposed to train I+L?
• Which novel datasets specific for I+L generation should we start developing, and how?
• Can build broad L+I training datasets which can be fine-tuned to specific tasks and
applications?
To create new datasets or help the systems use existing ones in the best way we need to develop
a theory of how I+L should be done. Possible questions to answer are:
• Use cases Which applications could profit the most from the integration of language and
vision, and which aspects of NLP should stick to words alone? An obvious application
6https://paperswithcode.com/dataset/ted-talks
7https://paperswithcode.com/dataset/ted-lium-3
8https://paperswithcode.com/dataset/ted-gesture-dataset
9https://phenaki.video/
is the generation of fiction : creating a mix of images and written narrative from a story
line, or turning a text-only story into a multimodal narration. We should of course ask
how useful this task really is. As discussed in the context of movies-from-books [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ],
depicting narrative can actually stifle mental imagery, standardize places and characters,
and ultimately reduce the appreciation of the story. However, if the generation system
is flexible enough to give the human story creator control over the way the narration is
depicted or the ability to customize the scenes and the look of the characters, the L+I
narration shifts in the direction of film-making, a respectable creative activity. A more
radical option is to give the viewer the possibility of customizing the imagery (changing
style, light, possibly background). This stretches the boundary between creation and
fruition, but also between images and textual narrative (if I can change the illustrations,
why not the story?). Scene customization is in fact a natural extension of the common
practice of letting video-game players choose the likeness of their avatars, though here
the choice is done by presenting compositional alternatives to choose from (“Do you
want this avatar? This style of hair”). Still, avatars represent the self — in the case of
stories generated via a L+I system, it’s much more likely that many viewers would accept
passive fruition. Ultimately, the trade of between the complexity of a system that allows
broad customization and the number of user interested in using it should be carefully
considered.
      </p>
      <p>
        Storytelling is just only one of many tasks that generation systems could tackle. At
least initially, a more likely candidate for practical I+L systems is the creation of how-to
instructional videos (e.g. “How to replace the battery of your phone.”), descriptions of
possible activities (e.g. a mountain walk created by combining 3D map imagery with
a commentary on dificult steps or things to see; a road trip guide that shows various
possible stops and their attractions), rendering of events for which we have only textual
descriptions, renderings of complex objects (e.g. the 3D view of a device extracted from a
patent description) and probably many other non-obvious applications that will emerge
as the systems mature.
• The focus of interactive communication. Where would the dialogic dimension
add the most value, compared to a static I+L system? Currently, the domain that best
combines NLP and image generation is image editing, where the focus is the image itself.
Current T2I models create images which may be far from their creator’s expectations; the
subfield of Multimodal Image Synthesis and Editing (MISE, see [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] for a recent review)
deals with the best approach to direct the generation process or edit the result. While
many systems use input diferent from text in the editing stage, high-level linguistic
instructions could be a powerful tool to refine them (e.g., in Fig. 2: “Make the teddy bear
smaller and move it to the center of the zebras.”). The system could answer by modifying
the image (see Sec. 2), but it is easy to see the usefulness of interaction, in clarifications
(“Should the skate also be smaller? Should I keep the bear on focus?”) or explanations
(User: “What is the object on the right?”; System: “A shop.”; “Can you show a specific
type of shop?”, “Which one?”).
      </p>
      <p>Despite the importance of MISE, an important yet undeveloped use case remains that of a
system that provides textual feedback to explain things beyond what an image shows: its
context, or any aspect that cannot be cast visually (see e.g. (1) below). This contrasts with
standard image captioning tasks, where the text describes what the image does show.</p>
      <p>
        The I+L problem is how text and image can be used together to provide information.
• How to deal with cases where language and images interact? A case in point is
anaphoric reference from text to the object in the image. Should for instance objects in
a scene flash or light up when mentioned by the system or by the user? In a modality
where the system sees the user as he or she looks at a screen, how should pointing be
detected and interpreted? (What is this? Can you show the other side of this? [pointing to
a detail in the image]).
• Image-Language Ratio. When contents are rendered as I+L, what is the ideal
images/language ratio? This question has been debated for a long time (see [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]), but it
acquires new depths with the range of possibilities ofered by generative AI. Human (of
line) examples go from illustrated novels (text, for the most part) to the intertitles that
occasionally interrupt silent movies (images, for the most part). The ratio could be a
user choice (see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), but varying it would forces the system to analyze and structure the
contents diferently. Moreover, should the text be voiced, written with/before/after the
image, or both? A theoretical analysis by Stenning and Oberlander [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] suggests that
images, being less abstract, can in many cases facilitate processing, but their conclusions
do not keep into account the cognitive cost of images that are too or not enough specific,
or consistent.
• Level of specificity In many cases, an I+L system could make use of a mix of realistic
and computer-generated materials. How can we clarify to the viewer/listener when an
image tries to give an accurate rendition of a particular scene or object, and when it is a
generic representation, or a mere filler? (In the description of the crash, the picture of the
car is specific, the lamp post it hits, made up).
• Consistency The images produced by the system should remain consistent throughout
the discourse, on pain of confusing the user. For instance, a story about a dog of unspecified
breed should prompt the AI to draw a particular type of dog, but the dog must then remain
of that breed in any subsequent image. There is a trade of between the dialogic dimension
and consistency: information that becomes available in the development of the dialogue
might contradict a random design choice made by the system early on (for instance, the
system might have drawn the mention of a dog as a German shepherd, to later discover
it was supposed to be a chihuahua). If the system has the possibility of processing the
information from whole text beforehand this pitfall can be avoided, but the range of use
cases may be reduced. Similarly, the system should avoid unwanted visual transformation
as it transitions from scene to scene; switching to a mirror-image of the previous scene can
be confusing, and so can changing the visual style and the appearence of the characters.
Yet, variation is the norm in current models, and current solutions around this problem
resolve in bags of tricks10.
      </p>
      <p>
        Lastly, there are technical aspects related to the input. Among them:
• What kind of input should a I+L generation system receive? Current image generators
create images from text, with a wide margin of variation (Fig. 1), due to the
underspec10https://mythicalai.substack.com/p/how-to-generate-consistent-characters
ification of textual information and to the non-deterministic nature of the generation
algorithms (e.g. difusion models). Images with a large part made up would be appropriate
when the I+L system is meant to provide a decoration for the text, not when an image
must be a faithful rendering of a description which is grasped more easily in graphic than
in textual form. In this case the input should include images of the specific object that
needs to be represented (see Sec. 4, esp. [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ])
• A textual input to the image generation system needs not be normal human language. It
could be language augmented with extremely detailed visual descriptions (on the model
of the text prompts11 used for image generation). But how can this type of input be
provided for training? What is the best image+language input combination?
The answer to these questions is of course conditioned on what is already (potentially) available
on the training side.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Dataset-building strategies</title>
      <p>
        It would of course be relatively easy to use existing T2I system to associate images to the text
generated by an automatic system (think of this as a DALL-E-over-GPT3-output approach). In
principle, the process could be carried out at diferent levels of granularity. At one end, we
could imagine generating images for all noun phrases in the text above a given threshold of
imageability [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ]12, then selecting some to be presented alongside the text. At the other end
of the spectrum, one could pass each generated sentence to the T2I generator, then filter the
11https://neuroflash.com/blog/ai-image-generation-prompt-examples/
12It does not make much sense to try and generate images of individual nouns, as opposed to noun phrases. The
ability to select existing images matching a conversation was demoed by Amazon in the re:MARS 2022 keynote.
output in some way. Neither of these extremes are likely to be very useful. What is needed
instead is a splitter that decides which aspects of the input text should be cast in graphic form
and which ones as text, starting from a required text/image ratio that will be low for, say, young
children literature and high for, e.g., philosophy talks. For illustration, consider (1).
(1) a. A teddy bear is standing on a skateboard in Times Square, thinking on how lonely it
appeared during COVID peaks.
b. A teddy bear is standing on a skateboard in Times Square, because it doesn’t have a
wallet to pay the cab.
      </p>
      <p>A graphic/textual splitter should presumably decide that the part in bold should be best express
by language, while the rest can be rendered as an image. Broadly speaking, when the modifier
is not imageable (“a sinful/nostalgic teddy bear”) or its relation with the rest is not transparent
(“Teddy is in Times Square because/when it wants/needs to see people”), it is likely to
be best rendered as text.13 Modifiers that embed disjunction or negation (“The teddy bear
or stufed dog that didn’t come yesterday reached Times Square instead of its brother,
despite having only a skateboard”) are essentially impossible to render as images alone.</p>
      <p>In other cases, the linguistic relation could be easily rendered by segmenting the text in
multiple images: before(/after) in (2) suggest ordering Fig. 2 before (respectively, after) Fig. 3. Or
in (2-b) suggests a composition of Fig. 2 and 3, and so forth.
(2) a. A teddy bear is standing on a skateboard in Times Square, before/after playing chemistry
with a friend.
b. A teddy bear is thinking of skateboarding in Times Square or playing chemistry with a
friend.
13This is not to say that it is impossible to add to the image elements that suggest nostalgia, or a specific reason.</p>
      <p>There is a ratio between making the information unambiguous and cluttering the image that must be carefully
balanced.</p>
      <p>
        As mentioned above, in the representation of a narrative the text needs to be segmented in a
sequence of images, and any image must be generated from the appropriate text segment. But
the image generation must be designed to preserve elements that should not vary from scene to
scene (the drawing style, the aspect and size of previously introduced objects). Solving image
consistency through time is a must for video generation, so it is a topic likely to receive a lot
of attention (recent attempts, like [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ], start from stills and learn movements from untagged
videos), but the problem of connecting the text to the individual entities or events shown in an
artificially generated image is specific to I+L. The techniques proposed in the Deep Captioning
(see Sec. 2) and the Image Paragraph literature are relevant for this point: a very detailed
and structurally explicit textual description would allows reference to the individual objects
introduced in the image, and make possible editing instructions like “Remove the skateboard
from Fig. 2”. It could also be used to avoid repeating in the text elements that are obvious from
the image.
      </p>
      <p>
        Fig. 4 illustrates a potential training path: a detailed textual description (b) is obtained from
an original image (a) via a pretrained I2T, then used by a pretrained T2I system to generate a
derived image (c) (at low-resolution, to keep the computation manageable; upscaling techniques
could be used for the final output, as in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]). The I2T phase could now be fine-tuned using
as error function the T-T distance between the description of the original image (b) and that
of the derived image (d). Alternatively (or concurrently), one could try to minimize a distance
metric (I-I) between the original image (a) and the derived one (d). This is obviously more
dificult, as the same textual description can generate very diferent images, but "diferent" does
not necessarily mean "incompatible": two pictures of the same object from diferent angles can
be graphically very diferent, yet serve the same purpose and be recognized as portraying “the
same” content. On the other hand, a nearly-identical framing of the same objects in diferent
styles can be disconcerting.
      </p>
      <p>
        Ruiz at al. [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] and Gal et al. [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] describe systems to generate the same object in diferent
context, by associating a small set of specific images to a unique token identifier (for instance,
in [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ], “[S]” is a unique identifier and the T2I is fine-tuned to associate “The [S] dog” to a
particular dog image, while preserving the association of “the dog” to a more diverse set of
dog images). The problem is how to do this starting from a single image, one which might
belong to the ground truth (the car that crashed in the narration of an actual accident), but
also be a decoration that the system has allucinated to fill up a realistic image (the ambulance,
the policemen). If the object has a familiar shape, one strategy might be to derived a set of 3D
views from it (see the technique in [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ], but without its very narrow domain) and use these
views to condition the T2I. One obvious limit of this approach is that it would require repeated
ifne-tuning of large T2I systems, as new objects appear and are picked up in the narration, so it
would be too slow for on-line deployment.
      </p>
      <p>
        It is interesting to consider here the diferences and similarities between humans, formal
representations of meaning and distribution-based representations, like all NN models. Humans
have no problem with the idea of persisting entities, be they physical objects or events, but
are also very good at forming general concepts or classes. Formal semantics, on the other
hand, has tackled the behavior of general classes of objects (types, or kinds in the terminology
of [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ] or [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ]) only after dealing with (quantified) token objects, and arguably with greater
dificulty. Distributional semantics has often been said to be more successful at capturing generic
knowledge (i.e. types; see [
        <xref ref-type="bibr" rid="ref40">40</xref>
        ]) rather than episodic events, or tokens [
        <xref ref-type="bibr" rid="ref41">41</xref>
        ]. The same seems to
apply in generation: T2I systems have no trouble generating individual tokens, but struggle
to generate the same token in a changing environment. This time, the problem is shared by
humans: a well-known dificulty for artists who start drawing comics is how to make each
character recognizable in diferent scenes.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Some linguistic issues</title>
      <p>Even the ability to neatly split what can be rendered as images and to render it consistently
across scenes is no guarantee that we are doing a service to communication, or attention. There
is a vast cognitive psychology literature on the relation between language and vision (see [42]
for a particular linguistic strand, and [43]), on how images afect attention and how languages
could shape perception. This literature should inform some of the issues touched above: what is
the best proportion of language and images for any given style (technical instruction, question
answering, reports, storytelling, etc.), and whether the text should be written as subtitles, voiced
over, voiced between the images or suspended as a balloon, comics-syle.14</p>
      <p>One important aspect is the possibility of training the I+L system to render graphically
features that are normally rendered by language alone. Consider for instance repeated entity
reference. Language is fond of introducing entity in the discourse, then go back to them using
pronouns or definite description; interpreting chains of reference is a complex issue in language
comprehension and production. A mini-discourse like (3) could be easily translated into a
correct sequence of pictures, but only if the referents of the boldfaced nouns are clear to the
system and can be made recognizable by the user.
(3) Two teddy bears are at a picnic table. One decides to chase a butterfly. The second one
runs after him. A third teddy bear occupies the picnic table left empty by them.
Suppose that the initial entities, two teddy bears, are first introduced graphically. Referring
back to them in the subsequent images may be done verbally, using a normal pronoun (they) or
a definite description ( the teddy bears). In a I+L system, however, these linguistic devices could
also be supplemented with arrows or bounding boxes, by making the referent flash or unfocus
the background — useful signals if we are dealing with many objects of the same type, or if it is
unclear how to designate an object that has only been introduced graphically (“the thing on the
left”).</p>
      <p>This is of course not the only way to put graphics to linguistic use. Another possibility is to
render the distinction between the true information that the I+L system is trying to convey and
any additional material. Consider for instance (4)
(4) Students were protesting against a judge.</p>
      <p>Here a judge may be a very specific one (part of the ground truth: the face might be recognizable,
for instance); the bare plural students is non-specific, so the system would probably draw a
random assortment of persons. But the system would also be entitled to draw a background
(the place where the students were protesting, maybe some journalists). These additions are
reasonable but entirely made up, and not in the prompt. Do we want to distinguish graphically
what now looks like a three-way distinction in determination? (specific type and token: the
judge; specific type, not token: the students; non specific type or token: everything else). If the
answer is positive, the distinction must be carried out consistently in all subsequent scenes.
Once again, we do not have a dataset of generated images where objects that were explicitly
mentioned in the prompt can be visually distinguished from object that have been added by the
generation process, but creating it does not seem to be an impossible task. Going back to the
loop in Fig. 4, it should be possible to start from a prompt, generate an image from it (with T2I),
then caption it (with I2T). The goal is now to have a captioning system that separates what it
guesses to have been in the prompt from what has been added (a discriminator). With this in
14It is worth noting that voice-over narration is not very popular is modern cinematography, possibly on the account
that it breaks the narrative identification. Still, narrative identification may not be a desideratum in many uses of
I+L dialogue.
place, the T2I system could be fine-tuned to introduce changes in the image which keep the
image recognizable but increase the efectiveness of the discriminator. The desired result would
be images where aspects or object that have been hallucinated are recognizable to the naked
eye.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>This piece is a call to arms for a new AI research task, the generation of mixed languages
and images in dialogue, which is starting to become possible due to advances in neighboring
tasks: text-to-text, image-to-image, text-to-image, image-to-text (in captioning),
text+image-toimage (in "guided" language generation), text+image-to-text (in VQA). The missing element,
text(+image)-to-text+image, is not absent by chance. While potentially useful (we do use images
in public presentations, we do have illustrated books and comics), it is a combination that our
species cannot produce at "on-line" speed: a speed suficient for dialogue. The consequence is
that none of the many datasets accumulated by AI researchers can be directly used to train it,
though some (instructional videos, comics, film+storyline) could be the base to create useful
data. I have suggested that the output of text-to-text generation system (or more image-ready
language variants) could be pre-processed in a way that separates those parts that can be
converted into images or image sequences and those that need to remain voice or text. This
could be carried out, in part, using hand-picked linguistic features, but could maybe also be
derived from human-annotated data. The mix of language and images that would result raises a
number of dificult but interesting linguistic and psychological questions — in part as a result of
entering the territory of superhuman abilities.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Acknowledgments</title>
      <p>The authors wishes to acknowledge Rafaella Bernardi, Paolo Rota and other CIMeC colleagues
for stimulating conversations, and the audience of the EviL seminar series plus two anonymous
TACL reviewers for feedback on early versions of these ideas.
Natural Language Processing, Association for Computational Linguistics, Lisbon, Portugal,
2015, pp. 22–32. URL: https://aclanthology.org/D15-1003. doi:10.18653/v1/D15-1003.
[42] R. Jackendof, On beyond zebra, Cognition 26 (1987) 89–114.
[43] M. Vulchanova, V. Vulchanov, I. Fritz, E. A. Milburn, Language and perception: Introduction
to the special issue "speakers and listeners in the visual world"., J. Cult Cogn Sci. 3 (2019)
103–112. doi:https://doi.org/10.1007/s41809-019-00047-z.
[44] J. Hankamer, I. Sag, Deep and surface anaphora, Linguistic Inquiry 7.3 (1976) 391–426.
[45] L. A. Friedman, Space, time, and person reference in american sign language., Language
51 (1975) 940–61. doi:https://doi.org/10.2307/412702.
[46] K. van Hoek, Conceptual spaces and pronominal reference in american sign language,
Nordic Journal of Linguistics 15 (1992) 183–199. doi:10.1017/S0332586500002596.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bassignana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramponi</surname>
          </string-name>
          , Preface to the
          <source>Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI)</source>
          ,
          <source>in: Proceedings of the Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI</source>
          <year>2023</year>
          )
          <article-title>co-located with 22th International Conference of the Italian Association for Artificial Intelligence (AI* IA</article-title>
          <year>2023</year>
          ),
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. F.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mattis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Mesnar</surname>
          </string-name>
          ,
          <article-title>Graphics and natural language as components of automatic explanation</article-title>
          ,
          <source>SIGCHI Bull</source>
          .
          <volume>20</volume>
          (
          <year>1988</year>
          )
          <article-title>76</article-title>
          . URL: https://doi.org/10.1145/49103.1046410. doi:
          <volume>10</volume>
          . 1145/49103.1046410.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Theune</surname>
          </string-name>
          , I. van der Sluis,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bachvarova</surname>
          </string-name>
          , E. André,
          <source>The AISB'08 Symposium on Multimodal Output Generation (MOG</source>
          <year>2008</year>
          ),
          <source>The Society for the Study of Artificial Intelligence and Simulation of Behaviour (AISB)</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>iii</fpage>
          -
          <lpage>iv</lpage>
          .
          <source>AISB Symposium on Multimodal Output Generation, MOG</source>
          <year>2008</year>
          , MOG ; Conference date:
          <fpage>03</fpage>
          -
          <lpage>04</lpage>
          -2008 Through 04-
          <fpage>04</fpage>
          -
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wasinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Spassova</surname>
          </string-name>
          ,
          <article-title>Three output planning strategies for use in context-aware computing scenarios (</article-title>
          <year>2008</year>
          )
          <fpage>46</fpage>
          -
          <lpage>49</lpage>
          . AISB Symposium on Multimodal Output Generation,
          <string-name>
            <surname>MOG</surname>
          </string-name>
          <year>2008</year>
          , MOG ; Conference date:
          <fpage>03</fpage>
          -
          <lpage>04</lpage>
          -2008 Through 04-
          <fpage>04</fpage>
          -
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Rieser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lemon</surname>
          </string-name>
          ,
          <article-title>Learning efective multimodal dialogue strategies from Wizard-of-Oz data: Bootstrapping and evaluation</article-title>
          ,
          <source>in: Proceedings of ACL-08: HLT</source>
          , Association for Computational Linguistics, Columbus, Ohio,
          <year>2008</year>
          , pp.
          <fpage>638</fpage>
          -
          <lpage>646</lpage>
          . URL: https://aclanthology. org/P08-1073.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ramesh</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hesse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , E. Sigler,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Litwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chess</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
          </string-name>
          ,
          <article-title>Language models are few-shot learners</article-title>
          , arXiv
          <year>2005</year>
          .
          <volume>14165</volume>
          ,
          <year>2020</year>
          . arXiv:
          <year>2005</year>
          .14165.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chowdhery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bosma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Barham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. W.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gehrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schuh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tsvyashchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Maynez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Barnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Prabhakaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Reif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hutchinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pope</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bradbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Austin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Isard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gur-Ari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Duke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Levskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghemawat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Michalewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Misra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fedus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ippolito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Spiridonov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sepassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Omernick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Pillai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pellat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lewkowycz</surname>
          </string-name>
          , E. Moreira,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Polozov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Saeta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Diaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Firat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Catasta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Meier-Hellstern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Eck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petrov</surname>
          </string-name>
          , ,
          <string-name>
            <surname>N. Fiedel.</surname>
          </string-name>
          , Palm:
          <article-title>Scaling language modeling with pathways</article-title>
          .,
          <year>2022</year>
          . ArXiv:
          <year>2001</year>
          .08361
          <string-name>
            <surname>and</surname>
          </string-name>
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Roller,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Artetxe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dewan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Diab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. V.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mihaylov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shleifer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shuster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Simig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Koura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sridhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , Opt: Open pre-trained
          <source>transformer language models</source>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2205.01068. doi:
          <volume>10</volume>
          .48550/ARXIV.2205.01068.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nichol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McGrew</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Chen, GLIDE: towards photorealistic image generation and editing with text-guided difusion models</article-title>
          ,
          <source>CoRR abs/2112</source>
          .10741 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2112.10741. arXiv:
          <volume>2112</volume>
          .
          <fpage>10741</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nichol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Hierarchical text-conditional image generation with clip latents</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2204.06125. doi:
          <volume>10</volume>
          .48550/ ARXIV.2204.06125.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>CIGLI: conditional image generation from language &amp; image</article-title>
          ,
          <source>CoRR abs/2108</source>
          .08955 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2108.08955. arXiv:
          <volume>2108</volume>
          .
          <fpage>08955</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tensmeyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Interactive image generation with natural-language feedback</article-title>
          , PREPRINT: https://www.aaai.org/AAAI22Papers/ AAAI-7081.ZhouY.pdf,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kawar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Tov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dekel</surname>
          </string-name>
          , I. Mosseri,
          <string-name>
            <given-names>M.</given-names>
            <surname>Irani</surname>
          </string-name>
          , Imagic:
          <article-title>Text-based real image editing with difusion models</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2210</volume>
          .
          <fpage>09276</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lu</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kortylewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Theobalt</surname>
          </string-name>
          , E. Xing,
          <article-title>Multimodal image synthesis and editing: A survey and taxonomy</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          . doi:
          <volume>10</volume>
          .1109/TPAMI.
          <year>2023</year>
          .
          <volume>3305243</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          ,
          <source>in: 2009 IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR.
          <year>2009</year>
          .
          <volume>5206848</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>T.-Y. Lin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Maire</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Belongie</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bourdev</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Hays</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Perona</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ramanan</surname>
            ,
            <given-names>C. L.</given-names>
          </string-name>
          <string-name>
            <surname>Zitnick</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dollár</surname>
          </string-name>
          , Microsoft coco: Common objects in context,
          <year>2014</year>
          . URL: https: //arxiv.org/abs/1405.0312. doi:
          <volume>10</volume>
          .48550/ARXIV.1405.0312.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Krueger</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <year>2021</year>
          . URL: https://arxiv.org/abs/2103.00020. doi:
          <volume>10</volume>
          .48550/ ARXIV.2103.00020.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bansal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.-W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Keutzer</surname>
          </string-name>
          ,
          <article-title>How much can clip benefit vision-and-language tasks</article-title>
          ?,
          <year>2021</year>
          . URL: https://arxiv.org/abs/2107.06383. doi:
          <volume>10</volume>
          .48550/ARXIV.2107.06383.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ghandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Pourreza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mahyar</surname>
          </string-name>
          ,
          <article-title>Deep learning approaches on image captioning: A review</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2201.12944. doi:
          <volume>10</volume>
          .48550/ARXIV.2201.12944.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Saharia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Saxena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Whang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Denton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K. S.</given-names>
            <surname>Ghasemipour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. K.</given-names>
            <surname>Ayan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Mahdavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. G.</given-names>
            <surname>Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salimans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Fleet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Norouzi</surname>
          </string-name>
          ,
          <article-title>Photorealistic text-to-image difusion models with deep language understanding</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv. org/abs/2205.11487.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karpathy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <article-title>Densecap: Fully convolutional localization networks for dense captioning</article-title>
          ,
          <year>2015</year>
          . URL: https://arxiv.org/abs/1511.07571. doi:
          <volume>10</volume>
          .48550/ARXIV. 1511.07571.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Krause</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Krishna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <article-title>A hierarchical approach for generating descriptive image paragraphs, in: Computer Vision and Patterm Recognition (CVPR</article-title>
          ),
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Fernandes</surname>
          </string-name>
          .,
          <string-name>
            <given-names>M. H. F.</given-names>
            <surname>Ribeiro.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. R.</given-names>
            <surname>Cerqueira</surname>
          </string-name>
          .,
          <string-name>
            <given-names>M. M.</given-names>
            <surname>Silva</surname>
          </string-name>
          .,
          <article-title>Describing image focused in cognitive and visual details for visually impaired people: An approach to generating inclusive paragraphs</article-title>
          ,
          <source>in: Proceedings of the 17th International Joint Conference on Computer Vision</source>
          , Imaging and Computer Graphics Theory and Applications - Volume
          <volume>5</volume>
          : VISAPP„ INSTICC, SciTePress,
          <year>2022</year>
          , pp.
          <fpage>526</fpage>
          -
          <lpage>534</lpage>
          . doi:
          <volume>10</volume>
          .5220/0010845700003124.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>You</surname>
          </string-name>
          ,
          <article-title>Colossal-ai: A unified deep learning system for large-scale parallel training</article-title>
          ,
          <year>2021</year>
          . URL: https://arxiv.org/abs/ 2110.14883. doi:
          <volume>10</volume>
          .48550/ARXIV.2110.14883.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>S.</given-names>
            <surname>Antol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Zitnick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <article-title>VQA: visual question answering</article-title>
          ,
          <source>CoRR abs/1505</source>
          .00468 (
          <year>2015</year>
          ). URL: http://arxiv.org/abs/1505.00468. arXiv:
          <volume>1505</volume>
          .
          <fpage>00468</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pezzelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Greco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gandolfi</surname>
          </string-name>
          , E. Gualdoni,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bernardi</surname>
          </string-name>
          , Be Diferent to Be Better!
          <article-title>A Benchmark to Leverage the Complementarity of Language and Vision, in: Findings of the Association for Computational Linguistics: EMNLP 2020, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>2751</fpage>
          -
          <lpage>2767</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          . ifndings-emnlp.
          <volume>248</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .findings-emnlp.
          <volume>248</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>B.</given-names>
            <surname>Mildenhall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Srinivasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tancik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Barron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ramamoorthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ng</surname>
          </string-name>
          , Nerf:
          <article-title>Representing scenes as neural radiance fields for view synthesis</article-title>
          , CoRR abs/
          <year>2003</year>
          .08934 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2003</year>
          .08934. arXiv:
          <year>2003</year>
          .08934.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Maharjan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P.</given-names>
            <surname>López-Monroy</surname>
          </string-name>
          , T. Solorio,
          <article-title>MPST: A corpus of movie plot synopses with tags</article-title>
          ,
          <source>in: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ),
          <article-title>European Language Resources Association (ELRA), Miyazaki</article-title>
          , Japan,
          <year>2018</year>
          . URL: https://aclanthology.org/L18-1274.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>R.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Babaeizadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.-J.</given-names>
            <surname>Kindermans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Moraldo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , M. T. Safar,
          <string-name>
            <given-names>S.</given-names>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kunze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Erhan</surname>
          </string-name>
          , Phenaki:
          <article-title>Variable length video generation from open domain textual descriptions</article-title>
          ,
          <source>in: International Conference on Learning Representations</source>
          ,
          <year>2023</year>
          . URL: https: //openreview.net/forum?id=vOEXS39nOF.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>T.</given-names>
            <surname>Leitch</surname>
          </string-name>
          , Film Adaptation and
          <article-title>Its Discontents: From Gone with the Wind to The Passion of the Christ</article-title>
          .
          <year>2007</year>
          . „
          <string-name>
            <surname>Project</surname>
            <given-names>MUSE</given-names>
          </string-name>
          , Johns Hopkins University Press,
          <year>2007</year>
          . doi:
          <volume>10</volume>
          .1353/ book.3302.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lu</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kortylewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Theobalt</surname>
          </string-name>
          , E. Xing,
          <article-title>Multimodal image synthesis and editing: The generative ai era</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2112</volume>
          .
          <fpage>13592</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>K.</given-names>
            <surname>Stenning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oberlander</surname>
          </string-name>
          ,
          <article-title>A cognitive theory of graphical and linguistic reasoning: Logic and implementation</article-title>
          ,
          <source>Cognitive Science 19</source>
          (
          <year>1995</year>
          )
          <fpage>97</fpage>
          -
          <lpage>140</lpage>
          . doi:https://doi.org/10. 1207/s15516709cog1901_
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Jampani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Pritch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rubinstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Aberman</surname>
          </string-name>
          , Dreambooth:
          <article-title>Fine tuning text-to-image difusion models for subject-driven generation</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/ abs/2208.12242. doi:
          <volume>10</volume>
          .48550/ARXIV.2208.12242.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ljubešić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fišer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Peti-Stantić</surname>
          </string-name>
          ,
          <article-title>Predicting concreteness and imageability of words within and across languages via word embeddings</article-title>
          ,
          <source>in: Proceedings of The Third Workshop on Representation Learning for NLP, Association for Computational Linguistics</source>
          , Melbourne, Australia,
          <year>2018</year>
          , pp.
          <fpage>217</fpage>
          -
          <lpage>222</lpage>
          . URL: https://aclanthology.org/W18-3028. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W18</fpage>
          -3028.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>U.</given-names>
            <surname>Singer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polyak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hayes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Ashual</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Gafni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Taigman</surname>
          </string-name>
          ,
          <article-title>Make-a-video: Text-to-video generation without text-video data</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2209.14792. doi:
          <volume>10</volume>
          .48550/ARXIV.2209.14792.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Alaluf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Atzmon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Patashnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Bermano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chechik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cohen-Or</surname>
          </string-name>
          ,
          <article-title>An image is worth one word: Personalizing text-to-image generation using textual inversion</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2208.01618. doi:
          <volume>10</volume>
          .48550/ARXIV.2208.01618.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>N.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Simonelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Porzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Bulò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nießner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kontschieder</surname>
          </string-name>
          ,
          <article-title>Autorf: Learning 3d object radiance fields from single view observations</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/ 2204.03593. doi:
          <volume>10</volume>
          .48550/ARXIV.2204.03593.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lawler</surname>
          </string-name>
          , Studies in English Generics, University of Michigan Papers in Linguistics 1 (
          <year>1973</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>G.</given-names>
            <surname>Carlson</surname>
          </string-name>
          ,
          <article-title>A unified analysis of the English bare plural</article-title>
          ,
          <source>Linguistics and Philosophy</source>
          <volume>1</volume>
          (
          <year>1977</year>
          )
          <fpage>413</fpage>
          -
          <lpage>457</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>L.</given-names>
            <surname>McNally</surname>
          </string-name>
          , Kinds, descriptions of kinds, concepts,
          <source>and distributions</source>
          ,
          <year>2015</year>
          .
          <article-title>Submitted to a volume of papers from the BRIDGE-</article-title>
          14 Workshop.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>A.</given-names>
            <surname>Herbelot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Vecchi</surname>
          </string-name>
          ,
          <article-title>Building a shared world: mapping distributional to modeltheoretic semantic spaces</article-title>
          ,
          <source>in: Proceedings of the 2015 Conference on Empirical Methods in</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>