<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>for Wikidata Image Completion</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Raia Abu Ahmad</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Critelli</string-name>
          <email>martin.critelli@unive.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Şefika</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Efeoğlu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eleonora Mancini</string-name>
          <email>e.mancini@unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Célian Ringwald</string-name>
          <email>celian.ringwald@inria.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xinyue Zhang</string-name>
          <email>xinyue.zhang@cs.ox.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff8">8</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Albert Meroño-Peñuela</string-name>
          <email>albert.merono@kcl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Generative AI, Image Generation, Automated Prompt Generation</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>(R. Abu Ahmad)</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI GmbH)</institution>
          ,
          <addr-line>Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Freie Universität Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>King's College London</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Technische Universität Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Université Côte d'Azur</institution>
          ,
          <addr-line>Inria, CNRS, I3S</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University Ca'Foscari of Venice</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff7">
          <label>7</label>
          <institution>University of Bologna</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff8">
          <label>8</label>
          <institution>University of Oxford</institution>
          ,
          <addr-line>Oxford</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Humans are critical for the creation and maintenance of high-quality Knowledge Graphs (KGs). However, creating and maintaining large KGs only with humans does not scale, especially for contributions based on multimedia (e.g. images) that are hard to find and reuse on the Web and expensive to generate by humans from scratch. Therefore, we leverage generative AI for the task of creating images for Wikidata items that do not have them. Our approach uses knowledge contained in Wikidata triples of items describing ifctional characters and uses the fine-tuned T5 model based on the WDV dataset to generate natural text descriptions of items about fictional characters with missing images. We use those natural text descriptions as prompts for a transformer-based text-to-image model, Stable Difusion v2.1, to generate plausible candidate images for Wikidata image completion. We design and implement quantitative and qualitative approaches to evaluate the plausibility of our methods, which include conducting a survey to assess the quality of the generated images.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Large knowledge bases (KBs) such as Wikidata are maintained by human editors in a
collaborative manner in order to provide structured data of high quality [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, given the size of
this platform, there is an evident problem of incompleteness that creates several content gaps
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We note that this is especially true for contributions based on multimedia (such as images,
nEvelop-O
CEUR
Workshop
Proceedings
audio, and video) since it is dificult for editors to find such high-quality contributions on the
Web, and even more dificult and expensive to create them from scratch.
      </p>
      <p>
        In this work, we examine the problem of missing images for a specific class of Wikidata entities:
ifctional characters . We motivate this choice by the fact that querying Wikidata shows that
only 7% out of the 83.7K instances of the fictional character class, including its sub-classes, have
an image 1. It is important to note that this class was specifically chosen due to ethical and
privacy concerns, as other classes (e.g. person) can have detrimental consequences if automatic
images are created to portray them. Although alternative methods for finding images for this
class of entities exist (e.g. fan-created images), they are unreliable in terms of the objective
representation of characters and demand thorough and manual research by editors to make
sure that they correctly align with Wikidata’s knowledge about each entity.
We propose a novel method of leveraging knowledge from Wikidata triples about each fictional
character entity in order to create a representative image for it using generative artificial
intelligence (AI) models. This is done by (1) extracting triples from Wikidata entities, (2) creating
English prompts to be fed into a generative text-to-image model such as Stable Difusion [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
and (3) generating a representative image of the character that could potentially be used on
Wikidata. We investigate the efectiveness of this approach by generating four diferent types
of prompts in English, including using triple verbalisation with large language models (LLMs),
for each character and comparing the resulting images. We evaluate our approach based on
a ground-truth dataset that consists of fictional characters which already have an image on
Wikidata. We select diferent metrics of automatic image comparison to measure how similar
each generated image is to the ground-truth one. Additionally, since automatic measures
for image comparison are limited, we conduct a human evaluation survey in which we ask
participants to evaluate image similarity.
      </p>
      <p>Our work addresses the following research questions (RQs):
• RQ1: To what extent can diferent types of prompts based on triples be used in
text-toimage models to produce high-quality images?
• RQ2: To what extent can the output of generative AI be used for Wikidata image
completion?
• RQ3: How can generative text-to-image models be evaluated?
To the best of our knowledge, no previous study has explored the realm of using Wikidata
as a source for creating prompts for generative text-to-image models. Our work 2 ofers the
following contributions:
• A framework that generates prompts for a text-to-image model (Stable Difusion v2.1 3)
with diferent levels of structure and natural language text based on Wikidata triples.
• A dataset of generated images for fictional characters extracted from Wikidata that can
potentially be used by editors for image completion.
• An evaluation strategy showing evidence of relevancy and adequacy of using AI-generated
images for our use case.
1This query was performed in June 2023.</p>
      <p>2The project and dataset are available at https://github.com/helemanc/gryffindor and at https://huggingface.co/
gryffindor-ISWS, respectively.</p>
      <p>3The model card is at https://huggingface.co/stabilityai/stable-diffusion-2-1</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related Work</title>
      <p>
        Generative AI: Current groundbreaking advances in AI enable machines to generate novel
and original content based on textual prompts. Such generative applications include
text-totext [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], text-to-image [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ], and even text-to-music [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Generally, these models can capture
complex patterns from the input text and produce coherent outputs. A recent survey [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] shows
that text-to-image applications specifically have been emerging since 2015, when AlignDRAW
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] pioneered the field by leveraging recurrent neural networks (RNNs) to encode textual
captions and produce corresponding images. Since then, end-to-end models started leveraging
architectures such as deep convolutional generative adversarial networks (GANs) [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10, 11, 12</xref>
        ],
autoregressive methods [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">13, 14, 15</xref>
        ], latent space models [
        <xref ref-type="bibr" rid="ref16 ref17 ref18">16, 17, 18</xref>
        ], and the current
state-ofthe-art difusion-based methods [
        <xref ref-type="bibr" rid="ref19 ref20 ref21">19, 20, 21</xref>
        ].
      </p>
      <p>
        Prompt Engineering: Because of the aforementioned advances, a novel area of prompt
engineering has emerged, in which humans interact with AI in an iterative process to produce
the best prompt (i.e. textual input) for a specific desired output [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Recent work has shed light
on prompt engineering for AI art specifically [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], concluding that simple and intuitive prompts
written by humans are not enough to get desired results. Rather, writing good prompts is a
learned skill that is enhanced by the usage of specific prompt templates 4 and modifiers [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
Automatic Prompt Generation: When it comes to automatic prompt generation, previous
studies tend to investigate using LLMs to construct prompts using techniques such as text
mining, text paraphrasing, and data augmentation [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. However, to the best of our knowledge,
no work has touched upon using large KBs such as Wikidata for prompt engineering and
generation.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Proposed Approach</title>
      <p>
        We conduct our study on instances of the class designated as fictional character with Q95074
item ID. The initial stage of our approach involves the extraction of relevant triples pertaining
to a specific entity. Subsequently, these triples are used to form various types of prompts in
English, functioning as inputs to Stable Difusion [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], a text-to-image AI model. We generate
diferent types of prompts related to the triples, including a verbalised triples prompt which uses
the T5 language model fine-tuned on the WDV dataset [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. This verbalisation model converts
triples into fluent language [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. The ultimate goal is to generate suitable images that can serve
as accurate visual identifiers for their corresponding Wikidata entities. This pipeline is shown
in Fig. 1.
      </p>
      <sec id="sec-4-1">
        <title>3.1. Triple Extraction</title>
        <p>Generating an image for a specific character requires a description that can be gathered from its
related triples. In Wikidata, these can be obtained through SPARQL queries 5, yielding all triples
with the character as the subject. Moreover, since properties and entities might be represented
4https://sweet-hall-e72.notion.site/A-Traveler-s-Guide-to-the-Latent-Space-85efba7e5e6a40e5bd3cae980f30235f
5All the SPARQL queries with detailed explanations are available at https://github.com/helemanc/gryffindor/
blob/main/src/data-collection/wiki_query_service.py.
by item IDs and property IDs, we also extract and translate each triple from these IDs to their
corresponding labels when the triple has an entity and property ID.</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Prompt Generation</title>
        <p>In order to investigate if prompts closer to natural language work better at generating better
images, we generate four distinct prompts for each character. The first three are created utilising
the set of triples that have been extracted for the respective entity from Wikidata, while the last
prompt utilises English DBpedia abstracts. These prompts are defined as follows (see Table 15
for prompt samples of an entity):
1. Basic Label: This prompt merely employs the “label” that Wikidata assigns to its entities.
2. Plain Triples: This prompt is derived by concatenating the subject, predicate, and object
of a triple to form a single sentence, utilising all available triples linked to a specific entity.</p>
        <p>
          Notably, sentences generated from plain triples may lack proper structure and grammar.
3. Verbalised Triples: Triple verbalisation is defined as the transformation of structured
data (i.e., triples) into human-readable formats (i.e., text). These serve as a summarised
paragraph of all input triples.
4. DBpedia Abstracts: We use DBpedia abstracts as prompts obtained by querying the
English chapter of DBpedia [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. Originally written by human editors on Wikipedia,
these abstracts are automatically extracted by DBpedia, preprocessed, and shortened.
Unlike previous prompt types, this is the only one originally written by a human in
natural language.
        </p>
        <p>When examining the triples for a single entity, we observe that triples sharing the same predicate
tend to contain redundant information. As a result, prompts generated directly from these
plain or verbalised triples will repetitively state the same facts. However, the “instance of”
predicate seems to provide distinct information for each triple. To avoid duplicating facts in
prompt types (2) and (3), we remove duplicate predicates, except for “instance of”, for the input
triples. Among the remaining triples that share the same predicate, we keep only the one with
the longest object, since longer objects likely contain more detailed information than shorter
objects with the same predicate.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Image Generation</title>
        <p>
          To ensure reproducibility in image generation, we utilise Stable Difusion version 2.1 6, an
open-source text-to-image model developed by Stability AI limited to the English language.
We chose version 2.1 because it supports all input shapes up to 1024x1024 and has a better
performance according to benchmark evaluation results [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ].
        </p>
        <p>It is important to note that this particular model has inherent limitations when it comes to
generating images related to the human body. To address this issue and enhance its image
generation capabilities, we employ the implementation of negative prompts that have been
suggested and shared on a public GitHub repository 7. By incorporating these negative prompts,
we aim to mitigate malformations in images (e.g. crossed eyes, more than five fingers, etc.).
Moreover, since Stable Difusion has a limitation on the number of tokens allowed in the prompt
sentence(s), we embed the prompt by utilising the encoder and tokenizer from Stable Difusion,
courtesy of the Compel library 8. The model runs positive prompts of 1500 fictional characters
without existing images on Wikidata and 1500 with images on Wikidata, the latter to be used
for building a ground-truth dataset for evaluation.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Collected Dataset</title>
      <p>We construct an extensive dataset 9 comprising 1500 fictional characters with images, as well
as 1500 fictional characters without images, which are randomly chosen from the entire set
of fictional characters on Wikidata. Our motivation for collecting data on fictional characters
rather than real people lies in our commitment to upholding ethical standards and safeguarding
privacy. Also, there is no available dataset about fictional characters, and our data collection
source codes 10 can be easily applied to diferent domains by changing parameter settings.
In addition, we extend our data by fetching the Wikipedia abstracts of each fictional character
from the English chapter of DBpedia. Although a majority of these fictional characters lack
information in DBpedia because it is constructed using English Wikipedia, this is not a problem
in our case since the Stable Difusion model can only use English text as input. Since most of
the fictional characters on Wikidata (ca. 78% 11) do not have any English Wikipedia page, we
only managed to gather DBpedia abstracts for 925 fictional characters with images on Wikidata
and 341 fictional characters without images (see Table 1).</p>
      <p>By analysing basic statistics from Table 1, we directly notice a big descriptive gap in terms of
triples, the number of extracted unique relations, and the length of the prompts between the
two datasets we constructed. Moreover, we notice that the length of the prompt is usually the
shortest for verbalised triples and the longest for plain triples. After gathering the data about
the fictional characters from Wikidata and DBpedia, four diferent prompts are automatically
6Stable Difusion v2.1 model card: https://huggingface.co/stabilityai/stable-diffusion-2-1
7The negative prompts: https://github.com/mikhail-bot/stable-diffusion-negative-prompts
8Compel encodes and decodes the portion of the prompt and available at https://github.com/damian0815/compel
9The entire dataset is available at https://huggingface.co/gryffindor-ISWS
10The data collection codes are available at https://github.com/helemanc/gryffindor/blob/main/src/
data-collection/wiki_query_service.py</p>
      <p>11The percentage is computed on Sept. 6, 2023
constructed by using the approach described in section 3. One example is shown in Figure 2,
which depicts the ground-truth image for the character Harlequin with its four generated
images. Based on this example, it is instantly clear that some of the prompts can produce images
more similar to the ground truth.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Evaluation</title>
      <p>In order to understand whether the generated images can plausibly be used for representing
ifctional characters based on their Wikidata triples, we employ two evaluation strategies. The
ifrst is an automatic evaluation of image similarity using diferent metrics, while the second is a
human evaluation survey. Since the task of identifying whether two images portray the same
character is subjective and dificult, we consider both qualitative and quantitative evaluation
approaches. This helps us better understand the efect of the prompt type on the quality of the
diferent generated images. In this section, we first describe the evaluation framework used,
explaining the diferent metrics we took into account. Then, we present the obtained results.</p>
      <sec id="sec-6-1">
        <title>5.1. Evaluation Framework</title>
        <sec id="sec-6-1-1">
          <title>5.1.1. Automatic evaluation</title>
          <p>
            We utilise automated evaluation methods based on three image comparison metrics:
• UQI [
            <xref ref-type="bibr" rid="ref30">30</xref>
            ]: computes a pixel-based similarity score by comparing generated images with
their corresponding ground-truth images. Notably, since the majority of the original
images are in grayscale, the similarity computation also takes into account their grayscale
versions. UQI evaluates “image quality based on factors such as loss of correlation,
luminance distortion, and contrast distortion” [
            <xref ref-type="bibr" rid="ref30">30</xref>
            ].
• CLIPscore [
            <xref ref-type="bibr" rid="ref31">31</xref>
            ]: leverages embeddings produced by a contrastive language-image
pretrained model [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ]. It is used for measuring image-caption compatibility by comparing
image and text embeddings using cosine similarity. CLIP embeddings can be used for
image-to-image comparisons as well, which we did by using the image encoder of the
CLIP-Visual Transformer model [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ]: ViT-L/14.
• FID [
            <xref ref-type="bibr" rid="ref33">33</xref>
            ]: is an improved version of Inception Distance (IS) proposed to measure the
quality of images produced by generative models.
          </p>
          <p>Since the computation of the FID metric is more time-consuming than the other two, we compute
it only on a small subset of our dataset consisting of images generated for ten random characters.
On the other hand, UQI and CLIPscores are computed on the entire dataset.
Additionally, we employ statistical methods for evaluating if the metrics above can measure
the impact of the prompt on the quality of the generated images. For this purpose, we perform
ANOVA to measure the efect of the prompt on the metric. We also perform Tukey’s HSD
(honestly significant diference) tests on the metrics to reflect the prompts’ efect on the generated
images. These statistical methods were computed on two subsets of our dataset: characters that
have DBpedia abstracts and characters that do not.</p>
          <p>Finally, we performed several Student tests to evaluate if a given property (e.g., the gender and
occupation of a character) could lead to better results, and we separately made the test only on
the values of the instance of property (P31). To carry out these tests we extract the types and
properties used more than 100 times. For each property, we build two subsets. The first one
includes evaluation metric results of the characters that contain the property, and the second is
built by randomly choosing characters that do not have the evaluated property.</p>
        </sec>
        <sec id="sec-6-1-2">
          <title>5.1.2. Human evaluation</title>
          <p>
            Although the above-mentioned evaluation metrics can provide automatic measures to compare
images, they are still unreliable in comparing whether the generated images successfully portray
the same characters as the ground-truth images. This is because noise such as the image style
or its color can afect the results of the automatic metrics. Therefore, we conduct a human
evaluation study in which we ask participants to rate how likely it is that a pair of images
(consisting of 1. the ground-truth image and 2. the generated image) portray the same character.
Additionally, we ask participants to list the criteria they think about when comparing two
images. The latter was done to get an idea of important features to look for when generating
images of fictional characters. For evaluating the agreement of the participants we compute
Krippendorf’s Alpha [
            <xref ref-type="bibr" rid="ref34">34</xref>
            ] on three levels: globally, per evaluated image, and per prompt type.
          </p>
        </sec>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Evaluation Results</title>
        <sec id="sec-6-2-1">
          <title>5.2.1. Automatic Evaluation</title>
          <p>The results we obtained from the automatic evaluation metrics show diferent outcomes. In
terms of UQI, all four prompt types yield a similar average similarity score of ca. 0.5, concluding
that this metric is not optimal for our purposes. FID results (Table 5) show that images created
from DBpedia are more similar to the ground-truth. When it comes to CLIPscores, we see
a hierarchy of prompt types in terms of the obtained average similarity scores, with images
generated by basic labels being the least similar (with a score of 0.48), followed by plain triples
(with a score of 0.55), verbalised triples (with a score of 0.56), and lastly DBpedia abstract
prompt seem to generate the most similar images to the ground-truth with a CLIPscore of 0.6.
Results for UQI ad CLIP are shown in Table 4. It is important to note that contrary to UQI and
CLIPscores, the FID metric is performed only on a subset of images generated for ten fictional
characters, which makes it hard to make any concrete conclusions about this method.
The ANOVA conducted on the UQI and the CLIPscores is shown in Table 6 and Table 7. The
results show that the UQI is not able to underline a significant diference between the prompt
type as a main fixed efect on the quality of the generated image. In contrast, CLIPscores are
able to reflect this efect with high confidence.</p>
          <p>The results of Tukey’s HSD test are shown in Tables 8 and 9. They highlight that the basic
prompt is generally the worst prompt strategy and that the DBpedia abstract prompt is always
the best one. However, for characters that do not have a DBpedia abstract, the verbalised triples
prompt is better than the basic label and the plain triples prompts (with a p-value of 0.05810).
Additionally, in order to understand if the number of relations and unique relation type attached
to a given entity in extracted triples has an efect on the generated image quality, we compute
the correlation between these variables with the CLIPscores. Results indicate that there is no
such correlation (see Table 10).</p>
          <p>Finally, we present the results of the Student tests related to the efect of the values of the
instance of properties attached to an entity on the quality of generated images in two parts.
The first displays the efect of values of the instance of properties on the generated images
in Table 11 for plain triples prompts, and Table 13 for verbalised triples prompts. We can see
that characters that already have a widely known visual representation (e.g. characters from
comics, cartoons, or movies) generally have low CLIPscores. On the other hand, characters that
do not have a visual representation (e.g. from written works such as novels) are usually more
similar to the ground-truth images. The second part of the Student tests deals with the efect of
properties on the quality of the generated images. Results are shown in Table 12 and Table 14.
These results show that the majority of relations impact CLIPscores negatively.</p>
        </sec>
        <sec id="sec-6-2-2">
          <title>5.2.2. Human Evaluation</title>
          <p>To measure how humans evaluate the similarity of generated and ground-truth images, we ran
an evaluation survey in which each participant is presented with images of ten diferent fictional
characters chosen randomly from our dataset (shown in Figure 5). For each character, four pairs
of images were displayed. Each pair consisted of the ground-truth image and a generated image.
Participants were asked to rate how likely it is that both images portray the same character on
a scale of 1-5, 1 being very unlikely and 5 being very likely. Figure 3 shows the distribution of
participant replies for all ten characters based on prompt types. We immediately notice that
images generated based on the three prompt types of basic labels, plain triples, and verbalised
triples are more likely to be evaluated as not similar to the ground-truth image (i.e. the most
frequent response for all three prompt types is one). On the other hand, images generated with
DBpedia abstract prompts are most frequently rated as 3 and 4, both having the same number
of responses. When examining the high numbers on the scale that indicate a high similarity
between the ground truth and the generated images (i.e. 4 and 5), we notice a specific trend.
The least frequent prompt type for those numbers is the basic label, followed by plain triples,
verbalised triples, and DBpedia abstracts.
Additionally, we analyse participant responses to the open question of which criteria they
consider when giving their responses. Figure 4 presents the top ten criteria mentioned by
participants. The analysis was done by extracting nouns and adjectives, and filtering out stop
words and generic terms such as ‘character’. We also manually grouped synonymous concepts
such as clothes, clothing, and outfit .</p>
          <p>In total, our survey had 101 participants ranging between the ages of 17-59 with an average
age of 30. About 57% of participants were male, 41% female, and 2% non-binary. 48% of the
participants had a master’s education level. We did not target any specific group since we
wanted to receive general responses regarding the similarity of images. Thus, we distributed
the survey among friends and colleagues both from within and outside the research community.
The cultural backgrounds of participants ranged from South and North America (ca. 8%) to</p>
          <p>Europe (ca. 63%), East Asia (ca. 10%), and the Middle East (ca. 20%). That being said, ca. 25% of
participants were Italian. We received 37 responses for the open question of listing relevant
criteria.</p>
          <p>Finally, we measured the agreement of participants using Krippendorf’s alpha. The global
score is equal to 0.17, meaning that no concrete agreement was found. The same conclusion
could be drawn at the level of the images presented to the participants and at the level of the
prompt types used for the image generation (cf. Table 2 and Table 3).</p>
        </sec>
        <sec id="sec-6-2-3">
          <title>5.2.3. Automatic and Human Evaluation Alignment</title>
          <p>As a last step in our evaluation, we want to assess if there is an alignment between the score of
the automatic metrics (UQI, CLIPscore, and FID) and the human evaluation. In order to be able
to normalise the participant evaluations, we standardise the score given by each participant.
For  ∈ {1, .., 101} , the unique ID of a given participant, and  ∈ {0, 39} , a given generated image,
the standardised score is computed as follows:</p>
          <p>,
=
 , −  
 
The alignment between automatic metrics and human evaluation scores is shown in figure
6. We see that CLIPscores seem to be most correlated to the human scores with a Pearson
correlation of 0.5 for the plain triples prompt, 0.6 for the verbalised triples prompt, and 0.7 for
the basic label prompt. Concerning the DBpedia abstract prompt, none of the metrics seem to
be correlated with the human evaluation. UQI and FID are not correlated to human evaluation,
results, both having scores close to zero.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Discussion</title>
      <p>Results of most automatic evaluation approaches we used (CLIPscores, ANOVA, and Tukey’s
HSD) as well as the human evaluation results suggest a clear trend: images generated using
DBpedia abstracts as prompts were rated as most similar to the ground truth images, followed
respectively by verbalised triples prompts, plain triples prompts, and basic label prompts. This
implies that DBpedia abstracts, which are written by human editors and contain more natural,
diverse, and fluent text, enable text-to-image generators to produce better results. The fact that
verbalised triple prompts produce the second-best results further emphasises the importance of
lfuent text on the quality of the generated image. These results directly answer RQ1.
When further analysing the obtained CLIPscores, we observe that the maximum CLIPscore
occurred for an image generated by using a basic label prompt, possibly indicating that the
text-to-image model had “seen” this character during training. This enabled it to create a similar
image to the ground-truth one without adding any additional context. However, the lowest
CLIPscore also occurred when using the prompt type of basic label, further emphasising that
for some characters, generating an image based only on their label is not enough.
We conclude that in order to automatically generate images for fictional characters that correctly
portray them, using natural text descriptions is the best option. When this text is available
(e.g. in DBpedia abstracts), it is best to use it, however, as we have observed when creating our
dataset, many entities of fictional characters (See Table 1) do not have a DBpedia abstract. To
create images for those characters, the best method seems to be extracting knowledge about
them in the form of triples, verbalising those triples using a large language model, and giving
the verbalised text as input to a text-to-image generative model. In this case, the content of
the triples is crucial for generating high-quality images. However, the quality of images is not
related to the number of triples or the number of unique relations contained in the triples. But is
highly dependent on object values, highlighting the impact of the value of instance of property
of fictional characters on the quality of generated images. Answering RQ2, generated images
can then be leveraged for completing missing images in Wikidata entities.</p>
      <p>Finally, addressing RQ3, when comparing the three automatic evaluation metrics we see that
only the CLIPscores align with the human evaluation scores. This is because, unlike the human
and CLIP evaluations which assess semantic similarity, the UQI and FID metrics only focus on
image quality. This limitation in evaluating semantic content likely explains the discrepancy in
results between the three automatic evaluation metrics.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Limitations and Risks</title>
      <p>Our work is limited in many aspects. First, we are currently dealing only with English data due
to the limitations of the verbalisation model and the Stable Difusion model we used. Future
work will consider dealing with multilingual datasets as well.</p>
      <p>Additionally, when designing prompts based on Wikidata triples, we had to make decisions such
as extracting triples based on subjects without considering objects. We also treated all triples
equally with no emphasis on properties or types of entities. As shown by the open question in
our human evaluation survey, it is evident that some properties are more important than others
when generating images to portray a specific character. Future work can potentially explore
in more depth which properties lead to better representations of characters. Further, when
encountered with triples that have the same predicate, we selected the one with the longest
object assuming it would contain more information. We are aware that this decision might
have removed important information for characters, and this can be addressed in future work
by concatenating object strings or summarising them automatically.</p>
      <p>Our usage of the Stable Difusion generative model means that our method is inheriting its
biases as well. Although directly leveraging information about each character from its triples is
supposed to limit biases when generating images, this cannot always be controlled (e.g., for
some female entities, the model generated images of male characters). Additionally, using a
predefined set of negative prompts for all characters (which includes terms such as mutilated and
disfigured ) is a considerable limitation of the model to correctly portray characters. A possible
solution for this could be to design specific negative prompts for each individual character in a
semi-automatic manner or to use another type of text-to-image model that does not require
negative prompting.</p>
      <p>Our work is also limited in terms of the ground-truth dataset constructed based on Wikidata
entities that already have images. This is because, for some of these entities, the images are not
reliably portraying the character, but the actor depicting the character.</p>
      <p>Finally, in order to mitigate any copyright and/or privacy risks, we stress that our method is not
suggested to be directly deployed into Wikidata, as we think that using AI-generated images can
potentially be very harmful. Should this method be used for image completion, we encourage
clearly watermarking images as AI-generated.</p>
    </sec>
    <sec id="sec-9">
      <title>8. Conclusion</title>
      <p>
        In this paper, we investigate four diferent methods for generating prompts based on extracting
knowledge in the form of triples. We then generate images based on each prompt using
Stable Difusion, a generative text-to-image model. We evaluate the diferent prompt types by
automatic as well as human evaluation approaches and conclude that the best-generated images
are based on natural language text that includes the context and background of the character.
When possible, this text can be extracted from a human-edited source such as DBpedia abstracts,
however, most characters do not have a DBpedia entity. This brings to light the need to verbalise
triples (i.e. transform them into natural text based on large language models) and use them
as prompts in order to receive the best visual representation of their corresponding fictional
characters. To the best of our knowledge, our work is novel in terms of utilising triples for
prompt engineering in order to complete missing information on Wikidata. Possible future
work includes finetuning the last Stable Difusion model via a Lora adaptation [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ], trying other
text-to-image models that rely on diferent architectures, and modifying prompts to include the
most significant triples by investigating which properties afect image quality the most. Our
approach is not intended to directly complete entities on Wikidata with AI-generated images,
rather it can be used by editors to further enrich entities such as fictional characters, fictional
places, or landscapes. Alternatively, instead of directly using the output of generative models,
they could be given to artists who can use them as inspiration to create depictions of entities.
      </p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>This project is the result of a research task force team at the International Semantic Web Summer School (ISWS) 2023. We would
like to thank the organisers and tutors. We especially thank our mentor Albert Meroño-Peñuela for his valuable advice and input
throughout this work.</p>
      <p>We made use of the central High Performance Computer system at Freie Universität Berlin to conduct the data collection and
image generation parts, and we would like to express our gratitude for the resources provided.</p>
      <p>The work of the author, Sefika Efeoglu, is funded by the German Federal Ministry of Education and Research (BMBF) and the state
of Berlin under the Excellence Strategy of the Federal Government and the Länder over the project.</p>
      <p>This paper has been developed within the HE project MuseIT, which has been cofounded by the European Union under the Grant
Agreement No 101061441. Views and opinions expressed are, however, those of the authors and do not necessarily reflect those of
the European Union or European Research Executive Agency.</p>
      <p>This work has been supported by the French government, through the 3IA Côte d’Azur Investments in the Future project managed
by the National Research Agency (ANR) with the reference number ANR-19-P3IA-0002.</p>
    </sec>
    <sec id="sec-11">
      <title>A. Appendix</title>
      <p>dif
lower
upper
q-value
p-value
basic prompt plain prompt 0.05848
basic prompt verbalised prompt 0.06724
basic prompt dbpedia abstract prompt 0.10023
plain prompt verbalised prompt 0.00875
plain prompt dbpedia abstract prompt 0.04174
verbalised prompt dbpedia abstract prompt 0.03298</p>
      <p>dif
lower
upper
q-value
p-value</p>
      <sec id="sec-11-1">
        <title>CLIP score Vs. Number of relations</title>
      </sec>
      <sec id="sec-11-2">
        <title>CLIP score Vs. Number of unique relations</title>
        <p>Figure 5: The 10 random generated images used for the Human Evaluation
Figure 6: Correlation plots of the automatic and the human evaluation</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Vrandečić</surname>
          </string-name>
          ,
          <article-title>Wikidata: A New Platform for Collaborative Data Collection</article-title>
          ,
          <source>in: Proceedings of the 21st international conference on world wide web</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>1063</fpage>
          -
          <lpage>1064</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Abián</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Meroño-Peñuela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Simperl</surname>
          </string-name>
          ,
          <article-title>An Analysis of Content Gaps Versus User Needs in the Wikidata Knowledge Graph</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2022</year>
          , pp.
          <fpage>354</fpage>
          -
          <lpage>374</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rombach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Blattmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lorenz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Esser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ommer</surname>
          </string-name>
          ,
          <article-title>High-Resolution Image Synthesis With Latent Difusion Models</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>10684</fpage>
          -
          <lpage>10695</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Introducing</surname>
            <given-names>ChatGPT</given-names>
          </string-name>
          ,
          <year>2023</year>
          . URL: https://openai.com/blog/chatgpt.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Krueger</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Learning Transferable Visual Models From Natural Language Supervision</article-title>
          , in: M.
          <string-name>
            <surname>Meila</surname>
          </string-name>
          , T. Zhang (Eds.),
          <source>Proceedings of the 38th International Conference on Machine Learning</source>
          ,
          <string-name>
            <surname>ICML</surname>
          </string-name>
          <year>2021</year>
          ,
          <volume>18</volume>
          -
          <issue>24</issue>
          <year>July 2021</year>
          ,
          <string-name>
            <given-names>Virtual</given-names>
            <surname>Event</surname>
          </string-name>
          , volume
          <volume>139</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>8748</fpage>
          -
          <lpage>8763</lpage>
          . URL: http://proceedings.mlr.press/v139/ radford21a.html.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nichol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Hierarchical Text-Conditional Image Generation with CLIP Latents</article-title>
          ,
          <source>ArXiv abs/2204</source>
          .06125 (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Schölkopf</surname>
          </string-name>
          , Moûsai:
          <article-title>Text-to-Music Generation with Long-Context Latent Difusion</article-title>
          ,
          <source>arXiv preprint arXiv:2301.11757</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C. Zhang,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. S.</given-names>
            <surname>Kweon</surname>
          </string-name>
          ,
          <article-title>Text-to-image Difusion Models in Generative AI: A Survey</article-title>
          ,
          <source>arXiv preprint arXiv:2303.07909</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Mansimov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Parisotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <article-title>Generating Images from Captions with Attention</article-title>
          ,
          <source>arXiv preprint arXiv:1511.02793</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Akata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Logeswaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Schiele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          , Generative Adversarial Text to Image Synthesis, in: International conference on machine learning,
          <source>PMLR</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1060</fpage>
          -
          <lpage>1069</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , T. Xu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. N.</given-names>
            <surname>Metaxas</surname>
          </string-name>
          ,
          <article-title>StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks</article-title>
          ,
          <source>in: Proceedings of the IEEE international conference on computer vision</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>5907</fpage>
          -
          <lpage>5915</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <surname>X. He,</surname>
          </string-name>
          <article-title>AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1316</fpage>
          -
          <lpage>1324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pavlov</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Voss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Sutskever</given-names>
            ,
            <surname>Zero-Shot</surname>
          </string-name>
          Text
          <article-title>-to-Image Generation</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>8821</fpage>
          -
          <lpage>8831</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          , et al.,
          <article-title>Cogview: Mastering text-to-image generation via transformers</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>34</volume>
          (
          <year>2021</year>
          )
          <fpage>19822</fpage>
          -
          <lpage>19835</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>C. Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Duan</surname>
          </string-name>
          , Nüwa:
          <article-title>Visual Synthesis Pre-training for Neural visUal World creAtion</article-title>
          , in: Computer Vision-ECCV
          <year>2022</year>
          : 17th European Conference, Tel Aviv, Israel,
          <source>October 23-27</source>
          ,
          <year>2022</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>XVI</given-names>
          </string-name>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>720</fpage>
          -
          <lpage>736</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Welling</surname>
          </string-name>
          ,
          <string-name>
            <surname>Auto-Encoding Variational</surname>
            <given-names>Bayes</given-names>
          </string-name>
          ,
          <source>CoRR abs/1312</source>
          .6114 (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vahdat</surname>
          </string-name>
          , J. Kautz, NVAE:
          <string-name>
            <given-names>A Deep</given-names>
            <surname>Hierarchical Variational Autoencoder</surname>
          </string-name>
          ,
          <year>2021</year>
          . arXiv:
          <year>2007</year>
          .03898.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <article-title>Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images</article-title>
          , ArXiv abs/
          <year>2011</year>
          .10650 (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nichol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McGrew</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Chen, GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Difusion Models</article-title>
          ,
          <source>arXiv preprint arXiv:2112.10741</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Saharia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Saxena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Whang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Denton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghasemipour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Gontijo</given-names>
            <surname>Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Karagol</given-names>
            <surname>Ayan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salimans</surname>
          </string-name>
          , et al.,
          <article-title>Photorealistic Text-to-Image Difusion Models with Deep Language Understanding</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>35</volume>
          (
          <year>2022</year>
          )
          <fpage>36479</fpage>
          -
          <lpage>36494</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rombach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Blattmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lorenz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Esser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ommer</surname>
          </string-name>
          ,
          <article-title>High-Resolution Image Synthesis with Latent Difusion Models</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>10684</fpage>
          -
          <lpage>10695</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>V.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. B.</given-names>
            <surname>Chilton</surname>
          </string-name>
          ,
          <article-title>Design Guidelines for Prompt Engineering Text-to-Image Generative Models</article-title>
          ,
          <source>in: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Oppenlaender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Linder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Silvennoinen</surname>
          </string-name>
          , Prompting AI Art:
          <article-title>An Investigation into the Creative Skill of Prompt Engineering</article-title>
          , arXiv preprint arXiv:
          <volume>2303</volume>
          .13534 (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>J.</given-names>
            <surname>Oppenlaender</surname>
          </string-name>
          ,
          <article-title>A Taxonomy of Prompt Modifiers for Text-to-Image Generation</article-title>
          ,
          <source>arXiv preprint arXiv:2204.13988</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          , C. Ma,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hu</surname>
          </string-name>
          , et al.,
          <source>Prompt Engineering for Healthcare: Methodologies and Applications</source>
          , arXiv preprint arXiv:
          <volume>2304</volume>
          .14670 (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>G.</given-names>
            <surname>Amaral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Rodrigues</surname>
          </string-name>
          , E. Simperl,
          <article-title>WDV: A Broad Data Verbalisation Dataset Built from Wikidata</article-title>
          , in: U. Sattler,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Keet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Presutti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P. A.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Takeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Monnin</surname>
          </string-name>
          , G. Pirrò, C. d'Amato (Eds.),
          <source>The Semantic Web - ISWC 2022</source>
          , Springer International Publishing, Cham,
          <year>2022</year>
          , pp.
          <fpage>556</fpage>
          -
          <lpage>574</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>L. F. R.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schütze</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Investigating Pretrained Language Models for Graph-to-Text Generation</article-title>
          ,
          <source>in: Proceedings of the 3rd Workshop on Natural Language Processing for Conversational AI</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>211</fpage>
          -
          <lpage>227</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .nlp4convai-
          <fpage>1</fpage>
          .20. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .nlp4convai-
          <fpage>1</fpage>
          .
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Brümmer</surname>
            , Martin and Dojchinovski, Milan and Hellmann, Sebastian, DBpedia
            <given-names>A</given-names>
            bstracts: A
          </string-name>
          <string-name>
            <surname>Large-Scale</surname>
          </string-name>
          , Open, Multilingual NLP Training Corpus,
          <source>in: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Portorož, Slovenia,
          <year>2016</year>
          , pp.
          <fpage>3339</fpage>
          -
          <lpage>3343</lpage>
          . URL: https://aclanthology.org/ L16-1532.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <surname>X-IQE</surname>
          </string-name>
          :
          <article-title>eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models</article-title>
          ,
          <source>arXiv preprint arXiv:2305.10843</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>D.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <article-title>Full-Reference Image Quality Assessment Based on an Optimal Linear Combination of Quality Measures Selected by Simulated Annealing</article-title>
          ,
          <source>Journal of Imaging</source>
          <volume>8</volume>
          (
          <year>2022</year>
          ). URL: https://www.mdpi.com/2313-433X/8/8/224. doi:
          <volume>10</volume>
          .3390/ jimaging8080224.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hessel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Holtzman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Forbes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Le</given-names>
            <surname>Bras</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Choi,</surname>
          </string-name>
          <article-title>CLIPScore: A Reference-free Evaluation Metric for Image Captioning</article-title>
          ,
          <source>in: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Online and
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>7514</fpage>
          -
          <lpage>7528</lpage>
          . URL: https: //aclanthology.org/
          <year>2021</year>
          .emnlp-main.
          <volume>595</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .emnlp- main.595.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weissenborn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Minderer</surname>
          </string-name>
          , G. Heigold,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Houlsby</surname>
          </string-name>
          ,
          <article-title>An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale</article-title>
          ,
          <source>in: 9th International Conference on Learning Representations, ICLR</source>
          <year>2021</year>
          ,
          <string-name>
            <given-names>Virtual</given-names>
            <surname>Event</surname>
          </string-name>
          , Austria, May 3-
          <issue>7</issue>
          ,
          <year>2021</year>
          , OpenReview.net,
          <year>2021</year>
          . URL: https://openreview.net/forum?id=YicbFdNTTy.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>M.</given-names>
            <surname>Heusel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ramsauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nessler</surname>
          </string-name>
          , S. Hochreiter,
          <article-title>GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <volume>1706</volume>
          .
          <fpage>08500</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Hayes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Krippendorf</surname>
          </string-name>
          ,
          <article-title>Answering the Call for a Standard Reliability Measure for Coding Data</article-title>
          ,
          <source>Communication Methods and Measures</source>
          <volume>1</volume>
          (
          <year>2007</year>
          )
          <fpage>77</fpage>
          -
          <lpage>89</lpage>
          . URL: https://doi.org/10.1080/19312450709336664. doi:
          <volume>10</volume>
          .1080/19312450709336664. arXiv:https://doi.org/10.1080/19312450709336664.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wallis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Allen-Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          , W. Chen, LoRA: Low-Rank
          <source>Adaptation of Large Language Models</source>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2106</volume>
          .
          <fpage>09685</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Lancelot (Q215681)</surname>
          </string-name>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>