<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>November</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Vividness Bias in Language Models Through Art Interpretations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Laura Samela</string-name>
          <email>ls.la.samela@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrico Daga</string-name>
          <email>enrico.daga@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul Mulholland</string-name>
          <email>paul.mulholland@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>AI and Art, Large Language Models, Social bias</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The Open University</institution>
          ,
          <addr-line>Walton Hall, Milton Keynes</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>2</volume>
      <fpage>6</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>Large language models (LLMs) play a crucial role in applications that require to tailor the content to user backgrounds and perspectives. In the context of cultural engagement, these models hold the promise of tailoring art interpretations to diverse audiences. However, LLMs are known to generate biased content, therefore, perpetrating stereotypes and inequality. Knowledge engineering methodologies can support the systematic observation of generative AI outputs. In this paper, we propose a method to identify these biases through personabased prompting. Crucially, we find evidence of vividness bias, a known phenomenon in social psychology where our decisions are driven by specific aspects in a given situation. Therefore, we pose the question of investigating such bias systematically, and propose a method based on in-context learning with pairwise association of persona features. Next, we represent LLM behaviour as a decision tree, to capture detailed evidence of bias. We investigate this phenomenon with artworks from the Irish Museum of Modern Art (IMMA) and Google Bard, focusing on features such as gender, race, age, profession, and sexual orientation. We discuss our findings and identify opportunities and challenges when dealing with vividness bias in persona-based, generated art interpretations.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In the dynamic intersection of technology and cultural heritage, museums strive to evolve, seeking
innovative ways to engage diverse audiences. Within the realm of museum culture, both staf and
visitors tend to be less diverse than the broader population [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Therefore, museums face the challenge
of making their collections and exhibitions relevant to a broader range of diverse communities often
underrepresented in society or among museum visitors. The advent of Large Language Models (LLMs)
ofers a promising avenue for enhancing museum interactions by tailoring content to the unique
preferences and backgrounds of individual visitors [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In theory, LLMs can dynamically process and
repurpose information tailoring it to specific audiences through the adoption of personas, therefore,
potentially transforming museums into interactive, personalised hubs of learning and engagement [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
LLMs hold the potential to fulfil this aspiration, but it comes with inherent risks related to biases and
stereotypes [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. In this research, we delve into the assessment of these risks and how generative
AI can help museums achieve their goal of providing customised, inclusive experiences. This raises
critical questions: 1) In what way bias of LLMs impact persona-based art interpretations? 2) How can
we identify systematically bias to inform mitigation strategies?
      </p>
      <p>In this paper we investigate bias in LLM through persona-based art interpretations. Crucially, we
empirically observe bias in art interpretations as vividness bias, a concept in social psychology where
the most striking information has a strong impact on our thinking and decision-making, making us
ignore other aspects that are equally or more important. In our scenario, certain persona features are
deemed more relevant than others by the language model and, more surprisingly, certain features’
categories seems to be more vivid than others. We pose the question on how to characterise vividness
bias in generated, persona-based art interpretations. In this work, we propose a method to identify</p>
      <p>CEUR
Workshop
Proceedings</p>
      <p>ceur-ws.org
ISSN1613-0073
and characterise such bias, experimenting with categories of features such as gender, race, ethnicity,
age, profession, and sexual orientation. Our investigation is performed with artworks from the Irish
Museum of Modern Art (IMMA) and Google Bard. With an approach based on in-context learning with
pairwise association of persona features, we acquire tailored art interpretations. We represent LLM
behaviour as a decision tree to generate fine-grained evidence. The derived rule sets expresses formally
the vividity bias of the large language model.</p>
      <p>The rest of the paper is structured as follows. We discuss background motivation in Section 2, where
we report on exploratory experiments in generating persona based art interpretations, which trigger
our research questions. Next, we introduce our methodology in Section 3 and apply it in experiments
in Section 4. Next, we discuss our findings in Section 5. We report on related work in Section 6 before
closing the paper with conclusions, limitations and future work (Section 7).</p>
    </sec>
    <sec id="sec-2">
      <title>2. Persona-based art interpretations</title>
      <p>
        In 2015, The Warwick Commission reported that UK residents from higher socioeconomic groups
accounted for 87% of museum visitors. One crucial challenge was “a mismatch between the public’s
taste and the publicly funded cultural ofer” [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Following this trend, museums are increasingly looking
into novel ways to engage visitors in finding closer, personal connections with artworks [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Artificial intelligence can have a huge impact in how art is produced, studied, and perceived [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Here, we explore persona-based art interpretations via a Large Language Model (LLM), specifically,
Google Bard, with the objective of gather evidence of potential bias, as it happens in many other
domains. However, here we ask ourselves what is the possible impact of such biases on museum visitors
interacting with an AI that generates tailored interpretations.
      </p>
      <p>
        We conducted initial experiments using persona-driven prompts to generate art interpretations1. To
this aim, we designed five personas combining diferent features. These are summarised in Table 2. In
our work, we consider the collection of the Irish Museum of Modern Art (IMMA), which is particularly
active in citizen engagement initiatives tailored to diverse communities [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The general scenario is akin
to citizen engagement systems such as deep viewpoints [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], where citizens interact with the collection by
exchanging personal interpretations. We select four artworks, diferent in technique, style, and content,
that have been previously used by museum curators in workshops with visitors. The artworks are:
• ”Little Miss Mufet” by Paula Rego 2: The image, a black and white pastel on paper, shows an
allegorical picture of a giant spider with a human face approaching a young girl from behind.
      </p>
      <p>The work touches psychoanalytic references with dark undertones.</p>
      <sec id="sec-2-1">
        <title>1The data was collected with Google Bard version 3 during Summer 2023. 2https://imma.ie/collection/little-miss-muffet/</title>
        <p>• ”Sthombe La Reunion” by Zanele Muholi3: The artist captures the essence of identity, gender, and
sexuality within South African culture. This is a photographic portrait of a man of African origin
on a plain black-point background.
• ”Berry Dress” by Alice Maher4: Created using mixed media, the work displays a dress-like shape
with abstract patterns, challenging conventional notions of reality and the human form.
• ”Margaret Thatcher’s Lunch” by Derek Jarman5: This piece is an experimental film still, focusing
on a political figure and a distinct context. Its inclusion allowed us to further explore the impact
3Featured in the exhibition The Narrow Gate of the Here-and Now: Queer Embodiment
https://imma.ie/whats-on/the-narrowgate-of-the-here-and-now-chapter-one-queer-embodiment/
4https://imma.ie/collection/berry-dress/
5https://imma.ie/collection/margaret-thatchers-lunch/</p>
        <p>Persona Background
1. Javier the Soft- Grew up in a middle-class family, pursued a degree in computer
ware Engineer science, is passionate about technology, successful software
engi</p>
        <p>neer, and enjoys playing soccer and volunteering.
2. Aisha the Mar- Comes from a multicultural background, achieved great success
keting Executive in the marketing industry, is known for creativity and innovative</p>
        <p>strategies, and advocates for LGBTQ+ rights.
3. Amir the High High school student with a passion for music and writing, comes
School Student from a working-class family, faces cultural challenges, actively
participates in LGBTQ+ support groups at school, and dreams of
pursuing a career in journalism.
4. Mei Ling the Environmental activist fighting for sustainability and climate
Environmental justice, has faced discrimination, remains determined to make a
Activist positive impact, enjoys hiking, photography, and advocating for</p>
        <p>transgender rights
5. Lisa the Retail Works as a cashier in a retail store, comes from a modest
backStore Cashier ground, did not pursue a university degree, enjoys interacting</p>
        <p>with customers, and has a passion for gardening.</p>
        <p>of personas on LLM-generated insights, especially in the absence of specific visual cues.</p>
        <p>The selected pieces vary in medium, theme, and style. Crucially, each artwork has distinct
features, stimulating a unique narrative and diferent visual complexity to better challenge the LLM’s
interpretative capacities.</p>
        <p>We designed two experiments combining each artwork with the two types of prompts: (a)
impersonating persona prompts, where the user adopted the persona’s perspective, providing a personal
introduction to the model, and (b) targeting persona prompts, where explicit instructions guided the LLM
to generate interpretations tailored to a specific demographics (e.g. a White - Transgender). Examples
are included in Figures 1 and 2.</p>
        <p>Impersonating persona prompt This approach involved crafting prompts wherein the experimenter,
adopted the role of the persona. For example, when exploring Paula Rego’s ”Little Miss Mufet,” the
prompt might read: ”Bard, I am an LGBTQ+ individual; please interpret ’Little Miss Mufet’ considering
my perspective and experiences”. Here, the LLM was encouraged to immerse itself in the persona’s
unique characteristics while generating interpretations.</p>
        <p>Targeting persona prompt An alternate approach revolved around providing explicit instructions to
the language model, directing it to generate an interpretation tailored to a specific persona. For instance,
when exploring Zanele Muholi’s ”Sthombe La Reunion,” the prompt is: ”Bard, interpret ’Sthombe La
Reunion’ for an African educator who identifies as non-binary” 6.</p>
        <p>
          We then performed a thematic analysis [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] of the generated interpretations. Interpretations were
organised in tables in accordance with the specific artwork, the persona and each visual element. More
specifically, themes covered visual elements of the artwork, emotions and meaning captured by LLM in
relation to each persona. These tables served as a repository for interpretations, allowing us to track
changes and variations in Bard’s responses across diferent conditions. This thematic analysis led to
preliminary insights into how personas influenced LLM interpretation by extrapolating general themes
of LLM related to personas’ diversity features.
        </p>
        <p>From this preliminary investigation, we can make several observations. The LLM demonstrated
varying levels of sensitivity to diferent persona attributes, with gender and sexual orientation playing</p>
      </sec>
      <sec id="sec-2-2">
        <title>6All personas interpretation can be reviewed in the supplementary material [9].</title>
        <p>significant roles in some cases while being less relevant in others. This variability points to inherent
biases in the interpretations, where certain attributes are highlighted over others based on the persona.</p>
        <p>Paula Rego’s ”Little Miss Mufet” When the persona included sexual diversity features, the
interpretations highlighted themes of identity and inclusivity, which align closely with the feminist undertones
of Rego’s work. When the persona was a man with a professional focus, the AI shifted the narrative
towards career aspects, downplaying the fear depicted in the artwork. This suggests that the professional
focus of the persona can skew the interpretation towards themes of achievement and career success,
overshadowing the original emotional content of the piece. In other interpretations, the content of the
artwork, which includes elements of fear and childhood, significantly influenced the interpretations.
The rich and layered nature of Paula Rego’s work likely contributed to the LLM’s tendency to focus on
specific features in the gender category.</p>
        <p>Zanele Muholi’s ”Sthombe La Reunion” is a poignant photographic portrait that explores identity,
gender, and sexuality within South African culture. The work is rich in socio-cultural symbolism,
influenced the AI’s interpretations by foregrounding themes of resilience and empowerment in black
and LGBTQ+ communities. The interpretation for a persona that included LGBTQ+ features emphasized
challenges related to race and gender identity. The AI tended to produce similar interpretations for
diferent personas, particularly for prompt two, indicating a bias towards certain socio-cultural themes
inherent in the artwork. In late trials, after an update, Bard refused to provide an interpretation when
uploading the image, answering that it could not process people’s photos yet (for us, anymore!).</p>
        <p>Alice Maher’s ”Berry Dress”, having a surreal and abstract nature, allowed for varied interpretations
based on diferent personas. For instance, interpretations for personas with a professional background
focused on themes of beauty and representation in media, while those with LGBTQ+ features highlighted
themes of diversity and resilience. The LLM demonstrated a significant sensitivity to the diversity
attributes of the personas, producing interpretations that reflected the intersection of gender, race,
and sexual orientation. This indicates that the abstract content of Maher’s artwork provided a flexible
canvas for the LLM, resulting in diverse interpretations that were better aligned with the features of
the persona.</p>
        <p>Derek Jarman’s ”Margaret Thatcher’s Lunch” is a politically charged artwork that provided a distinct
context for analysing LLM interpretations. The contemporary political figure and the socio-political
context of the piece influenced the LLM to generate interpretations that were highly critical and focused
on themes of power, privilege, and oppression. For example, when the persona included features
related to LGBTQ+ identity, the interpretations emphasized political activism and the social impact of
Thatcher’s policies.</p>
        <p>
          Summary of preliminary findings Generally, our findings confirms current research on the role of
persona in enriching the quality of generated text [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Except for the case of La Reunion, interpretations
made relevant associations between the artwork and persona characteristics. However, we observe how
interpretations coming from impersonating persona prompts are generally richer and more engaging.
Targeting persona prompt often lead to more generalised interpretations that may not deeply connect
with personal themes, resulting in outputs that, while broadly relevant, lack the insight and depth found
in interpretations coming from impersonating persona prompts. By providing the LLM with detailed,
tailored prompts, the LLM can generate interpretations that align with the unique perspectives and
experiences of each persona. Impersonating persona prompts generally produced more consistent and
targeted interpretations. Except for the case of La Reunion, we provided the image of the artwork. It
is not our objective to assess to what extent the information accountable to the artwork comes from
the processing of the image or from background knowledge captured from web resources (e.g. the
IMMA website). However, within both types of interpretations, features are used in an inconsistent,
unpredictable way. To a certain extent, this may be caused by the relation of specific features with
the artwork content. However, in many other cases we could not determine the reason of certain
features being neglected in favour of others. Generally, we make the observation that the LLM exhibits
a vividness bias, in the sense that the model prioritises certain persona features over others when
generating art interpretations. Therefore, we pose the question: How can we systematically capture
and analyse vividness bias in the interpretative behaviour of a large language model, when generating
persona-based art interpretations?
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>Using personas in LLM prompts will produce tailored interpretations of artworks. Impersonating the
persona seemed to provide better results. However, we also observed how persona features are not
treated equally by the language models. In order to systematically study this phenomenon, we designed
a methodology that combines a given set of characteristics and generates personas by combining them
pairwise. In this way, we can check for each combination, which one is the prominent feature considered
by the generated interpretation. The method is structured as follows:
Features design The first step is to identify a comprehensive set of categories (e.g. Profession, Gender,
etc...) and associated diversity features (Nurse, Clerical, Transgender, etc...).</p>
      <p>Artwork Selection In this phase, our objective is twofold. On the one hand, we need for the artwork
to be a stable variable in our experiments. Therefore, we select a single artwork from the museum
collection. On the other hand, to minimise the impact of the artwork content on the interpretation, we
select one found in the preliminary study to be most open to alternative interpretations for diferent
personas.</p>
      <p>Persona generation All the features are combined in pairs to generate personas to be used in the
prompts. In this way, we aim to observe how changes in persona features afect the LLM’s interpretation.
Prompts generation In this phase, we generate impersonating persona prompts that incorporate each
one pair of features. These prompts should guide the LLM to consider the designated attributes during
interpretation.</p>
      <p>Data Collection We administer the prompts to the LLM, pairing each prompt with the selected artwork.
Next, collect the LLM’s interpretations.</p>
      <p>Data Annotation In this phase, we categorise the interpretations based on the diversity features
emphasised in the responses. We organise data into a table where each row represents an interpretation,
and each column corresponds to a diversity feature. Next, we mark the primary and secondary features
of each interpretation based on their prominence in the LLM’s output.</p>
      <p>Vividness Analysis In this phase, we conduct two types of analysis. First, we rank the diversity
categories, from most to least influential based on their frequency of prominence annotations. Next, we
organise the data into a decision tree to visualise and further analyse the hierarchy of influences among
the categories and related diversity features, in detail. This approach allows to create a deterministic,
data-driven representation of vividness selectivity bias of the LLM on persona features of art
interpretations. The Decision Tree method is therefore being used analytically to provide a description of the
responses made by the LLM rather than a model used to predict future responses. Next, we analyse
the hierarchy to derive insights about the behaviour of LLM interpretations. We read the decision tree
from the bottom up of each end node, discussing the impact of vividness bias on the diferent persona
characteristics.</p>
      <p>The methodology was designed to evaluate the influence hierarchy of various persona characteristics on
LLM-generated art interpretations systematically. In the next section, we apply this method to identify
the precedence of diferent features in the generation of interpretations by the LLM.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>In this section, we apply our methodology to capture how the combination of diverse personas within
an art interpretation prompt allows to identify and characterise vividness selection bias of generative
AI. We selected Google Bard for its ability to support an image as part of the prompt context7.</p>
      <p>Features design Findings from our exploratory analysis underscore the need for a broader range of
persona features to capture a more diverse spectrum of interpretations. For example, we added the
category Age and increased the values of others. Table 1 shows the set of features used, extending the
ones originally used in the design of the persona for the exploratory analysis (Section 2).</p>
      <p>Artwork selection We select one of the artworks used by the previous experiments (Section 2) as the
base for investigating the influence of diversity features pair. The artwork used is ”Berry Dress” by
Alice Maher. This conclusion was drawn from analysis results 8 which had shown very variegated
results with both impersonating and targeting persona prompts. In contrast, for example, to La Reunion,
yielded identical results for persona-specific prompts, lacking suficient diversity, and to Little miss
Mufet, who showed an excessive impact of artwork content. This consistent context of reference
provided by the Berry dress helps to minimise the noise introduced by comparing interpretations from
diferent artworks, thereby ensuring a more controlled and reliable experimental setting.</p>
      <p>Persona generation In this phase, we combined all identified features pairwise to generate unique
combinations. Personas include, for example, a Male Transgender, a 18-30Age Black person, an Asian
Nurse, and so forth. A total of 205 unique binary combinations were tested, each serving as a lens
through which the LLM would interpret the artwork.</p>
      <p>Prompt generation Next, we generate impersonating prompts in the following form:
I am a ( f e a t u r e 1 ) − ( f e a t u r e 2 ) , i n t e r p r e t t h e A l i c e Maher ’ s B e r r y</p>
      <p>D r e s s f o r me .</p>
      <p>In addition, we provided the image as part of the context. The prompts have always been run once.</p>
      <p>Data collection We administered a series of mixed questions to LLM, featuring prompts that included
the designated pair of diversity features, for example: ”Bard, I am a Male - African, interpret the Alice
Maher’s Berry Dress for me.”. This is repeated across all possible binary combinations of the features to
generate a range of interpretations.</p>
      <p>
        Data annotation The interpretations generated by the LLM were systematically documented and
categorised based on the diversity features emphasised in the responses. Firstly, the initial dataset
consists of the collected interpretations, each associated with specific diversity features including gender,
race/ethnicity, age, profession, and sexual orientation. Primary and secondary features were marked
based on their prominence in the LLM’s interpretation. For example, for the interpretation tailored on a
7This was true at the time of our data collection during Summer 2023. More details can be found in the supplementary
material [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
8Extensive results can be reviewed in the supplementary material [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
male transgender nurse, the result was clearly designed around the profession feature (see Fig. 3). We
collect the data into a table, where features are marked as primary or secondary factors based on their
occurrence in the AI’s interpretation.
      </p>
      <p>Vividness analysis In this phase, we analyse the data to capture information about the vividness of
persona attributes in the interpretations, by ranking the categories considered primary features and by
organising the data in a decision tree.</p>
      <p>Ranking. Based on the frequency data, we rank the categories from most to least influential. The
ranking ofers a quantitative assessment of the impact of various diversity features on AI-driven art
interpretations: 1. Profession (81 mentions) 2. Gender (39 mentions) 3. Sexual Orientation (35 mentions)
4. Age (25 mentions) 5. Race (25 mentions).</p>
      <p>Decision tree We analyse the data using a decision tree classifier and generate a hierarchy of
priorities between features. To calculate the decision tree each persona feature was represented as a
binary variable, e.g. gender-male, gender-female, etc. The One-Hot Encoding technique in scikit-learn
Python was then used to generate the decision tree, illustrated in Figure 4.</p>
      <p>This decision tree allows us to represent how the AI systematically prioritises personal attributes
when interpreting artworks, ultimately leading to a tailored outcome based on the combination of
features it has determined to be relevant and features it has decided are not. The AI’s interpretations
change depending on the presence or absence, as well as the specific features, of the categories such
as profession, gender, race, age, and sexual orientation. The decision tree describes the vividness of
features as a step-by-step choice, systematically including or excluding persona features to guide the
AI’s interpretation of artworks. Although such representation does not reflect the way generative
language models operate, it provides a useful description of their behaviour. In what follows, we
describe the results via a decision-making metaphor, for illustrative purposes, as-if the decision tree
would be the process that leads to the result:
• Undefined Profession: At first, the model checks whether Profession is defined for the persona.</p>
      <p>This is the largest exclusion criterion, accounting for 112 out of the 205 combinations. If the
profession is defined, it moves to assess sexual orientation. If profession is undefined it moves
onto assess gender.
• Undefined Orientation: If sexual orientation and profession are both defined, the model proceeds
to consider whether the persona has a heterosexual orientation. If sexual orientation is undefined
the model proceeds to consider whether the persona is a transgender male.
• Heterosexual orientation: If the persona does not have a heterosexual orientation, then sexual
orientation is the primary feature used in the art interpretation. If the person does have a have a
heterosexual orientation then profession is used in the art interpretation.
• Transgender male: From this point on, the model goes on to consider whether the person is
transgender male or transgender female. If either is the case, then gender is the primary feature
used in art interpretation.
• Age and Profession: If the persona if not transgender then either age or profession is selected as
the primary feature used in art interpretation. The selection of age is associated with personas
having a profession of retired, clerical or student.
• Gender and race: On the right hand side of the decision tree, if race and gender are both defined
in the persona then race tends to be selected if the persona is not white. Gender tends to be
selected if the persona is transgender or female.
• Age and sexual orientation: On the right hand side of the decision tree, similarly, age tends to be
associated with younger personas (18-30) and orientation with personas who are not heterosexual.</p>
      <p>Analysis of the structure of the tree complements the ranking of the features presented earlier,
shedding light on how the features were used to make a selection. Essentially, the decision making
process tends to exclude demographic features that could be considered dominant or default for example
male gender, heterosexual sexuality or white race. Analysis of the model suggests how the vividness of a
characteristic, for example a gender, sexuality or race, reflects the non-dominance of that characteristic
as represented in the LLM. For example, the art interpretation would not reference gender for a male
persona, sexuality for a heterosexual persona or race for a white persona. This method explains in
detail the vividness selectivity bias of the language model, indicating that certain personal attributes
are deemed more attractive for shaping the generated interpretation of art.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>To further understand the implications of our findings, it is crucial to delve into the biases revealed by
the AI’s interpretations, especially regarding how diferent aspects are emphasised or overlooked.</p>
      <p>
        Vividness bias and Race There is a notable pattern where the interpretation of AI art often references
race when it involves non-White individuals (e.g., African, Asian). When diversity features includes
White, the interpretation seems less inclined to focus on racial background and instead gives attention
to other attributes such as gender or profession. This can be seen as an example of what Dabiri [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
describe as the ”invisibility” of white people, who are presented as just people and assumed to be the
default norm. Additionally, while the term ”caucasian” may be repeated in the output when included
in the prompt, this repetition does not imply that race was actively considered as a dimension in the
interpretation of the artwork. The mention of racial categories does not necessarily indicate their
influence on the interpretive process, aligning with the findings related to vividness bias, particularly
when diversity features such as ”White” are present. Moreover, such patterns indicate a problematic
inclination to view race as a defining characteristic of non-White individuals, diminishing other aspects
such as age or gender.
      </p>
      <p>Vividness bias and Gender With Non-Heterosexual features, gender appears to be a more common
point of focus in the AI interpretations. The inclination to foreground gender in the case of White
individuals may potentially reinforce the notion of Whiteness as a ’default’ or ’norm’ in the context of
museum engagement, which goes against the objective of inclusion and diversity that cultural heritage
institution strive to achieve.</p>
      <p>Vividness bias and Profession The profession is a highly predominant feature across almost all
intersections, indicating a bias towards perceiving an individual’s occupation as a equally defining characteristic
in AI art interpretations. In reality, some people consider their profession more or less defining in term
of identity. This might lead to a narrow interpretation of individuals, overshadowing other equally
significant aspects of their identity (see example in Fig. 3).</p>
      <p>Vividness bias and Age Age is a significant feature in interpretations, but it seems less emphasised
when paired with race/ethnicity for non-White individuals, suggesting that the AI may be defaulting to
race as the primary diversity feature over age in such cases. The lesser emphasis on age in interpretations
involving non-White individuals points to a potential bias where race/ethnicity is considered the more
dominant diversity feature. This can result in an incomplete representation of an individual’s identity,
especially in a multicultural ageing society, while it is relevant for a straight persona. For instance, when
prompted by a 60+-year-old heterosexual, the resulting interpretation centered around transformation,
endurance, and the beauty of aging, while neglecting to address the sexual orientation of the individual.</p>
      <p>Vividness bias and Sexual Orientation Sexual Orientation tends to be a focus more so when the subjects
are paired with gender but is less emphasised when combined with race/ethnicity or profession.
Heterosexual orientation may be perceived as the norm or default setting within the AI interpretations, which
aligns with a common social bias. Furthermore, Non-heterosexual orientations such as homosexuality
and bisexuality seem to be more in the foreground, except when in association with Profession. When
sexual orientation intersects with other diversity features like race or gender, there appears to be a
nuanced shift in focus. For example, if AI more frequently mentions sexual orientation when the subject
is non-heterosexual, this may reflect societal biases where non-heterosexual orientations are often
unduly spotlighted or ”othered”. Such a tendency to consider non-heterosexual orientations more vivid
can be problematic if it unnecessarily overemphasizes one aspect of a person’s identity, which can
contribute to stereotyping or reductionist views.</p>
      <p>Vividness and Intersectional Bias Art interpretations give an exciting opportunity to reflect a challenge
in adequately capturing the complexity of intersectionality. For example, when features are combined,
some aspects of identity seem to overshadow others, which may oversimplify or misrepresent the
multidimensional nature of identity. When diverse features are combined, some aspects, like race or
profession, often overshadow others, leading to oversimplified or inaccurate representations.</p>
      <p>
        In summary, these biases in AI interpretations underscore the need for more nuanced, context-aware
algorithms that can better represent the multifaceted nature of human identity. AI development needs
to incorporate a deeper understanding of social and cultural contexts to avoid perpetuating stereotypes
and to provide more balanced and comprehensive interpretations. The experiment results, associated
software and data can be found at [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Related work</title>
      <p>
        In recent decades, extensive digitization eforts have opened new opportunities for AI research in
art [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Existing approaches analyse collections to derive similarities, including classifications of
authorship, time period, and genre [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], or for tasks such as object recognition [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], sometimes
according to established taxonomies such as Iconclass [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Ontology engineering methods allow to
break down the art interpretation process in detail, to support advanced large-scale analysis of
museum collections [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Ontology engineering methods are applied to art interpretations to support
interactive narratives [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and formalise iconography [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Generative AI can potentially automate
activities with any artwork in the digital collection and reach out to groups that are less able to access
museums and engage with them in a personalised experience. Specifying persona allows adjusting the
communication style of the generated text [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and can support cases such as personalised dialogue
systems [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and explainable recommendations [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. These family of techniques can be easily applied
to the case of personalised artwork interpretations. However, generative AI systems may be afected
by information bias, in the way they promote or demote relevant information [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Applying large
language models encompasses considering the biases originating from the distribution of training data
and hidden in the learnt model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Specifically, identifying social biases in language models is an
important and dificult problem [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Stereotypical bias can be measured by Context Association Test
(CAT), where a fill-in-the-blank question is paired with a number of choices, including stereotypical
ones [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Bias assessment can be performed by developing large benchmarks encompassing target
groups and associated stereotypes [
        <xref ref-type="bibr" rid="ref23 ref24">23, 24</xref>
        ]. Tackling bias is one objectives of semantic and knowledge
engineering [
        <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
        ]. Using personas has been shown to reduce biased responses compared to not using
any personas [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. However, it has also been found that the choice of persona can impact the extent
of negative outcomes in the responses generated, so it’s important to assess personas systematically
before deploying them [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. Several metrics have been produced by the literature to quantify bias [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
Our objective is not to measure bias but to assess its indirect efect on art interpretations.
      </p>
      <p>
        In our work, we do not analyse what type of societal or cultural bias is expressed by the AI systems,
nor what is the source of such behaviour. Instead, we develop a methodology to collect AI answers
systematically and cluster responses to support their explanation. Crucially, we make the hypothesis
that information bias in LLM can be analysed and described as vividness bias. Per Nisbett and Ross’
widely accepted definition [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], vivid information is characterised as “likely to attract and hold our
attention and to excite the imagination to the extent that it is emotionally interesting, concrete and
imageryprovoking, and proximate in a sensory, temporal, or spatial way”. Our results leads us into hypothesising
a vividity bias in the way art interpretations are produced. To the best of our knowledge, this is the
ifrst work focusing on analysing bias of large language models when developing persona-based art
interpretations and the first one related to vividness bias in such models.
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions, limitations, and future work</title>
      <p>Large language models provide an exciting opportunity to open up museum collection to a diverse
range of personal backgrounds and experiences. However, implicit bias may produce unwanted results,
stigmatising individuals towards specific personal characteristics. In our empirical analysis, we identified
this problem as a way of expressing vividness selective bias, a known phenomenon in social psychology
where our decisions are irrationally driven by the most emotionally engaging aspects of a situation. Our
experiments revealed a pronounced variability in AI-generated art interpretations, heavily influenced
by the specific characteristics of the personas employed. Key findings included Bard’s tendency to
prioritise professional background, sexual orientation, age, gender, and race/ethnicity, in that order,
which significantly afected its analysis of artworks.</p>
      <p>
        In this paper, for the first time, we use art interpretations as a lens to uncover implicit bias of a large
language model, through a knowledge engineering approach. Limitations of the present work include
the scale of the analysis, which included a limited set of features. Future work includes analysing a
broader range of characteristics. Furthermore, we performed experiments with a single large language
model, Google Bard, and a specific version, which is currently being superseeded by Google Gemini.
This problem is a current issue with recent research on large language models [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. Finally, although our
methodology allows to generate a set of decision rules that characterise vividness bias systematically, the
interpretation and analysis of such rules require qualitative method with significant human intervention.
Future work includes strategies for scaling up the methodology and combine human judgement with
semi-automatic methods, as well as compare vividness bias across multiple large language models.
      </p>
      <p>This research contributes to the growing field of AI in the arts, it highlights the potential of AI to
democratize art appreciation through personalized interpretations, but underscores the need for ethical
considerations in AI development to avoid biases. We plan to verify our method and findings and
aligning it to current research in social psychology, especially on the relationship between personal
characteristics, social expectations, and personal identity.</p>
      <p>Furthermore, there is scope for developing sophisticated prompt engineering techniques as mitigation
strategies towards developing tailored, accurate and, ethically responsible art interpretations. In the
future, we plan to expand and detail the feature categories, separating race and ethnicity as two distinct
dimensions to enhance the precision and depth of analysis. Additionally, the prompts in this study have
always been run once; future work should include multiple runs of the LLM to improve robustness
and consistency in outputs. Collaborations between technologists, artists, ethicists, and curators will
be crucial in advancing the field and ensuring that AI serves as a tool for enhancing, rather than
diminishing, the richness and diversity of human cultural expression.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work was supported by the EPSRC Training Grant DTP 2022–2024 of The Open University and
by the EU’s Horizon Europe research and innovation programme under the Polifonia project (Grant
Agreement No. 101004746).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Neelands</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Belfiore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Firth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Perrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Brock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Holdaway</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Woddis</surname>
          </string-name>
          , Enriching Britain: culture, creativity and growth, University of Warwick,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Hettmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wölfel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Butz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Torner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Finken</surname>
          </string-name>
          ,
          <article-title>Engaging museum visitors with ai-generated narration and gameplay</article-title>
          , in: International Conference on ArtsIT,
          <source>Interactivity and Game Creation</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Constantinides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Constantinides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Koukopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fidas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Belk</surname>
          </string-name>
          , Culturai:
          <article-title>Exploring mixed reality art exhibitions with large language models for personalized immersive experiences</article-title>
          ,
          <source>in: Adjunct Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>102</fpage>
          -
          <lpage>105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-P.</given-names>
            <surname>Morency</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <article-title>Towards understanding and mitigating social biases in language models</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>6565</fpage>
          -
          <lpage>6576</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , et al.,
          <article-title>A survey on evaluation of large language models</article-title>
          ,
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Mulholland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stoneman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          , E. Daga,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maguire</surname>
          </string-name>
          ,
          <article-title>Deep viewpoints: Scripted support for the citizen curation of museum artworks</article-title>
          ,
          <source>in: Proceedings of the 34th ACM Conference on Hypertext and Social Media</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hutson</surname>
          </string-name>
          ,
          <article-title>Integrating art and ai: Evaluating the educational impact of ai tools in digital art history learning, in: Forum for Art Studies</article-title>
          , volume
          <volume>1</volume>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Stoneman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          , E. Daga,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maguire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mulholland</surname>
          </string-name>
          ,
          <article-title>Uncomfortable revelations: can citizen curation widen access to museums?</article-title>
          ,
          <source>Museum Ireland</source>
          <volume>28</volume>
          (
          <year>2021</year>
          )
          <fpage>64</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Samela</surname>
          </string-name>
          ,
          <article-title>Supplementary material of ”Investigating vividness bias in language models through art interpretations</article-title>
          ,
          <year>2024</year>
          . URL: https://doi.org/10.5281/zenodo.12723687.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>V.</given-names>
            <surname>Braun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <article-title>Thematic analysis</article-title>
          ., American Psychological Association,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Chen,
          <article-title>Personalized prompt learning for explainable recommendation</article-title>
          ,
          <source>ACM Transactions on Information Systems</source>
          <volume>41</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Dabiri</surname>
          </string-name>
          ,
          <article-title>What white people can do next: From allyship to coalition</article-title>
          ,
          <string-name>
            <surname>Penguin</surname>
            <given-names>UK</given-names>
          </string-name>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Cetinic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>She</surname>
          </string-name>
          ,
          <article-title>Understanding and creating art with ai: Review and outlook</article-title>
          ,
          <source>ACM Transactions on Multimedia Computing</source>
          , Communications, and
          <string-name>
            <surname>Applications</surname>
          </string-name>
          (TOMM)
          <volume>18</volume>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Renoust</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nakashima</surname>
          </string-name>
          ,
          <article-title>Contextnet: representation and exploration for painting classification and retrieval in context</article-title>
          ,
          <source>International Journal of Multimedia Information Retrieval</source>
          <volume>9</volume>
          (
          <year>2020</year>
          )
          <fpage>17</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mensink</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Van Gemert</surname>
          </string-name>
          ,
          <article-title>The rijksmuseum challenge: Museum-centered visual recognition</article-title>
          ,
          <source>in: Proceedings of international conference on multimedia retrieval</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>451</fpage>
          -
          <lpage>454</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>N.</given-names>
            <surname>Banar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Daelemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <article-title>Transfer learning for the visual arts: The multi-modal retrieval of iconclass codes</article-title>
          ,
          <source>ACM Journal on Computing and Cultural Heritage</source>
          <volume>16</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sartini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Baroncini</surname>
          </string-name>
          , M. van
          <string-name>
            <surname>Erp</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Tomasi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Gangemi</surname>
            ,
            <given-names>Icon:</given-names>
          </string-name>
          <article-title>An ontology for comprehensive artistic interpretations</article-title>
          ,
          <source>ACM Journal on Computing and Cultural Heritage</source>
          <volume>16</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Baroncini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Steels</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. van Trijp</surname>
          </string-name>
          ,
          <article-title>Semantic data retrieval and integration for supporting artworks interpretation through integrative narrative networks</article-title>
          ,
          <source>Proceedings of the International Workshop on Semantic Web and Ontology Design for Cultural Heritage</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <article-title>A pre-training based personalized dialogue generation model with persona-sparse data</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>34</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>9693</fpage>
          -
          <lpage>9700</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yun</surname>
          </string-name>
          , G. Kim, Mpchat:
          <article-title>Towards multimodal persona-grounded conversation</article-title>
          ,
          <source>arXiv preprint arXiv:2305.17388</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>E.</given-names>
            <surname>Marx</surname>
          </string-name>
          ,
          <article-title>Assessing bias on entity retrieval models through conjunctive fallacies</article-title>
          ,
          <source>in: 2023 IEEE 17th International Conference on Semantic Computing (ICSC)</source>
          , IEEE,
          <year>2023</year>
          , pp.
          <fpage>260</fpage>
          -
          <lpage>261</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nadeem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bethke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Reddy</surname>
          </string-name>
          , Stereoset:
          <article-title>Measuring stereotypical bias in pretrained language models</article-title>
          , arXiv preprint arXiv:
          <year>2004</year>
          .
          <volume>09456</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Lyu</surname>
          </string-name>
          ,
          <article-title>Biasasker: Measuring the bias in conversational ai system</article-title>
          ,
          <source>in: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>515</fpage>
          -
          <lpage>527</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>D.</given-names>
            <surname>Esiobu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hosseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fernandes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dwivedi-Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Presani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <surname>Robbie:</surname>
          </string-name>
          <article-title>Robust bias evaluation of large generative language models</article-title>
          ,
          <source>in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>3764</fpage>
          -
          <lpage>3814</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>P.</given-names>
            <surname>Reyero Lobo</surname>
          </string-name>
          , E. Daga,
          <string-name>
            <given-names>H.</given-names>
            <surname>Alani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <article-title>Semantic web technologies and bias in artificial intelligence: A systematic literature review</article-title>
          ,
          <source>Semantic Web</source>
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>745</fpage>
          -
          <lpage>770</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>P.</given-names>
            <surname>Reyero Lobo</surname>
          </string-name>
          , E. Daga,
          <string-name>
            <given-names>H.</given-names>
            <surname>Alani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <article-title>Enhancing hate speech annotations with background semantics (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Arnold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.-W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <article-title>Revealing persona biases in dialogue systems</article-title>
          ,
          <source>arXiv preprint arXiv:2104.08728</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>J.</given-names>
            <surname>Shin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jeong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Park</surname>
          </string-name>
          , Ask llms directly,“
          <article-title>what shapes your bias?”: Measuring social bias in large language models</article-title>
          , arXiv e-prints (
          <year>2024</year>
          ) arXiv-
          <fpage>2406</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>I. O.</given-names>
            <surname>Gallegos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Rossi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Barrow</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Tanjim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Dernoncourt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>N. K.</given-names>
          </string-name>
          <string-name>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <article-title>Bias and fairness in large language models: A survey, Computational Linguistics (</article-title>
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>79</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Nisbett</surname>
          </string-name>
          , L. Ross,
          <article-title>Human inference: Strategies and shortcomings of social judgment (</article-title>
          <year>1980</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>S.</given-names>
            <surname>Biderman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schoelkopf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sutawika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Abbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Aji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Ammanamanchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Black</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clive</surname>
          </string-name>
          , et al.,
          <article-title>Lessons from the trenches on reproducible evaluation of language models</article-title>
          ,
          <source>arXiv preprint arXiv:2405.14782</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>