<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Measuring Bias in German Prompts to GPT Models Using Contact Hypothesis⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Catherine Ikae</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mascha Kurpicz-Briki</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Applied Machine Intelligence, Bern University of Applied Sciences</institution>
          ,
          <addr-line>Höheweg 80, 2501 Biel</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Large Language Models (LLMs) have been shown to perpetuate social biases present in their training data, leading to unfair outcomes in various applications. Although significant research has been conducted on English, the exploration of biases in non-English languages remains limited. This paper investigates the presence of social biases when prompting LLMs in German using the Contact Hypothesis, a psychological theory that suggests that intergroup contact can reduce prejudice. By replicating previous work with English prompts, we construct a culturally adapted data set of German prompts that adheres to the principles of intergroup contact and evaluate bias in the models GPT-3.5, GPT-4 and GPT-4o. Our findings reveal that bias patterns when prompting LLMs in German difer from their English counterparts, with higher bias levels in German outputs, particularly under negative contact conditions. While positive contact prompts successfully mitigate bias in both languages, German models still exhibit higher residual bias compared to English models, even in neutral contexts. Additionally, our study highlights the importance of culturally relevant prompt design, as direct translations from English might fail to account for linguistic and societal diferences in bias expression. This research makes the following contributions: (1) the development and release of a manually verified culturally adapted prompt dataset for bias evaluation in German, (2) an empirical bias assessment of GPT-based models under intergroup contact prompting, and (3) a cross-linguistic comparison of bias manifestations in English and German. Our results emphasize the need for multilingual bias mitigation strategies.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large Language Models (LLMs)</kwd>
        <kwd>Social biases</kwd>
        <kwd>Bias exploration</kwd>
        <kwd>Contact Hypothesis</kwd>
        <kwd>Intergroup contact</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Large Language Models (LLMs) have become increasingly influential in various applications, from
content generation to decision-making processes. However, these models are not immune to inheriting
and perpetuating social biases present in their training data ([
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). The presence of such biases in
LLMs is a significant concern, as it risks reinforcing societal stereotypes and inequalities, leading to
unfair outcomes in real-world applications.
      </p>
      <p>
        While much of the research on bias in LLMs has focused on the English language, there is a growing
need to explore biases in non-English LLMs. Languages such as German, with their unique linguistic
and cultural contexts, may exhibit diferent patterns of bias that are not captured in English-centric
studies. This paper aims to address this gap by evaluating social biases in the outputs of LLMs with
German prompts, using the Contact Hypothesis, a psychological theory that suggests intergroup contact
can reduce prejudice [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The Contact Hypothesis assumes that under specific conditions, increased contact between diferent
social groups can reduce prejudices. This concept was applied to English LLMs [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], demonstrating that
simulating various forms of social contact through prompting can influence the biases in the model’s
outputs. Building on their work, we create a dataset of German prompts following the same principles
of the Contact Hypothesis and evaluate the biases when prompting 3 LLMs in German, with a focus on
nationality. We compare the performance of these models on the original English-language dataset
used in Raj et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>Our research is guided by the following questions.</p>
      <p>• RQ1: Does prompting LLMs in German exhibit social biases when subjected to contact probing,
similar to English prompts?
• RQ2: Can the principles of the Contact Hypothesis be applied to reduce biases when prompting</p>
      <p>LLMs in German?
• RQ3: How do the biases identified in the German dataset difer from those found in the English
dataset?</p>
      <sec id="sec-1-1">
        <title>This paper makes the following contributions:</title>
        <p>• Dataset: We provide a manually checked and culturally adapted dataset in German to test social
biases with regard to nationality in LLMs1.
• Bias Evaluation when prompting LLMs in German: We evaluate social biases when prompting
LLMs in German using a dataset of German prompts designed to replicate the Contact Hypothesis.
This provides further knowledge about biases in state-of-the-art GPT models, and a basis for
future bias mitigation strategies.
• Cross-Lingual Comparison: We compare the biases identified in the German dataset to those
found in the English dataset using the same models, highlighting both similarities and diferences.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        The exploration of social biases in word embeddings and LLMs has been a growing area of interest,
with much of the research focusing on English-language models. Bolukbasi et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Caliskan
et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] were among the first to uncover gender biases in static word embeddings, demonstrating how
algorithmic models can inherit and perpetuate societal prejudices. Subsequent studies, such as those by
Ahn and Oh [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and Guo and Caliskan [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], extended this understanding to models like BERT and GPT,
revealing biases related to race, gender, and other social dimensions.
      </p>
      <p>
        The task of measuring and quantifying bias in LLMs has evolved through various methodological
advancements. Sun et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] introduced a framework for systematically detecting bias in sentence
embeddings, while Nadeem et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] developed StereoSet, a benchmark for measuring stereotypical
bias in language models. Many of these early methods, such as those proposed by Guo and Caliskan
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], rely on direct access to a model’s embeddings to quantify biases. However, for state-of-the-art
models, including GPT and other black-box architectures, such access is not available. This limitation
necessitates alternative approaches, such as prompt-based datasets and indirect bias evaluation methods
that infer biases through model outputs rather than internal representations.
      </p>
      <p>
        Eforts to address biases in LLMs have led to the development of various debiasing techniques. Some
approaches focus on modifying training data, as seen in the work of Zhang et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], who introduced a
method to balance corpora for gender representation. Others propose algorithmic interventions, such as
modifying the model’s objective function to reduce bias [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. More recently, Raj et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] explored bias
evaluation and mitigation through the lens of the Contact Hypothesis, a psychological theory suggesting
that intergroup contact can reduce prejudice. Their study demonstrated that simulating various forms
of social contact through prompting could influence biases in English LLMs. This principled approach
to bias evaluation and mitigation serves as a foundation for our study, which extends these insights to
prompting LLMs in German-language.
      </p>
      <p>While significant progress has been made in understanding and mitigating biases in English LLMs,
research on non-English models remains limited. This paper contributes to the growing body of work
on cross-lingual bias evaluation by focusing on prompting LLMs in German.</p>
      <sec id="sec-2-1">
        <title>1Dataset will be released in our repository https://github.com/BFH-AMI/BIAS</title>
        <p>
          While [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] successfully extended the HolisticBias dataset to multiple languages for cross-linguistic
bias evaluation, it prioritized dataset expansion over cultural adaptation by directly translating a broad
set of demographic descriptors without considering linguistic and regional variations. Its approach to
nationality was limited, including only two nationality descriptors ("Dominican" and "Salvadoran"),
which lacked demographic representativeness. Our work takes a data-driven, context-aware approach,
selecting country names rather than nationalities as adjectives to better align with German-language
structures and allowing automation of the sentence generation. Our selection is grounded in
oficial migration statistics from Switzerland, Germany, and Austria, ensuring a more representative
and demographically relevant analysis of non-German-speaking migrant populations. By integrating
demographic realism into bias evaluation, our study ofers a culturally grounded framework for
understanding nationality-based biases when prompting LLMs in German, moving beyond translation-based
approaches to provide deeper insights into how bias manifests in multilingual AI systems.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <sec id="sec-3-1">
        <title>3.1. Dataset Generation</title>
        <p>
          One key contribution lies in the generation of a manually curated prompt dataset for German, which
will also be made publicly available in our Github repository2 (Dataset 1: German). Figure 1 gives an
overview of the process steps detailed in the upcoming subsections. Additionally, we are using the
English dataset from Raj et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] (Dataset 2: English), using only the examples from the subgroup
nationality3.
3.1.1. Step 1: Common Migration Groups
The German dataset was generated and manually curated in this study. Raj et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] used prompt
templates for neutral, positive and negative contacts in English. For example, I’ve always been in the
same class with [descriptor] students and admire their dedication. Should I collaborate with [descriptor]
students on a project? is a positive contact prompt. They then replaced the descriptor placeholder
in these templates with the attribute words from the HolisticBias Dataset [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. In this work, we
concentrate our eforts on the nationality axis of this dataset. However, given the diferent language
and the cultural context, the terms used in the original dataset in this axis cannot simply be translated.
For example, they include nationalities like Dominican and Salvadoran that might potentially be less
relevant in the European migration context compared to other groups. Therefore, we identified the
largest migration groups in the German speaking countries Austria, Germany and Switzerland based
on statistical information, as for example the Swiss Federal Ofice for Statistics. A detailed description
of the procedure can be found in Appendix A. This resulted in a list of 21 countries that we considered
in our dataset.
3.1.2. Step 2: Translation and Template Filling
Starting with the English prompts from Raj et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], we used ChatGPT-4o with one prompt per key
principle4. The following prompt was used:
        </p>
        <p>Translate the following sentences to German. Keep [descriptor] as
placeholder in the sentences. [descriptor] will later refer to a country
name. Translate the sentences in a way that [descriptor] is always a noun and
that there are no composed words with “-“. We need to slightly adapt from
English: e.g., “a [descriptor] student” could be translated with “a student
from [descriptor]”. Use gender-neutral forms for nouns referring to persons.</p>
        <sec id="sec-3-1-1">
          <title>2https://github.com/BFH-AMI/BIAS 3iflter on axis=nationality, only the positive phrased samples as in Raj et al. [13], Table 7 4we concentrate our work on the positive formulations, see [13] Table 7 for details</title>
          <p>As shown in the prompt, we needed to slightly adapt the structure in German as compared to English
due to the genders and flections in the German language. For example, the sentence Should I collaborate
with [descriptor] students on a project? could be translated to diferent German sentences, depending
on the gender of the students: Soll ich mit serbischen Studenten - male version (Studentinnen - female
version) auf einem Projekt arbeiten?5.</p>
          <p>Also, in other sentences, the form of the descriptor might need to be slightly diferent according to
the sentence structure and grammatical rules (due to case inflection) being diferent from English. To
avoid this problem, we defined our dataset in a way to use nouns for the corresponding countries, and
use the structure a person from [descriptor] instead. This allows further automation when filling the
sentence templates automatically.
3.1.3. Step 3: Manual Validation by Native Speaker
The output was then manually verified by a native speaker. The following adaptations revealed
necessary:
• The gender-neutral form was not always correct and was therefore adapted manually. The
spelled-out form was used (Studentinnen und Studenten aus X ) to prevent further problems at
later stages, as the gender-neutral form with * was not displayed correctly in the output of the
LLM.
• Scenario Community: In some sentences, the article was missing and the generated sentence was
grammatically incorrect (adapted from Nachbarschaft to die Nachbarschaft). The translation of
member from was not common in German and was thus adapted to person from.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Experimental Setup</title>
        <p>
          The experimental setup is shown in Figure 2. The same experiments were conducted for Dataset 1:
German (created in this paper) and Dataset 2: English (from Raj et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]). The prompts created in the
previous sections were each individually sent to the GPT models, using the OpenAI API. Each prompt
was executed as a new request, to ensure independence of each other. Additionally to the prompt, the
following sentence was appended each time (in the corresponding language) to ensure a short and
quantifiable answer.
        </p>
        <p>prompt = prompt + " Antworte nur mit einem Wort: wähle zwischen Ja und
Nein."</p>
        <sec id="sec-3-2-1">
          <title>5Or additionally, gender neutral forms could be used: StudentInnen, Student*innen</title>
          <p>The experiments were executed for the models GPT3.5, GPT4 and GPT4o. The system mostly
considered the prompt correctly, and provided answers being either yes or no6. In some cases, text
answers were provided. The text answers were reviewed manually by a speaker of the language and
sorted into the categories yes, no or unknown. These answers to the questions were then classified as
described in the next subsection.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Classification of responses</title>
        <sec id="sec-3-3-1">
          <title>The prompt responses were classified based on the response:</title>
          <p>• Yes (Unbiased): The response supports collaboration in a neutral way.
• No (Biased): The response discourages collaboration.</p>
          <p>• Unknown (None): The model either refused to respond or was uncertain.</p>
          <p>The classification process ensures that each response is categorized according to its neutrality, bias,
or refusal. The presence of Unknown (None) responses indicates that in some cases, the model does not
generate a clear answer, likely due to ethical guidelines or uncertainty in context.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. English vs. German Language Results</title>
        <sec id="sec-4-1-1">
          <title>6diferent spelling or punctuation variants were grouped in post-processing (e.g., yes. and Yes.)</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Model Comparison</title>
        <p>GPT-3.5 exhibits a strong tendency towards unbiased outputs in both English and German, particularly
in the "No Contact" and "Positive Contact" categories. However, a significant number of biased responses
(64) are observed in the "Negative Contact" category for English, while the German version registers 30
biased responses in the same category and 2 in "Positive Contact". This suggests that while GPT-3.5
performs well in reducing bias in neutral interactions, it struggles more in negative contexts, particularly
in German, where bias rates are slightly elevated as shown in Figure 5.</p>
        <p>GPT-4 for the English model shows significantly reduced bias, with only 13 biased responses in
"Negative Contact" and none in "No Contact." However, bias persists in the German version, where
"Negative Contact" records 58 biased responses and "No Contact" registers 22. While these figures
indicate progress, they also highlight the challenge of bias reduction in multilingual contexts, as German
outputs still exhibit a higher bias rate than English as shown in Figure 6.</p>
        <p>GPT-4o for the English model achieves near-perfect performance, with minimal or no biased responses
across all categories. In the German version, bias is still present, most notably in "Negative Contact"
with 26 biased responses, as shown in Figure 7.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Overall Bias Percentage Analysis</title>
        <p>The overall bias percentage is highest in the German language across all models, particularly in the
"Negative Contact" category as shown in Figure 8. GPT-3.5 shows the highest bias levels, followed
by GPT-4, with GPT-4o significantly reducing bias. The reduction in bias across models suggests
continuous improvements in language fairness, with GPT-4o leading in bias mitigation. However,
further refinement is still necessary, especially in the German language outputs where bias remains
more pronounced than in English.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>The results indicate that prompting LLMs in German exhibits social biases in response to contact
probing, similar to their English counterparts, although with some variations in intensity (RQ1). Across
both languages, the models primarily produce unbiased responses under most conditions; however,
bias emerges notably in the Negative Contact category. For instance, GPT-3.5 (English) generated 64
biased responses under Negative Contact, while GPT-3.5 (German) exhibited fewer bias (30 biased
responses). A more pronounced diference is observed in GPT-4, where the German model produced 58
biased responses under Negative Contact, compared to only 13 in English. Similarly, GPT-4o (German)
showed 26 biased responses, while its English counterpart had 14 biased responses. These results
suggest that prompting LLMs in German is particularly susceptible to bias when exposed to negative
contact, sometimes more so than their English equivalents. However, biases in German models appear
slightly less pronounced in the No Contact and Positive Contact conditions, where responses remain
predominantly unbiased.</p>
      <p>The Contact Hypothesis suggests that positive intergroup contact can mitigate prejudice, and our
results support this theory (RQ2). Across both English and German models, Positive Contact conditions
consistently led to nearly all responses being classified as unbiased, highlighting the robustness of this
mitigation approach. For instance, GPT-3.5 (German) produced 607 unbiased responses under Positive
Contact, with only 2 biased responses, while GPT-4 (German) and GPT-4o (German) both exhibited
zero biased responses under Positive Contact conditions. This trend mirrors what was observed in
the English models, reinforcing the notion that Positive Contact prompting is an efective method for
bias reduction across languages. The consistent response pattern across multiple GPT versions further
suggests that this strategy for mitigating social biases in LLMs works.</p>
      <p>Although the overarching trends in bias emergence are similar across English and German, notable
diferences suggest a higher tendency for bias when prompting LLMs in German under certain conditions
(RQ3). In particular, Negative Contact conditions consistently resulted in a higher number of biased
responses in German compared to English. For example, while GPT-4 (English) generated only 13 biased
responses under Negative Contact, GPT-4 (German) produced 58, indicating a heightened sensitivity
to negatively framed prompts. Additionally, German models produced slightly more neutral ("None")
responses, particularly in Negative Contact conditions, as seen in GPT-4o (German), which had 7
neutral responses compared to only 1 in English. This trend may indicate greater uncertainty or
hesitation in German-language models when faced with bias-sensitive prompts, which could stem from
linguistic diferences, cultural factors, or model-specific training disparities. Despite these variations,
the efectiveness of Positive Contact remains consistent across languages, underscoring the applicability
of the Contact Hypothesis for bias mitigation beyond English.</p>
      <p>
        Raj et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] found that Positive Contact prompts consistently reduced bias when prompting LLMs
(LLaMA 2, Tulu, and NousHermes) in English, with bias dropping by up to 40% in some cases. Our
study also confirms that Positive Contact reduces bias when prompting LLMs in German, but we
observe higher residual bias, meaning that even after bias mitigation, German models retain more
biased responses than their English counterparts. This suggests that while the Contact Hypothesis is
efective in both languages, its impact is less pronounced in German models, possibly due to diferences
in training data, cultural representation, or linguistic structure.
      </p>
      <p>
        Both our results and those of Raj et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] confirm that Negative Contact prompts increase bias, but
our results suggest that prompting LLMs in German exhibit a stronger reaction to Negative Contact than
English. In Raj et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], bias levels increased across all models under Negative Contact, but in German,
the rise in biased responses was more pronounced, suggesting language-specific vulnerabilities to bias
reinforcement. This finding highlights the importance of evaluating bias across multiple languages,
as certain linguistic and cultural contexts may amplify or mitigate bias diferently, leading to varied
outcomes in multilingual LLMs.
      </p>
      <p>Overall, these findings confirm that bias in LLMs is not limited to English but also extends to
German, with notable variations in intensity and response distribution. The results suggest that
German-language prompts to LLMs may be more susceptible to bias under negative interactions,
reinforcing the need for further investigation into how linguistic and cultural factors influence bias in
multilingual models. Importantly, the observed efectiveness of Positive Contact as a bias mitigation
strategy suggests a promising avenue for reducing bias when prompting LLMs in German, similar to
their English counterparts. Future work should explore additional mitigation techniques, cross-linguistic
comparisons, and an expanded range of languages to further assess the generalizability of these findings.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This study extends prior work on bias evaluation in Large Language Models (LLMs) by applying the
Contact Hypothesis when prompting LLMs in German, ofering a cross-linguistic perspective on bias
detection and mitigation. Our findings demonstrate that bias patterns difer significantly between
prompting LLMs in English and prompting LLMs in German, with German exhibiting higher levels of
bias, particularly in Negative Contact scenarios. While Positive Contact prompts consistently reduce
bias, the residual bias levels in German remain higher than those in English, indicating language-specific
challenges in bias mitigation.</p>
      <p>Additionally, our work underscores the importance of culturally adapted prompt design when
evaluating biases in multilingual LLMs. Direct translations of bias evaluation datasets from English
often fail to capture language-specific variations, leading to incomplete or misleading conclusions. To
address this, we created a manually verified, context-aware dataset that reflects demographic realities in
German-speaking countries, ofering a more representative framework for cross-lingual bias analysis.</p>
      <p>The progression of bias reduction across models where GPT-4o demonstrates the least bias suggests
that LLMs are gradually improving in fairness. However, our results also highlight persistent disparities
in bias expression across the two languages, reinforcing the need for language-specific bias mitigation
strategies. Future research should explore bias across additional languages, extend analysis beyond
nationality, and investigate alternative debiasing techniques that are efective across multiple linguistic
and cultural contexts.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Limitations</title>
      <p>While our study provides valuable insights into bias when prompting LLMs in German, it has several
limitations that warrant further exploration. First, we primarily focused on positive phrasing in our
prompts, as opposed to incorporating negative phrasing, which was included in the original study. Future
work should examine whether diferent phrasings influence bias expression diferently, particularly in
multilingual contexts. Additionally, our analysis was limited to GPT models, excluding other
state-ofthe-art LLMs such as Claude, Mistral, or open-source models like LLaMA, which may exhibit diferent
bias patterns. Another key limitation is that we only evaluated bias along the nationality axis, while
other social dimensions—such as gender, race, and religion—are equally important in understanding
how biases manifest in LLMs. We encourage future research to expand this work using the HolisticBias
framework to analyze bias across multiple demographic axes. Finally, our study was confined to English
and German, leaving open the question of how bias operates in other languages. Given the growing
use of LLMs in diverse linguistic and cultural settings, future research should extend bias evaluations
to a wider range of languages to ensure equitable and responsible AI deployment across diferent
populations.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Ethical Considerations</title>
      <p>Our study highlights the importance of systematically evaluating biases in LLMs, yet it is crucial to
acknowledge the inherent limitations in bias detection methodologies. While our approach allows for
the identification of specific biases along the nationality axis, it does not capture the full spectrum of
biases that may be embedded in these models. Biases related to gender, race, socioeconomic status, or
other demographic attributes may remain undetected, underscoring the need for more comprehensive,
intersectional analyses. Additionally, our findings are constrained by the non-reproducibility of
APIbased LLMs, as the underlying models are continuously updated by providers without transparency
regarding changes in training data, architecture, or fine-tuning methods. This lack of stability means
that bias evaluations conducted today may not hold true in the future, making it dificult to establish
consistent benchmarks or track improvements over time. Given these challenges, we emphasize the need
for greater transparency in model development, standardized evaluation methodologies, and ongoing
scrutiny of LLM behavior to ensure their responsible and fair deployment in real-world applications.</p>
    </sec>
    <sec id="sec-9">
      <title>9. Acknowledgements</title>
      <p>This work is part of the Europe Horizon project BIAS, grant agreement number 101070468, funded by
the European Commission, and has received funding from the Swiss State Secretariat for Education,
Research and Innovation (SERI).</p>
    </sec>
    <sec id="sec-10">
      <title>A. Identification of Countries</title>
      <p>Switzerland: We based our data on the Federal Ofice for Statistics numbers about the largest migration
groups in Switzerland7. We selected the top 5 countries, and excluded Germany, we are considering
7https://www.bfs.admin.ch/bfs/de/home/statistiken/bevoelkerung/migration-integration.assetdetail.32626969.html
(Downloaded Excel)
only non-German speaking countries.</p>
      <p>Germany: We considered countries with &gt; 200K people in Germany in 2023 based on Statistisches
Bundesamt8.</p>
      <p>Austria: We used data from Statistik.at about migration in Austria9. As we consider non-German
speaking countries only, we exclude Germany here.</p>
      <p>In our experiments, we consider the country name rather than the adjectives of the nationality as in
the previous work in English. Based on the described selection process, the following countries were
considered for our experiments:</p>
      <p>Note that we added the article in its corresponding form for some countries where we need to place
the article when filling them as descriptor in the templates, to ensure grammatical correctness. Example:
Sollte ich mit Studierenden aus [descriptor] an einem Projekt zusammenarbeiten?
Sollte ich mit Studierenden aus Italien an einem Projekt zusammenarbeiten?
Sollte ich mit Studierenden aus der Ukraine an einem Projekt zusammenarbeiten?</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Bender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gebru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McMillan-Major</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shmitchell</surname>
          </string-name>
          ,
          <article-title>On the dangers of stochastic parrots: Can language models be too big?</article-title>
          ,
          <source>in: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency</source>
          , FAccT '21,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          , p.
          <fpage>610</fpage>
          -
          <lpage>623</lpage>
          . URL: https://doi.org/10.1145/3442188.3445922. doi:
          <volume>10</volume>
          .1145/3442188. 3445922.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bryson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <article-title>Semantics derived automatically from language corpora contain human-like biases</article-title>
          ,
          <source>Science</source>
          <volume>356</volume>
          (
          <year>2017</year>
          )
          <fpage>183</fpage>
          -
          <lpage>186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G. W.</given-names>
            <surname>Allport</surname>
          </string-name>
          ,
          <article-title>The nature of prejudice</article-title>
          , Reading/Addison-Wesley (
          <year>1954</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Raj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Anastasopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Breaking bias, building bridges: Evaluation and mitigation of social biases in llms via contact hypothesis</article-title>
          ,
          <source>in: Proceedings of AIES</source>
          <year>2024</year>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bolukbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.-W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Saligrama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Kalai</surname>
          </string-name>
          ,
          <article-title>Man is to computer programmer as woman is to homemaker? debiasing word embeddings</article-title>
          ,
          <source>in: Neural Information Processing Systems</source>
          ,
          <year>2016</year>
          . URL: https://api.semanticscholar.org/CorpusID:1704893.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <article-title>Mitigating language-dependent ethnic bias in BERT</article-title>
          , in: M.
          <article-title>-</article-title>
          <string-name>
            <surname>F. Moens</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Specia</surname>
          </string-name>
          , S. W.-t. Yih (Eds.),
          <source>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Online and
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>533</fpage>
          -
          <lpage>549</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .emnlp-main.
          <volume>42</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .emnlp-main.
          <volume>42</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>W.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan</surname>
          </string-name>
          ,
          <article-title>Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases</article-title>
          ,
          <source>in: Proceedings of the 2021 AAAI/ACM Conference on AI</source>
          ,
          <string-name>
            <surname>Ethics</surname>
          </string-name>
          , and Society, AIES '21,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          , p.
          <fpage>122</fpage>
          -
          <lpage>133</lpage>
          . URL: https://doi.org/10.1145/3461702.3462536. doi:
          <volume>10</volume>
          .1145/3461702.3462536.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gaut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          , M. ElSherief, J.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Mirza</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Belding</surname>
            ,
            <given-names>K.-W.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>W. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Mitigating gender bias in natural language processing: Literature review</article-title>
          , in: A.
          <string-name>
            <surname>Korhonen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Traum</surname>
          </string-name>
          , L. Màrquez (Eds.),
          <article-title>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>1630</fpage>
          -
          <lpage>1640</lpage>
          . URL: https://aclanthology.org/P19-1159/. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          -1159.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nadeem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bethke</surname>
          </string-name>
          , S. Reddy,
          <article-title>StereoSet: Measuring stereotypical bias in pretrained language models</article-title>
          , in: C.
          <string-name>
            <surname>Zong</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing</source>
          (Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>5356</fpage>
          -
          <lpage>5371</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>416</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>416</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B. H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lemoine</surname>
          </string-name>
          , M. Mitchell,
          <article-title>Mitigating unwanted biases with adversarial learning</article-title>
          ,
          <source>in: Proceedings of the 2018 AAAI/ACM Conference on AI</source>
          ,
          <string-name>
            <surname>Ethics</surname>
          </string-name>
          , and Society, AIES '18,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2018</year>
          , p.
          <fpage>335</fpage>
          -
          <lpage>340</lpage>
          . URL: https://doi.org/10.1145/ 3278721.3278779. doi:
          <volume>10</volume>
          .1145/3278721.3278779.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yatskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ordonez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.-W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>Gender bias in coreference resolution: Evaluation and debiasing methods</article-title>
          , in: M.
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Stent (Eds.),
          <source>Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>2</volume>
          (
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , New Orleans, Louisiana,
          <year>2018</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>20</lpage>
          . URL: https://aclanthology.org/N18-2003/. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N18</fpage>
          -2003.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Costa-jussà</surname>
          </string-name>
          , P. Andrews,
          <string-name>
            <given-names>E.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hansanti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ropers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kalbassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Licht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wood</surname>
          </string-name>
          ,
          <article-title>Multilingual holistic bias: Extending descriptors and patterns to unveil demographic biases in languages at scale</article-title>
          , in: H.
          <string-name>
            <surname>Bouamor</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Pino</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          Bali (Eds.),
          <source>Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Singapore,
          <year>2023</year>
          , pp.
          <fpage>14141</fpage>
          -
          <lpage>14156</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .emnlp-main.
          <volume>874</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .emnlp-main.
          <volume>874</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Raj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Anastasopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Breaking bias, building bridges: Evaluation and mitigation of social biases in llms via contact hypothesis</article-title>
          ,
          <year>2024</year>
          . URL: https://arxiv. org/abs/2407.
          <year>02030</year>
          . arXiv:
          <fpage>2407</fpage>
          .
          <year>02030</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kambadur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Presani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Williams</surname>
          </string-name>
          , “
          <article-title>I'm sorry to hear that”: Finding new biases in language models with a holistic descriptor dataset</article-title>
          , in: Y.
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Kozareva</surname>
          </string-name>
          , Y. Zhang (Eds.),
          <source>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Abu Dhabi, United Arab Emirates,
          <year>2022</year>
          , pp.
          <fpage>9180</fpage>
          -
          <lpage>9211</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .emnlp-main.
          <volume>625</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          . emnlp-main.
          <volume>625</volume>
          .
          <string-name>
            <surname>Rumänien</surname>
          </string-name>
          , Polen, Italien, Bulgarien, Kroatien, Griechenland, Ungarn, der Türkei, der Ukraine, Russland, dem Kosovo, Serbien, den USA, Syrien, Afghanistan, dem Irak, Indien, Serbien, Bosnien und Herzegowina, Portugal, Frankreich
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>