<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>From Inclusive Language to Inclusive AI: A Proof-of-Concept Study into Pre-Trained Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marion Bartl</string-name>
          <email>marion.bartl@insight-centre.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Susan Leavy</string-name>
          <email>susan.leavy@ucd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight SFI Research Centre for Data Analytics</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information and Communication Studies, University College Dublin</institution>
          ,
          <addr-line>Belfield, Dublin 4</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Pre-trained language models are central to today's AI landscape. However, harmful and outdated gender stereotypes can be learned from training data and ingrained into these models. Since pre-trained models are used in many everyday language-based technologies, the deployment of unchecked systems risks the perpetuation of stereotypical and heteronormative conceptualizations of gender in society and could result in biased AI-driven decisions. In this work, we present a study into the efects of data curation to mitigate such gender bias. We use language that counteracts male-centric expressions and structures in favor of inclusivity across all gender identities. This line of interdisciplinary research has received little attention in NLP in the past, despite the fact that gender-inclusive language has been a central tenet within feminist linguistics over five decades. For this study we rewrite gender-specific pronouns using the gender-neutral they pronoun and replace gendered role nouns for gender-inclusive variants. Our findings show a reduction in gender stereotyping for English word embedding models and a disruption of latent gender associations of gender-neutral words. This work demonstrates how incorporating principles of gender inclusive language can mitigate risks of bias in AI.</p>
      </abstract>
      <kwd-group>
        <kwd>gender-inclusive language</kwd>
        <kwd>feminist AI</kwd>
        <kwd>gender bias</kwd>
        <kwd>pre-trained models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Language models significantly impact society. They are ubiquitous in applications ranging from search
engines to hiring systems. State-of-the-art models like GPT-4 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and LLama2 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] dominate current
research due to their high performance. However, earlier models, such as classic pre-trained embeddings
(Word2Vec [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], GloVe [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) and smaller-scale language models (BERT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]), remain in industrial use for
their cost-efectiveness due to fast computation and memory eficiency [
      </p>
      <p>CEUR</p>
      <p>ceur-ws.org</p>
      <p>original text
after rewriting</p>
      <p>As a fireman, Zachary is always ready to help people, but since his parents’</p>
      <p>relationship was marked by conflict, he is opposed to commitments.</p>
      <p>As a firefighter, Zachary is always ready to help people, but since their parents’</p>
      <p>
        relationship was marked by conflict, they are opposed to commitments.
up by LLMs through fine-tuning [ 17]. However, fine-tuning an LLM necessarily invites interference
from the pre-trained model, which might obscure conclusions on how gender-inclusive language is
incorporated into model representations of gender. This work therefore presents a foundation-level
proof-of-concept study with classic pre-trained embeddings. These allow us to train a model with
gender-neutral text from scratch. By contrast, training an LLM from scratch goes beyond our and many
other institution’s resources [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Further, word embedding models might still be used in small-scale
industry settings due to their low computational costs, which makes them relevant [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        We train two Word2Vec embedding models [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] on unchanged vs. gender-neutral English text,
additionally comparing against a common post-hoc debiasing technique [19]. The code for our experiments
is openly accessible1. In the experiments we find that the use of gender-neutral terminology reduces
gender stereotyping as measured by the Word Embedding Association Test [20] and the Embedding
Coherence Test [21] as well as reducing latent gender information in the embeddings of gender-neutral
words. These results demonstrate how incorporating principles of gender-inclusive language, which
were designed to help people avoid bias or discrimination in how they speak or write, can have the
same efect on how gender is represented in word embedding models.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <sec id="sec-2-1">
        <title>2.1. Data Collection</title>
        <p>Our experiments were conducted on a corpus introduced by [22], the Small Heap. The corpus is made up
of random subsections of three popular LLM training corpora: OpenWebText2 [23] (50%), CC-News [24]
(30%) and English Wikipedia (20%). The final sub-corpus contains ~250 million tokens, or 1.5 GB of text.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Gender-neutral Rewriting</title>
        <p>
          The corpus was edited using the NeuTralRewriter [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. This involved replacing gender-specific pronouns
(he, she, him etc.) with the corresponding variant of the gender-neutral pronoun they. Additionally, 91
gender-specific nouns ( headmaster, mankind, etc.) including plural and spelling variants were replaced
by neutral versions (principal, humankind, etc.; for full set see 14). Table 1 shows an example of rewritten
text.
        </p>
        <p>
          There are two implementations of the NeuTral Rewriter, a rule-based version and a neural, machine
translation-based model. While the neural model performed better in the original experiments [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], it
proved to be very susceptible to noise in our data (email addresses, digits, etc.) as well as low-frequency
words, often translating them into unintelligible text. We therefore used the rule-based implementation,
which uses a combination of word, part-of-speech and dependency information to derive the correct
replacement of pronouns.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Embedding Models</title>
        <p>
          In order to evaluate the efects of gender-neutral language on representations of gender within the
corpus, we built three diferent Word2Vec models [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The first was trained on the original and the
second on the rewritten corpus. Each Word2Vec model was trained using the Continuous Bag of Words
(CBOW) algorithm with the default hyperparameters of the gensim library’s Word2Vec class [25]. The
1https://github.com/marionbartl/ILIA
third model was created by performing hard debiasing [19] on our original model in order to compare
our method to an existing, model-based debiasing method. Hard debiasing modifies embeddings in such
a way that gender-neutral words (e.g. babysit) are equidistant to gendered word pairs (e.g. grandfather
– grandmother ). Additionally, the gendered component of embeddings of gender-neutral words, as
defined by what is termed the ‘gender subspace’, is set to zero.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Bias Evaluation</title>
        <p>The three trained embedding models were analyzed for underlying gender bias using three methods.
Previous research found that bias measures are not always consistent [26, 27]. Using a composition of
metrics therefore allows for a more comprehensive evaluation.</p>
        <p>
          The Word Embedding Association Test (WEAT) is one of the most commonly applied bias
measures for word embeddings [20]. The test is modelled after a psychological assessment, the Implicit
Associations Test [28], and measures bias by computing the mean association between two sets of target
and attribute words. We used a WEAT implementation by Lauscher et al. [
          <xref ref-type="bibr" rid="ref17">29</xref>
          ]. Each WEAT test (i.e.
the specific combination of target and attribute terms) is identified as W  with  corresponding to its
position in the original WEAT paper. W9 and W10 were added by Lauscher et al. [
          <xref ref-type="bibr" rid="ref17">29</xref>
          ]. We additionally
added two tests using attribute words related to male- and female-stereotypical professions (W ) as
well as words related to computer science and childcare (W , cf. Table 4).
        </p>
        <p>Clustering and Classification into two groups was used by Gonen and Goldberg [26] to show that
embedding spaces retain gender information despite application of debiasing. We measured cluster
integrity after K-Means clustering (averaged over 50 runs) as well as classification accuracy with an
SVM (trained for 20 epochs) in order to find out how well gender information can be recovered from
the embedding space. We use the original model’s 500 most male-/female-biased words according to
their similarity to the element-wise mean of the male/female attribute embeddings  1 and  2 of W8 (cf.
Table 5).</p>
        <p>
          The Embedding Coherence Test (ECT) calculates distances between two sets of gendered target
words  1, 2 and a set of attribute words  that relate to a societal gender imbalance (e.g. captain,
football) [21]. Instead of relying on the absolute distances, the ECT calculates the Spearman coeficient
between the ranked distances for  1 and  vs.  2 and  . A high coeficient indicates similar ranks
between the two gendered sets, signifying reduced bias. We used the ECT implementation by [
          <xref ref-type="bibr" rid="ref17">29</xref>
          ].
        </p>
        <p>
          Semantic Quality is evaluated following Lauscher et al. [
          <xref ref-type="bibr" rid="ref17">29</xref>
          ], by using the similarity benchmarks
SimLex-999 [
          <xref ref-type="bibr" rid="ref18">30</xref>
          ] and WordSim-353 [
          <xref ref-type="bibr" rid="ref19">31</xref>
          ] and computing the Pearson and Spearman correlation
coefifcients between the benchmark term-pair similarities and cosine similarities of the corresponding
embedding pairs of the respective models.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Findings and Discussion</title>
      <p>We will discuss the results for our chosen gender bias metrics: WEAT [20], ECT [21] and Clustering
and Classification [ 26]. Finally, we contextualize these findings with the performance of our models on
two semantic quality benchmarks.</p>
      <p>WEAT: Results indicated a reduction in gender bias related to stereotypical associations of women
with arts, domestic work, and childcare, and men with (computer) science, maths, and careers,
respectively. All five tests measuring gender bias, W6, W7, W8, W  , and W , showed a reduction in the
statistic after rewriting (Table 2). W additionally shows that there is a reduction in the association
of feminine/masculine words with traditionally gendered professions. Comparing the WEAT scores
after rewriting to the hard debiased embeddings, one can see that the scores for the hard debiased
embeddings are approaching zero, which indicates equal association of male/female attributes with
the respective targets. Thus, on one hand, training with gender-neutral language generally leads to
a reduction in WEAT bias scores, indicating that this change in the language can lead to a reduction
in associations based on stereotypes. On the other hand, stereotyped associations can be specifically
targeted and mostly removed post-hoc.</p>
      <p>#
W6
W7
W8
W
W
(a) ECT results marked * significant with  &lt; 0.05 . CS
= computer science.</p>
      <sec id="sec-3-1">
        <title>Pearson</title>
      </sec>
      <sec id="sec-3-2">
        <title>Spearman</title>
      </sec>
      <sec id="sec-3-3">
        <title>SimLex 999</title>
      </sec>
      <sec id="sec-3-4">
        <title>WordSim 353</title>
      </sec>
      <sec id="sec-3-5">
        <title>SimLex 999 WordSim 353 pre</title>
        <p>(b) Semantic quality of W2V embeddings before and
after re-writing. All results significant with  &lt; 0.01 .</p>
        <p>ECT: A reduction in bias was demonstrated in relation to gendered associations with professions.
Table 3a shows that for arts vs science and profession attributes, the ECT scores increase, both after
rewriting and hard debiasing. However, the ECT scores show a higher increase for the hard debiased
model, suggesting an advantage of this method over neutral rewriting. For the computer science vs.
childcare attributes however, both methods show a reduction in ECT scores, which could indicate that
neither neutral rewriting nor hard debiasing are suficiently afecting words in these semantic fields.</p>
        <p>Clustering and Classification: Our results demonstrated that while the embeddings clearly encode
gender information that is very salient to a binary classifier, rewriting with gender-neutral terminology
has a more comprehensive efect than focusing on removing a limited ‘gender subspace’, as done by
hard debiasing. In fact, hard debiasing improves the clustering accuracy, indicating that gender in word
embeddings is more intricately encoded than can be captured by a gender subspace. These results mirror
the findings of Gonen and Goldberg [26]. Both clustering and classification accuracy marginally decline
in the model trained on gender-neutral text, as shown in Table 3a. The SVM shows a very high accuracy
of 98% in separating words with male and female direct bias in the unchanged and hard-debiased model,
which is reduced to 96% in the gender-neutral model.</p>
        <p>
          Semantic Quality: The semantic quality of the word embeddings as measured by the SimLex 999 [
          <xref ref-type="bibr" rid="ref18">30</xref>
          ]
and WordSim 353 [
          <xref ref-type="bibr" rid="ref19">31</xref>
          ] benchmarks dropped only minimally by at most 0.01 points after rewriting
(Table 3b). Overall, these results fall only slightly behind larger embedding models. According to the
SimLex 999 leaderboard, a Word2Vec model trained on one billion words of Wikipedia text reached a
Spearman correlation of 0.372, which is similar to our model.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Limitations and Future Work</title>
      <p>
        The first limitation pertains to the size of the data used. Our corpus contains 250 million tokens,
which is less than 25% of the training data size for a common embedding model [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We will explore
in future work whether our findings hold for larger datasets and whether the measured reduction in
gender stereotyping in the embedding model can translate to LLMs if fine-tuned on gender-inclusive
text. Secondly, our research is focused on gender-inclusive language in English and not directly
applicable to other languages. The NeuTralRewriter [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] was specifically developed for English, and
since the specific characteristics of gender-neutral terminology are language-dependent, applying
the method to other languages would require the development of a language-specific version of the
Rewriter. We leave this to future research. A third limitation of our research lies in the erasure of
word embeddings for he/she pronouns due to the replacement with they. However, since we are
presenting a proof-of-concept study and the measurement of gender stereotyping is not dependent
on these pronouns, we accepted this. Future work could rewrite pronouns only in a percentage of
cases or only in cases where masculine/feminine pronouns are used generically. Lastly, our research is
limited by a narrow focus on binary male and female genders when assessing model bias. There is a
significant gap in NLP research regarding the incorporation of non-binary gender identities in both
measuring and mitigating bias [
        <xref ref-type="bibr" rid="ref20">32</xref>
        ]. Due to the nature of this proof-of-concept study, we adhered to
commonly employed binary metrics. Future work will need to examine progress made regarding the
integration of non-binary gender identities in embedding models through inclusive terminology.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>This research explored the efects of gender-neutral language on gender stereotyping and latent gender
information in classic embedding models. We found that training on text with gender-neutral singular
pronouns and role nouns efected a reduction in stereotyping as measured by WEAT [ 20] and ECT [21].
These reductions do not surpass those that can be achieved by targeted, post-hoc debiasing [19].
However, gender-neutral training data showed an advantage when measuring latent gender information
in embeddings through classification and clustering. This demonstrates a more comprehensive efect of
gender-neutral language in the removal of unnecessarily gendered associations, which is in line with
the aims of gender-inclusive language.</p>
      <p>While future work will need to investigate whether our results hold at scale and can be transferred
to LLMs, our exploratory findings suggest that adjusting training data to be more gender-inclusive
can improve gender representations in pre-trained models toward a more equitable conceptualization
of gender. This research presents a promising approach to the incorporation of principles of
genderinclusive language to ensure fairness and inclusivity in AI systems.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This publication has emanated from research conducted with the financial support of Science Foundation
Ireland under Grant number 12/RC/2289_P2. For the purpose of Open Access, the authors have
applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this
submission.
Machine Translation: from Theoretical Foundations to Open Challenges, 2023. URL: http://arxiv.
org/abs/2301.10075. doi:10.48550/arXiv.2301.10075, arXiv:2301.10075 [cs].
[17] H. Thakur, A. Jain, P. Vaddamanu, P. P. Liang, L.-P. Morency, Language Models Get a Gender
Makeover: Mitigating Gender Bias with Few-Shot Data Interventions, in: Proceedings of the
61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers),
Association for Computational Linguistics, Toronto, Canada, 2023, pp. 340–351. URL: https://
aclanthology.org/2023.acl-short.30.
[18] Z. Fatemi, C. Xing, W. Liu, C. Xiong, Improving Gender Fairness of Pre-Trained Language Models
without Catastrophic Forgetting, in: Proceedings of the 61st Annual Meeting of the Association for
Computational Linguistics (Volume 2: Short Papers), Association for Computational Linguistics,
Toronto, Canada, 2023, pp. 1249–1262. URL: https://aclanthology.org/2023.acl-short.108.
[19] T. Bolukbasi, K.-W. Chang, J. Y. Zou, V. Saligrama, A. T. Kalai, Man is to Computer Programmer
as Woman is to Homemaker? Debiasing Word Embeddings, in: Advances in Neural Information
Processing Systems, volume 29, Curran Associates, Inc., 2016. URL: https://proceedings.neurips.cc/
paper_files/paper/2016/hash/a486cd07e4ac3d270571622f4f316ec5-Abstract.html.
[20] A. Caliskan, J. J. Bryson, A. Narayanan, Semantics derived automatically from language corpora
contain human-like biases, Science 356 (2017) 183–186. Publisher: American Association for the
Advancement of Science.
[21] S. Dev, J. Phillips, Attenuating Bias in Word vectors, in: Proceedings of the Twenty-Second
International Conference on Artificial Intelligence and Statistics, PMLR, 2019, pp. 879–887. URL:
https://proceedings.mlr.press/v89/dev19a.html, iSSN: 2640-3498.
[22] M. Bartl, S. Leavy, From ‘Showgirls’ to ‘Performers’: Fine-tuning with Gender-inclusive Language
for Bias Reduction in LLMs, in: A. Faleńska, C. Basta, M. Costa-jussà, S. Goldfarb-Tarrant,
D. Nozza (Eds.), Proceedings of the 5th Workshop on Gender Bias in Natural Language Processing
(GeBNLP), Association for Computational Linguistics, Bangkok, Thailand, 2024, pp. 280–294. URL:
https://aclanthology.org/2024.gebnlp-1.18.
[23] L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite,
N. Nabeshima, S. Presser, C. Leahy, The Pile: An 800GB Dataset of Diverse Text for
Language Modeling, 2020. URL: http://arxiv.org/abs/2101.00027. doi:10.48550/arXiv.2101.00027,
arXiv:2101.00027 [cs].
[24] J. Mackenzie, R. Benham, M. Petri, J. R. Trippas, J. S. Culpepper, A. Mofat, CC-News-En: A
Large English News Corpus, in: Proceedings of the 29th ACM International Conference on
Information &amp; Knowledge Management, ACM, Virtual Event Ireland, 2020, pp. 3077–3084. URL:
https://dl.acm.org/doi/10.1145/3340531.3412762. doi:10.1145/3340531.3412762.
[25] R. Řehůřek, P. Sojka, Software Framework for Topic Modelling with Large Corpora, in: Proceedings
of the LREC 2010 Workshop on New Challenges for NLP Frameworks, ELRA, Valletta, Malta, 2010,
pp. 45–50.
[26] H. Gonen, Y. Goldberg, Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender
Biases in Word Embeddings But do not Remove Them, in: J. Burstein, C. Doran, T. Solorio
(Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for
Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers),
Association for Computational Linguistics, Minneapolis, Minnesota, 2019, pp. 609–614. URL:
https://aclanthology.org/N19-1061. doi:10.18653/v1/N19- 1061.
[27] S. Goldfarb-Tarrant, R. Marchant, R. Muñoz Sánchez, M. Pandya, A. Lopez, Intrinsic Bias Metrics
Do Not Correlate with Application Bias, in: Proceedings of the 59th Annual Meeting of the
Association for Computational Linguistics and the 11th International Joint Conference on Natural
Language Processing (Volume 1: Long Papers), Association for Computational Linguistics, Online,
2021, pp. 1926–1940. URL: https://aclanthology.org/2021.acl-long.150. doi:10.18653/v1/2021.
acl- long.150.
[28] A. G. Greenwald, D. E. McGhee, J. L. Schwartz, Measuring individual diferences in implicit
cognition: the implicit association test, Journal of personality and social psychology 74 (1998)
1464. Publisher: American Psychological Association.</p>
    </sec>
    <sec id="sec-7">
      <title>A. WEAT Target and Attribute Terms</title>
      <p>present study
category
male-dominated
professions
female-dominated
professions
computer
science
childcare
words
manager, executive, doctor, lawyer, programmer, scientist,
soldier, supervisor, rancher, janitor, firefighter, oficer
secretary, nurse, clerk, artist, homemaker, dancer, singer,
librarian, maid, hairdresser, stylist, receptionist, counselor
children, babysitter, daycare, homemaker, newborn, baby,
toddler, parenting</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] OpenAI, GPT-4
          <source>Technical Report</source>
          ,
          <year>2024</year>
          . URL: http://arxiv.org/abs/2303.08774. doi:
          <volume>10</volume>
          .48550/arXiv. 2303.08774, arXiv:
          <fpage>2303</fpage>
          .08774 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Izacard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Martinet</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Rozière</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Hambro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Azhar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Rodriguez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Joulin</surname>
            , E. Grave, G. Lample, LLaMA: Open and
            <given-names>Eficient</given-names>
          </string-name>
          <string-name>
            <surname>Foundation Language Models</surname>
          </string-name>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2302.13971. doi:
          <volume>10</volume>
          .48550/arXiv. 2302.13971, arXiv:
          <fpage>2302</fpage>
          .13971 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <source>Eficient Estimation of Word Representations in Vector Space</source>
          ,
          <year>2013</year>
          . URL: http://arxiv.org/abs/1301.3781, arXiv:
          <fpage>1301</fpage>
          .3781 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Manning</surname>
          </string-name>
          , Glove:
          <article-title>Global Vectors for Word Representation</article-title>
          ,
          <source>in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Doha, Qatar,
          <year>2014</year>
          , pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          . URL: http://aclweb.org/anthology/D14-1162. doi:
          <volume>10</volume>
          .3115/v1/
          <fpage>D14</fpage>
          -1162.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kenton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          ,
          <source>in: Proceedings of NAACL-HLT</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>May</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C. Ré, Contextual Embeddings:
          <article-title>When Are They Worth It?</article-title>
          , in: D.
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chai</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Schluter</surname>
          </string-name>
          , J. Tetreault (Eds.),
          <article-title>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>2650</fpage>
          -
          <lpage>2663</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>236</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          . acl-main.
          <volume>236</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I. O.</given-names>
            <surname>Gallegos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Rossi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Barrow</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Tanjim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Dernoncourt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>N. K.</given-names>
          </string-name>
          <string-name>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <article-title>Bias and Fairness in Large Language Models: A Survey, Computational Linguistics (</article-title>
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>79</lpage>
          . URL: https://doi.org/10.1162/coli_a_00524. doi:
          <volume>10</volume>
          .1162/coli_a_
          <fpage>00524</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Bender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gebru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McMillan-Major</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shmitchell</surname>
          </string-name>
          ,
          <article-title>On the dangers of stochastic parrots: Can language models be too big?</article-title>
          ,
          <source>in: FAccT 2021 - Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>610</fpage>
          -
          <lpage>623</lpage>
          . doi:
          <volume>10</volume>
          .1145/3442188.3445922, conference Proceedings.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kramer</surname>
          </string-name>
          , Feminist Linguistics and Linguistic Feminisms, in: Ellen Lewin,
          <string-name>
            <surname>Leni M. Silverstein</surname>
          </string-name>
          (Eds.),
          <article-title>Mapping Feminist Anthropology in the Twenty-First Century</article-title>
          , Rutgers University Press,
          <year>2016</year>
          , p.
          <fpage>65</fpage>
          . URL: https://go.exlibris.link/J2p0HbgK.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Saguy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Little</given-names>
            <surname>Word That Means A Lot:</surname>
          </string-name>
          <article-title>A Reassessment of Singular They in a New Era of Gender Politics</article-title>
          ,
          <source>Gender &amp; Society</source>
          <volume>36</volume>
          (
          <year>2022</year>
          )
          <fpage>5</fpage>
          -
          <lpage>31</lpage>
          . URL: http://journals.sagepub.com/ doi/10.1177/08912432211057921. doi:
          <volume>10</volume>
          .1177/08912432211057921.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Lakof</surname>
          </string-name>
          ,
          <article-title>Language and Woman's Place, Language in Society 2 (</article-title>
          <year>1973</year>
          )
          <fpage>45</fpage>
          -
          <lpage>80</lpage>
          . URL: http://www. jstor.org/stable/4166707, publisher: Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pauwels</surname>
          </string-name>
          ,
          <article-title>Linguistic Sexism and Feminist Linguistic Activism</article-title>
          , in: J.
          <string-name>
            <surname>Holmes</surname>
          </string-name>
          , M. Meyerhof (Eds.),
          <source>The Handbook of Language and Gender</source>
          , Blackwell Publishing Ltd, Oxford, UK,
          <year>2003</year>
          , pp.
          <fpage>550</fpage>
          -
          <lpage>570</lpage>
          . doi:
          <volume>10</volume>
          .1002/9780470756942.ch24.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. F.</given-names>
            <surname>Kiesling</surname>
          </string-name>
          , Language, gender, and
          <article-title>sexuality: an introduction</article-title>
          , Book, Whole,
          <volume>1</volume>
          ;1st; ed.,
          <source>Routledge</source>
          , London;New York;,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .4324/9781351042420.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E.</given-names>
            <surname>Vanmassenhove</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Emmery</surname>
          </string-name>
          , D. Shterionov, NeuTral Rewriter:
          <article-title>A Rule-Based and Neural Approach to Automatic Rewriting into Gender Neutral Alternatives</article-title>
          ,
          <source>in: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Online and
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>8940</fpage>
          -
          <lpage>8948</lpage>
          . URL: https:// aclanthology.org/
          <year>2021</year>
          .emnlp-main.
          <volume>704</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Amrhein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Schottmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sennrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Läubli</surname>
          </string-name>
          , Exploiting Biased Models to De-bias
          <string-name>
            <surname>Text</surname>
          </string-name>
          :
          <article-title>A Gender-Fair Rewriting Model</article-title>
          , in: A.
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Boyd-Graber</surname>
          </string-name>
          , N. Okazaki (Eds.),
          <source>Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>4486</fpage>
          -
          <lpage>4506</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>246</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>246</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Piergentili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Savoldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bentivogli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Negri</surname>
          </string-name>
          ,
          <article-title>Gender Neutralization for an Inclusive</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lauscher</surname>
          </string-name>
          , G. Glavaš,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Vulić</surname>
          </string-name>
          , A
          <article-title>General Framework for Implicit and Explicit Debiasing of Distributional Word Vector Spaces</article-title>
          ,
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          <volume>34</volume>
          (
          <year>2020</year>
          )
          <fpage>8131</fpage>
          -
          <lpage>8138</lpage>
          . URL: https://ojs.aaai.org/index.php/AAAI/article/view/6325. doi:
          <volume>10</volume>
          .1609/aaai.v34i05.6325, number:
          <fpage>05</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Reichart</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Korhonen, SimLex-999:
          <article-title>Evaluating Semantic Models with (Genuine) Similarity Estimation</article-title>
          ,
          <year>2014</year>
          . URL: http://arxiv.org/abs/1408.3456. doi:
          <volume>10</volume>
          .48550/arXiv.1408.3456, arXiv:
          <fpage>1408</fpage>
          .3456 [cs] version:
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>L.</given-names>
            <surname>Finkelstein</surname>
          </string-name>
          , E. Gabrilovich,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rivlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Solan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wolfman</surname>
          </string-name>
          , E. Ruppin,
          <article-title>Placing search in context: the concept revisited</article-title>
          ,
          <source>ACM Transactions on Information Systems</source>
          <volume>20</volume>
          (
          <year>2002</year>
          )
          <fpage>116</fpage>
          -
          <lpage>131</lpage>
          . URL: https://doi.org/10.1145/503104.503110. doi:
          <volume>10</volume>
          .1145/503104.503110.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>H.</given-names>
            <surname>Devinney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Björklund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Björklund</surname>
          </string-name>
          , Theories of ”Gender” in NLP Bias Research, in: ACM FAccT Conference 2022, Conference on Fairness, Accountability, and
          <string-name>
            <surname>Transparency</surname>
          </string-name>
          , Hybrid via Seoul, Soth Korea, June 21-14,
          <year>2022</year>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>T.</given-names>
            <surname>Manzini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Yao</given-names>
            <surname>Chong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. W.</given-names>
            <surname>Black</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tsvetkov</surname>
          </string-name>
          ,
          <article-title>Black is to Criminal as Caucasian is to Police: Detecting and Removing Multiclass Bias in Word Embeddings, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics</article-title>
          , Minneapolis, Minnesota,
          <year>2019</year>
          , pp.
          <fpage>615</fpage>
          -
          <lpage>621</lpage>
          . URL: https://aclanthology.org/N19-1062. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          - 1062.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>