<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Graphical Identification of Gender Bias in BERT with a Weakly Supervised Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michele Dusi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Arici</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alfonso E. Gerevini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Putelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ivan Serina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Università degli Studi di Brescia</institution>
          ,
          <addr-line>Brescia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Transformer-based algorithms such as BERT are typically trained with large corpora of documents, extracted directly from the Internet. As reported by several studies, these data can contain biases, stereotypes and other properties which are transferred also to the machine learning models, potentially leading them to a discriminatory behaviour which should be identified and corrected. A very intuitive technique for bias identification in NLP models is the visualization of word embeddings. Exploiting the concept of that a short distance between two word vectors means a semantic similarity between these two words; for instance, a closeness between the terms nurse and woman could be an indicator of gender bias in the model. These techniques however were designed for static word embedding algorithms such as Word2Vec. Instead, BERT does not guarantee the same relation between semantic similarity and short distance, making the visualization techniques more dificult to apply. In this work, we propose a weakly supervised approach, which only requires a list of gendered words that can be easily found in online lexical resources, for visualizing the gender bias present in the English base model of BERT. Our approach is based on a Linear Support Vector Classifier and Principal Component Analysis (PCA) and obtains better results with respect to standard PCA.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Gender Bias</kwd>
        <kwd>Ethics</kwd>
        <kwd>Fairness</kwd>
        <kwd>Model Interpretability</kwd>
        <kwd>BERT</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        With the afirmation of Artificial Intelligence in everyday user experience, many AI systems
exhibited a behaviour that can be legally defined, when acted by humans, discriminatory.
In particular, for Natural Language Processing (NLP) and Machine Learning techniques, the
problem of algorithmic discrimination is mainly caused by prejudiced data involved in the
learning process and it produces an uneven outcome for demographic minorities, i.e. subgroups
of people difering by gender, race, religion, sexual orientation, disability, etc [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In order to
solve this issue, at first we have to identify the presence of a bias, possibly with an intuitive
technique that can be understood not only by AI experts but also from the people who simply
will use the system.
      </p>
      <p>In NLP systems, an intuitive way to assess bias is visualization. As can be seen in the simple
example showed in Figure 1, more stereotypical masculine jobs such engineer or mechanic
occupy a specific region of the two-dimensional space. Instead, stereotypical feminine jobs such
electrician(1%)
mechanic(1%)</p>
      <p>engineer(2%)
as nurse or hairdresser are grouped together in another region.</p>
      <p>
        In the last few years, the state of the art for many NLP tasks has been profoundly changed by
Transformer-based algorithms [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. One of the most known of such models, BERT (Bidirectional
Encoder Representations from Transformers) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], is trained on a language modelling task into
which the model has the goal of predicting words from context. In order to do that, exactly as
in typical word embedding models such as Word2Vec [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], BERT represents words as vectors
of real numbers. However, while the former produces a unique, static representation for each
word, in BERT the representation can significantly difer depending on the entire context of the
sentence or the document into which the word appears.
      </p>
      <p>
        Considering static word embeddings, there are several studies for measuring and visualizing
the gender bias [
        <xref ref-type="bibr" rid="ref7 ref8 ref9">7, 8, 9</xref>
        ] in a intuitive way, similarly to the example in Figure 1. Considering
BERT, although several studies have shown its intrinsic bias and its diferent results for male and
female subjects in NLP tasks [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref14">10, 11, 12, 13, 14</xref>
        ], at the best of our knowledge these visualization
techniques have not been applied yet.
      </p>
      <p>
        In our opinion, this is mainly due to two factors. First, while Word2Vec vectors have a length
between 100 and 300 [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], for BERT typically the length is 768, making it harder to apply
dimensionality reduction techniques that allow an efective visualization. Second, it is proven
that while static word embedding vectors have an isotropic distribution, BERT vectors occupy a
narrow cone of the 768-dimensional space [16, 17] causing several issues with typical metrics
for measuring the distance among vectors and for reducing their dimensions [18].
      </p>
      <p>
        In this work, we focus on the problem of visualizing gender prejudice in human occupations,
analysing the word representation produced by BERT. We propose a new weakly supervised
method which requires a simple and minimal dataset for obtaining a graphical representation
of biased word embeddings by reducing the BERT vector space to a two-dimensional plane
showing the gender distortion. We use a linear Support Vector Classifier [ 19], trained on a
simple list of gendered English words (e.g. woman, man, sister, brother, etc.), to decide which
features are more involved in the gender definition. This training process does not require
any time-consuming data collecting or labelling tasks, as these words can be easily found in
many online lexical resources. Secondly, we apply the Principal Component Analysis (PCA)
[20] in order to further reduce the word representation to a two-dimensional space which can
be visualized. With respect to standard PCA, we show that our approach produces a better
visualization, providing an intuitive understanding of gender bias in human occupations, as
captured by the classical Masked Language Model for bias identification [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and by the actual
statistics of gendered employment rate for several occupations.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Work</title>
      <p>In the last few years, several scientific papers [ 21, 22] have pointed out the presence of bias or
discriminatory behavior in machine learning algorithms and in Natural Language Processing
systems. The term bias is often used in reference to an algorithm to indicate a systematic
distortion of outputs that produces unfair results, such as favoring or discriminating against
certain groups of people [23]. For an algorithm, bias is something unwelcome, the absence
of which is necessary to satisfy the fairness property. A statistical, and commonly adopted,
definition of bias is presented in [ 24] and it can be summarized as the distance between two
conditional probabilities ( |) and ( |) that a word  denoting a stereotype  will
appear in a sentence, given words  and  characterizing two distinct categories  and 
of subjects. In order to address this issue, the same work outlined a standard three-steps
approach: (1) definition, (2) identification and (3) bias correction. The focus of the current work
is visualization, therefore a part of the step 2.</p>
      <p>
        Based on this and other similar definitions, several techniques for identifying the presence of
bias in models have been developed over the years with variable efectiveness. For example, in
2017 Caliskan et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] introduced the use of associative tests (WEAT and WEFAT) to estimate
the closeness of a target word to two terminological reference groups. These tests, in addition
to depending heavily on the choice of reference terms, are specifically designed for static word
embedding algorithms, that encode each word uniquely, regardless of the sentence into which
it is embedded.
      </p>
      <p>
        Another interesting work is the study by Bolukbasi et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which detected discriminatory
behaviour in the worldwide known algorithm Word2vec [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Their identification process is
done through the study of the geometric distribution of static embeddings. Similar approaches
were presented in Zhou et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] which focuses on bilingual spaces, and in Maudslay et al. [25]
which exploits clustering algorithms for identifying stereotyping classes. The basic assumption
overall these approaches is that a geometric proximity corresponds to a semantic similarity of
the two terms.
      </p>
      <p>
        However, with the advent of more complex deep learning architectures, such as Transformer
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or BERT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and their vectorial representation of words, such assumption has been
challenged. In fact, several studies [16, 17, 18] claim that in BERT a high cosine similarity between
two words (i.e. a very high closeness in the vectorial space) do not necessarily correspond to
a semantic similarity. Moreover, in BERT there is no unique correspondence between a word
and its embedding, therefore the bias should be measured by placing the word in a pre-defined
artificial context or template, like in CEAT [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] or SEAT [26]. Consequently, these issues make
the aforementioned visualization methods, such as the ones proposed in [
        <xref ref-type="bibr" rid="ref7 ref9">7, 9</xref>
        ], not suitable for
BERT embeddings as the ones we analyse in this work.
      </p>
      <p>Thus, for BERT a typical approach to bias identification is not based on visualization. Instead,</p>
      <sec id="sec-2-1">
        <title>Phase 1</title>
        <p>R</p>
      </sec>
      <sec id="sec-2-2">
        <title>Phase 2</title>
        <p>BERT</p>
        <sec id="sec-2-2-1">
          <title>BERT</title>
          <p>trains</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>LSVC</title>
          <p>relevant
features</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>Features selection</title>
        </sec>
        <sec id="sec-2-2-4">
          <title>Features selection</title>
          <p>trains</p>
          <p>
            PCA
 ^
to the features selection; the  most relevant features related to the gender information are identified
by the LVSC model trained over the dataset  of gendered words. In the second phase, the PCA model
is trained over  and applied to , obtaining the final resulting two-dimensional vector ^ ⊆ R2.
it focuses on evaluating the behaviour of the so-called Masked Language Modeling (MLM) task
[
            <xref ref-type="bibr" rid="ref11 ref14">11, 14</xref>
            ]. For instance, in the sentence “[MASK] is a programmer”, the model could predict both
the pronouns he and she and form a correct sentence. However, if the model predicts he with a
significantly greater probability than the one associated with the prediction of she, the model
presents a gender bias for the word programmer. However, this approach based on language
modeling is generally less intuitive with respect to the word embedding visualization (like the
one we propose in this work), which can be understood by a glance also by people who are
not expert in NLP architectures. Nonetheless, as we show in Section 4, graphical and language
modeling methods can be complementary and the intuition provided by the former can be
confirmed, more quantitatively, by the latter.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology for Weakly Supervised Visualization</title>
      <p>Our objective is to visualize the distribution of BERT-encoded words, i.e. to reduce high
dimensional vectors to a 2D space, in order to highlight the gender spectrum. More specifically,
we use the base version of BERT for the English language1 which has an embedding space of
length 768. We do this in two steps: we first select the  most relevant features for gender,
and secondly we reduce them while preserving their variance. This procedure, called Weakly
Supervised Visualization (WSV) is showed in Figure 2.</p>
      <p>We apply this method to the 1678 single-word job titles appearing in the JNeidel dataset2.
Given a word  describing an occupation, we exploit BERT for calculating an embedding
representation ¯. However, given that this representation depends also on the overall context of
the sentence into which  appears, we use a basic template which will form the context of every
word in the dataset. This template is composed by three tokens: the special token [cls], which
is used by BERT to calculate an entire representation of the sentence,  and [sep], which marks
the end of the sequence. Although BERT will produce a vector for each token in this simple
sequence, we extract only the one representing , which is composed by 768 real numbers.</p>
      <p>In order to measure potential bias of the embedded representation of job titles, we need also
to encode gendered words, i.e. relationships, titles, pronouns and other expressions clearly
indicating the gender of the subject such as mother, brother, grandpa, girlfriend, he, him, she, her,
sir, madam, queen, king, etc. This list consists in 102 terms coming from the WinoBias dataset
[27], 51 for each considered gender. In order to obtain a vector representation of these words,
we apply the same technique used for the job titles. Once the gendered words (denoted as )
and the occupation words (denoted as ) are encoded, we begin the dimensionality reduction.</p>
      <p>The first phase exploits a Support Vector Classifier [ 19] with a linear kernel (LSVC). The
model is trained over the  dataset labeled with a binary classification value representing the
word gender3. The LSVC model then learns a separator hyperplane in R768: w¯ −  = 0. From
the vector w of weights, we extract the  higher absolute values, deducing the most relevant
dimensions of the vector for the gender identification. After this procedure, the remaining
768 −  features are cut of from the vectors.</p>
      <p>This features selection procedure is applied both to  and to , obtaining respectively
 ⊆ R and  ⊆ R, i.e. two compressed representations which would contain most of
the gender-related information. Please note that, in order to train the LSVC, we need only a
small list of about 100 gendered words which can be easily found online, as in the WinoBias
dataset or in lexical databases such as WordNet. No actual labelling process or data collecting is
required, therefore we claim that our approach is weakly supervised.</p>
      <p>The second phase consists in the Principal Component Analysis (PCA) [20] of . PCA
defines the optimal linear transformation  : R ↦→ R2 such that the maximum percentage of
variance in the starting samples is preserved. Thus, applying the same  to , we get ^ ⊆ R2
summarising the gender bias in the evaluation set . Given that we chose the most important
dimensions of the vector related to the gender prediction, the variance preserved by the PCA in
 should capture whether there is a gender distortion in the job titles.</p>
      <p>The WSV implementation requires to set a single hyperparameter , namely, the vectors
dimension extracted in the first phase:
 : R76→8−− ↦− LSVC</p>
      <p>R→−↦− PCA R2
As we describe in Section 4, the best choice of  are from middle values, while very low (such
as  = 2) or high ( &gt; 300) values do not obtain satisfying results.</p>
      <p>In general, we can see our approach as a linear transformation from a high dimensional vector
space to a two-dimensional plane. This requires just a single training of a simple classifier. In
our case, we have trained a LSVC under the hypothesis (later confirmed by the results, as we
show in Section 4) that the gendered words are linearly separable. However, in our opinion
this is not necessary, and diferent kernels or other models (such as XGBoost or Feed-Forward
Neural networks) can be used alongside with a technique for extracting the most important
features (e.g. SHAP [28]) and then applying the PCA. In all these cases, after than the classifier
has learned which dimensions are mostly related to the gender, this information can be exploited
3For now, we improperly simplify the social perception of gender by considering only the male and female classes.
[mask] works as a [job].
[mask] worked as a [job].
[mask] was a [job].
[mask] will soon be a [job].
[mask] has a job as [job].
[mask] is a [job].</p>
      <p>[mask] should be [job] soon.
[mask] has studied for years to become a [job].</p>
      <p>One day [mask] will be a [job].</p>
      <p>From tomorrow, [mask]’s going to work as a [job]
[mask] is studying to be a [job].</p>
      <p>[mask] has always wanted to become a [job].
for any set of words. Moreover, this whole process could be easily generalized to diferent types
of bias, as long as we provide a labeled set of training samples. As future work, we will perform
a more in-depth study considering diferent classifiers, sets of words and types of bias.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Results</title>
      <p>In this section we present the results achieved by the proposed model, comparing it to the bare
use of the Principal Component Analysis (PCA). Our Linear Support Vector Classifier has been
trained using the default hyperparameters in the implementation provided by the Scikit-learn4
library [29] and with the hinge squared loss function.</p>
      <p>Considering the 1678 jobs in the JNeidel dataset, first we calculated the gender distortion with
the standard Masked Language Modeling (MLM) technique: we chose 12 templates regarding the
career domain (Table 1) and we measure for each job the diference between the probabilities of
“[mask]” being “she” or “he”. With this procedure, we obtain a score between − 1, that indicates
an only man profession, and +1, that indicates an only female profession. From now on, we
will refer to this score as the MLM score. The MLM score is incorporated in our visualization as
follows. In Figure 3, each point represents a diferent occupation, obtained by only applying
PCA (on top) and by our method (on the bottom) selecting  = 50. The colour of each point
indicates the MLM score: jobs with a MLM score very close to − 1 are represented as a blue
point, while those ones very close to +1 are represented as a magenta point.</p>
      <p>As it can be seen in the first chart in Figure 3, the PCA results have no spatial relations with the
bias detected by the MLM score. In fact, the pink points are sparse across all the space, without
any noticeable logic. On the other hand, samples in the second chart are placed according to
the stereotypical gender with most of the pink and magenta points on the left region of the
space, highlighting the gender spectrum of prejudiced jobs. While for PCA reducing vectors
in 2D is not enough to intuitively show a bias, being only suficient to show the distortion of
the most extreme samples, our approach visualizes a more evident trend. For instance, jobs
like nurse or hostess are strongly related to the female gender and occupies the left region of
the two-dimensional space, while typically masculine jobs like priest or infantryman are in the
opposite region.
4https://scikit-learn.org/stable/modules/generated/sklearn.svm.LinearSVC.html</p>
      <p>chairman
infanptroypumoltlaricynemwaonman
midwife</p>
      <p>PCA
hostess
nuprrsieest
bankman</p>
      <p>busgirl
hostess
nurse
midwife</p>
      <p>busgirl
policewoman</p>
      <p>WSV for n = 50
chairman</p>
      <p>priest
bankman
poultryman
infantryman</p>
      <p>G
e
n
d
e
r
p
e
r
c
e
i
v
e
d
w
i
t
h
M
L
M
G
e
n
d
e
r
p
e
r
c
e
i
v
e
d
w
i
t
h
M
L
M</p>
      <p>In order to provide a more quantitative indication of the ability of PCA and WSV (selecting
diferent values of ) to show the bias, we have calculated the Pearson Correlation Coeficient,
in absolute value, among the two components extracted by both methods and the MLM score.
We present the results in Table 2. Results in terms of correlation for PCA are very low (0.09
for the first component, 0.05 for the second one), demonstrating that this method is not able
to represent the gender bias in a two-dimensional space. However, simply extracting the two
most relevant features ( = 2) without applying the PCA algorithm does not provide satisfying
results. Instead, combining the selection of a relatively small number of features with the
PCA provide the best results. In fact, the correlation between the first component (plotted
Value of 
1st comp.
2nd comp.
as the horizontal axis in Figure 3) and the MLM score with  = 50 is 0.42, confirming the
intuition provided by the plot. In general, the choice of the hyperparameter , namely the
number of elements extracted from the BERT embeddings is a fundamental vector for evaluating
the eficacy of WSV. In fact, a value too low could cut of important gender information, but
a value too high would make the selection useless and admit a lot of unrelated information.
Experimental tests (like the ones reported in Table 2) showed us that values between 20 and
100 are usually good options, however these results may depend also on the dataset considered.</p>
      <p>
        In Figure 4, we performed the same experiment but considering the WinoGender dataset [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
which is made by 60 occupations and their gender employment rate. In this case, the samples
are not coloured on the basis of the MLM score; instead, pink points indicate occupations into
which the female workers are the vast majority, while blue points represent jobs typically done
by men. For Figure 3, the diference between the left chart (visualization using only PCA) and
the right chart (WSV) suggests that our method is valid also in this case, showing more pink and
violet points on the left of the two-dimensional space, while using only PCA no clear pattern
can be identified.
      </p>
      <p>
        Considering the WinoGender dataset, we have evaluated how PCA and WSV are correlated
with the real world statistics for gender employment rate. While the most correlated component
of standard PCA has a Pearson Coeficient of 0.24, our approach reaches 0.56. Given that, for
the same dataset, the MLM score has a correlation with the gender employment rate of 0.59, this
result is particularly important. In fact, we are able to produce a visual plot which has a very
similar correlation to the one obtained by the standard technique for bias identification [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <p>We presented a new weakly supervised graphical approach to identify the gender bias in a
BERT model. The results indicate that our method gives better representations of prejudiced
gender than the standard PCA approach, ofering the possibility to easily grasp distortion.
When used on real world data, analysing the relation between the extracted bias and the gender
employment rate, it provides comparable results with respect to the commonly used Masked
Language Model method. An import characteristic of our work is that we propose an algorithm
which only needs a list of common gendered words for training, without any expensive data
collecting or labelling processes.</p>
      <p>As future work, the first thing to do will be to apply WSV to other types of bias, such as
ethnicity or religion. This can be done by identifying a set of words related to the bias in
question to quickly train our proposed weakly supervised system. However, the number of
training words could vary depending on the kind of bias and on the dificulty of the identification
process. In adapting our approach to other fields and datasets, another fundamental aspect
which has to be considered is the tuning of the hyperparameter , which could require several
trials. Another possible development will be to apply the technique to models which analyse
the Italian language. Since this language, like Spanish or French, has a rich morphology for
nouns and adjectives and often has terms directly specifying the gender (such as studente or
studentessa, which identify respectively a male and a female student), we may need to modify
our bias detection strategy. Finally, we will try to apply and to adapt our method to newer and
more complex NLP models, such as GPT.
embeddings, in: G. Kondrak, T. Watanabe (Eds.), Proceedings of the Eighth International
Joint Conference on Natural Language Processing, IJCNLP 2017, Taipei, Taiwan, November
27 - December 1, 2017, Volume 2: Short Papers, Asian Federation of Natural Language
Processing, 2017, pp. 31–36.
[16] K. Ethayarajh, How contextual are contextualized word representations? comparing the
geometry of bert, elmo, and GPT-2 embeddings, in: K. Inui, J. Jiang, V. Ng, X. Wan (Eds.),
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing
and the 9th International Joint Conference on Natural Language Processing,
EMNLPIJCNLP 2019, Hong Kong, China, November 3-7, 2019, Association for Computational
Linguistics, 2019, pp. 55–65.
[17] S. Rajaee, M. T. Pilehvar, An isotropy analysis in the multilingual BERT embedding
space, in: S. Muresan, P. Nakov, A. Villavicencio (Eds.), Findings of the Association for
Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022, Association for
Computational Linguistics, 2022, pp. 1309–1316.
[18] N. Reimers, I. Gurevych, Sentence-bert: Sentence embeddings using siamese bert-networks,
in: K. Inui, J. Jiang, V. Ng, X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical
Methods in Natural Language Processing and the 9th International Joint Conference on
Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7,
2019, Association for Computational Linguistics, 2019, pp. 3980–3990.
[19] C. Cortes, V. Vapnik, Support-vector networks, Mach. Learn. 20 (1995) 273–297.
[20] K. P. F.R.S., Liii. on lines and planes of closest fit to systems of points in space, The London,</p>
      <p>Edinburgh, and Dublin Philosophical Magazine and Journal of Science 2 (1901) 559–572.
[21] L. Sweeney, Discrimination in online ad delivery, Commun. ACM 56 (2013) 44–54.
[22] I. Zliobaite, A survey on measuring indirect discrimination in machine learning, CoRR
abs/1511.00148 (2015).
[23] D. Rozado, Wide range screening of algorithmic bias in word embedding models using
large sentiment lexicons reveals underreported bias types, PLOS ONE 15 (2020) 1–26.
[24] I. Garrido-Muñoz, A. Montejo-Ráez, F. Martínez-Santiago, L. A. Ureña-López, A survey on
bias in deep nlp, Applied Sciences 11 (2021). doi:10.3390/app11073184.
[25] R. H. Maudslay, H. Gonen, R. Cotterell, S. Teufel, It’s all in the name: Mitigating gender bias
with name-based counterfactual data substitution, in: K. Inui, J. Jiang, V. Ng, X. Wan (Eds.),
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing
and the 9th International Joint Conference on Natural Language Processing,
EMNLPIJCNLP 2019, Hong Kong, China, November 3-7, 2019, Association for Computational
Linguistics, 2019, pp. 5266–5274.
[26] C. May, A. Wang, S. Bordia, S. R. Bowman, R. Rudinger, On measuring social biases in
sentence encoders, in: J. Burstein, C. Doran, T. Solorio (Eds.), Proceedings of the 2019
Conference of the North American Chapter of the Association for Computational
Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7,
2019, Volume 1 (Long and Short Papers), Association for Computational Linguistics, 2019,
pp. 622–628.
[27] J. Zhao, T. Wang, M. Yatskar, V. Ordonez, K.-W. Chang, Gender bias in coreference
resolution: Evaluation and debiasing methods, in: Proceedings of the 2018 Conference of the
North American Chapter of the Association for Computational Linguistics: Human
Language Technologies, Volume 2 (Short Papers), Association for Computational Linguistics,
New Orleans, Louisiana, 2018, pp. 15–20.
[28] S. M. Lundberg, S. Lee, A unified approach to interpreting model predictions, in:
I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N.
Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems 30:
Annual Conference on Neural Information Processing Systems 2017, 4-9 December
2017, Long Beach, CA, USA, 2017, pp. 4765–4774. URL: http://papers.nips.cc/paper/
7062-a-unified-approach-to-interpreting-model-predictions.
[29] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel,
P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher,
M. Perrot, E. Duchesnay, Scikit-learn: Machine learning in Python, Journal of Machine
Learning Research 12 (2011) 2825–2830.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Passaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          , Preface to the
          <source>Sixth Workshop on Natural Language for Artificial Intelligence (NL4AI)</source>
          , in: D.
          <string-name>
            <surname>Nozza</surname>
            ,
            <given-names>L. C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          , M. Polignano (Eds.),
          <source>Proceedings of the Sixth Workshop on Natural Language for Artificial Intelligence (NL4AI</source>
          <year>2022</year>
          )
          <article-title>co-located with 21th International Conference of the Italian Association for Artificial Intelligence (AI*IA</article-title>
          <year>2022</year>
          ), November 30,
          <year>2022</year>
          , CEUR-WS.org,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hovy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Prabhumoye</surname>
          </string-name>
          ,
          <article-title>Five sources of bias in natural language processing</article-title>
          ,
          <source>Lang. Linguistics Compass</source>
          <volume>15</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          , in: I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>30</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          , in: J.
          <string-name>
            <surname>Burstein</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Doran</surname>
          </string-name>
          , T. Solorio (Eds.),
          <source>Proceedings of the</source>
          <year>2019</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis</article-title>
          , MN, USA, June 2-7,
          <year>2019</year>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <source>Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Eficient estimation of word representations in vector space</article-title>
          , in: Y. Bengio, Y. LeCun (Eds.),
          <source>1st International Conference on Learning Representations, ICLR</source>
          <year>2013</year>
          , Scottsdale, Arizona, USA, May 2-
          <issue>4</issue>
          ,
          <year>2013</year>
          , Workshop Track Proceedings,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rudinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Naradowsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Leonard</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Van Durme</surname>
          </string-name>
          ,
          <article-title>Gender bias in coreference resolution, in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</article-title>
          , New Orleans, Louisiana,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bolukbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Saligrama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Kalai</surname>
          </string-name>
          ,
          <article-title>Man is to computer programmer as woman is to homemaker? debiasing word embeddings</article-title>
          , in: D. D.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sugiyama</surname>
            , U. von Luxburg,
            <given-names>I. Guyon</given-names>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems</source>
          <year>2016</year>
          , December 5-
          <issue>10</issue>
          ,
          <year>2016</year>
          , Barcelona, Spain,
          <year>2016</year>
          , pp.
          <fpage>4349</fpage>
          -
          <lpage>4357</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Bryson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <article-title>Semantics derived automatically from language corpora contain human-like biases</article-title>
          ,
          <source>Science</source>
          <volume>356</volume>
          (
          <year>2017</year>
          )
          <fpage>183</fpage>
          -
          <lpage>186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cotterell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>Examining gender bias in languages with grammatical gender</article-title>
          , in: K. Inui,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          Wan (Eds.),
          <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLPIJCNLP</source>
          <year>2019</year>
          ,
          <string-name>
            <given-names>Hong</given-names>
            <surname>Kong</surname>
          </string-name>
          , China, November 3-
          <issue>7</issue>
          ,
          <year>2019</year>
          , Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>5275</fpage>
          -
          <lpage>5283</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Stanovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <article-title>Evaluating gender bias in machine translation</article-title>
          , in: A.
          <string-name>
            <surname>Korhonen</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          <string-name>
            <surname>Traum</surname>
          </string-name>
          , L. Màrquez (Eds.),
          <source>Proceedings of the 57th Conference of the Association for Computational Linguistics</source>
          ,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          <year>2019</year>
          , Florence, Italy,
          <source>July 28- August 2</source>
          ,
          <year>2019</year>
          , Volume
          <volume>1</volume>
          :
          <string-name>
            <given-names>Long</given-names>
            <surname>Papers</surname>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1679</fpage>
          -
          <lpage>1684</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Kurita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Vyas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pareek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. W.</given-names>
            <surname>Black</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tsvetkov</surname>
          </string-name>
          ,
          <article-title>Measuring bias in contextualized word representations</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Gender Bias in Natural Language Processing</source>
          , Association for Computational Linguistics, Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>166</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>W.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan</surname>
          </string-name>
          ,
          <article-title>Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases</article-title>
          ,
          <source>in: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society</source>
          , ACM,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lauscher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lüken</surname>
          </string-name>
          , G. Glavas,
          <article-title>Sustainable modular debiasing of language models</article-title>
          , in: M.
          <string-name>
            <surname>Moens</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Specia</surname>
          </string-name>
          , S. W. Yih (Eds.),
          <source>Findings of the Association for Computational Linguistics: EMNLP</source>
          <year>2021</year>
          , Virtual Event / Punta Cana, Dominican Republic,
          <fpage>16</fpage>
          -
          <lpage>20</lpage>
          November,
          <year>2021</year>
          , Association for Computational Linguistics,
          <year>2021</year>
          , pp.
          <fpage>4782</fpage>
          -
          <lpage>4797</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bartl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nissim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gatt</surname>
          </string-name>
          ,
          <article-title>Unmasking contextual stereotypes: Measuring and mitigating bert's gender bias</article-title>
          ,
          <source>in: Proceedings of the Second Workshop on Gender Bias in Natural Language Processing</source>
          , arXiv,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>K.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          ,
          <article-title>Towards lower bounds on number of dimensions for word</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>