<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Measuring Gender Stereotype Reinforcement in Information Retrieval Systems*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Fabris</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alberto Purpura</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianmaria Silvello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gian Antonio Susto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <addr-line>Padua</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Can we measure the tendency of an Information Retrieval (IR) system to reinforce gender stereotypes in its users? In this abstract, we define the construct of Gender Stereotype Reinforcement (GSR) in the context of IR and propose a measure for it based on Word Embeddings. We briefly discuss the validity of our measure and summarize our experiments on diferent families of IR systems. IIR 2021 - 11th Italian Information Retrieval Workshop, September 13-15, 2021, Bari, Italy " fabrisal@dei.unipd.it (A. Fabris); purpuraa@dei.unipd.it (A. Purpura); silvello@dei.unipd.it (G. Silvello); gianantonio.susto@unipd.it (G. A. Susto) © 2021 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). CPWrEooUrckResehdoinpgs 2IhStTpN:/c1e6u1r3-w-h0s.o7r3g e CprEeUseRnWtboirnkasrhyofpraPmroincgeeodfinggesnd(CerEiUsRa-WcoSn.soerqgu)ence of this fact and a clear limitation of our work. This is a common weakness for work in this space; addressing it is far from trivial. * Extended abstract of Fabris et al. [4].</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Fairness</kwd>
        <kwd>Gender Stereotypes</kwd>
        <kwd>Information Retrieval</kwd>
        <kwd>Search Engines</kwd>
        <kwd>Word Embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Search Engines (SEs) increasingly act as the gatekeepers of information. Their role in information
access is undisputed, with a user base exceeding 90% of all people connected to the internet
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. SEs inevitably influence users, helping them map concepts and link entities across queries
and documents. For this reason, they can play an important role in countering or reinforcing
stereotypical associations [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Stereotypes are generalised beliefs about groups of individuals, held widely in a population
of interest. They arise from a co-occurrence of features, such as membership to a group and
display of certain traits and roles. The extent to which an individual believes a stereotypical
trait to be common in a given group is often measured through an association test between
groups and traits [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Male and female are highly salient categories in human cognition, available from an early age
for stereotypical associations.2 As a result, western societies maintain a wide range of gender
stereotypes, relating e.g. to professions, career, competence, care, predisposition for science
and mathematics. The same stereotypes are also found in artifacts and technology produced by
the same societies. For example, the search results of popular image SEs were found to contain
gender-stereotypical associations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and to influence users’ cognition accordingly [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Only
recently, novel approaches to measure gender bias in text-based SEs have been proposed [
        <xref ref-type="bibr" rid="ref4 ref7">4, 7</xref>
        ].
-0.4
      </p>
      <p>
        -0.2 0.2 g(w) 0.4 x
him man jobascg_amesrncecieeyrnce commaruftansmioinly jobs_f woman her
(a) Gender direction
In this work we provide an overview of the Gender Stereotype Reinforcement (GSR) construct
and measure of Fabris et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. GSR: Construct, Measure and Validity</title>
      <p>
        In the context of IR, we define GSR as a SE’s tendency to reinforce (or counter)
genderstereotypical associations in its users. Direct measurement of this construct would require
impractical longitudinal user studies of counterfactual nature. Fabris et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] propose a
computational approach, based on Word Embeddings (WEs).
      </p>
      <p>
        Indeed, WEs have been found to reliably encode several gender stereotypes [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], typically
along a single direction of the embedded space. More precisely, Bolukbasi et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] show
how to isolate a problematic direction, called gender subspace, where gender-related concepts
are clustered in accordance with gender stereotypes. To illustrate this concept, Figure 1a
depicts the gender direction  of Word2vec embeddings [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] along the  axis. A sample word,
 = beauty, is projected onto the gender direction, where it is closer to intrinsically female
words (her, woman) and to stereotypically female concepts than to their male counterparts.
      </p>
      <p>
        Let us indicate by () = ( · )/(||||) the function associating a word  with its
normalized scalar projection on the gender direction . By extension, (· ) maps a query 
into the average projection of its words (). Let us call () the genderedness of . Moreover,
we apply function (· ) to the ranked list of documents ℒ returned by an IR system  in
response to query . We indicate it by (ℒ) and define it as the average projection of words
in ranked documents , weighted according to the rank of each document in ℒ. In symbols
(ℒ) = ∑︀∈ℒ  · (), with  computed according to a DCG-like logarithmic discount
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].3 Given a set of  queries  and a collection of documents  available for retrieval,
we define the GSR of an IR system  over (, ) in terms of the correlation between the
genderedness of queries in (),  ∈  and the genderedness of ranked lists of documents
(ℒ) produced in response. More precisely, GSR is defined as
(, ) =
1
      </p>
      <p>1 ∑︁(() −  )((ℒ) −  ℒ),
 2()  =1
(1)
where  ,  ℒ represent the average genderedness of queries and ranked lists, while  2
() is a
scaling factor to go from correlation to slope coeficient. Informally, Equation 1 captures the
agreement between the language of queries and documents along stereotypically gendered
lines, induced by an IR system .</p>
      <p>
        A thorough assessment of the suitability of this equation to measure the GSR construct is
an important and complex endeavour undertaken in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Here we show the precision of the
projection function (· ) in finding interesting queries for the study of gender stereotypes in SEs.
Figure 1b (1c) shows the ten queries with lowest (highest) genderedness () in the Robust04
collection [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which are the most associated with women (men) according to (). Indeed these
are gendered queries, ranging from intrinsically gendered (mentioning women), to biologically
gendered (mentioning menopause), to stereotypically gendered (with quilts and child
among words in stereotypically female queries, dangerous and heroic in stereotypically
male queries).
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments and Discussion</title>
      <p>
        Our experiments on the Robut04 collection [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], omitted here for brevity, compare IR ranking
algorithms from diferent families. We consider lexical models (e.g. BM25 - [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]), semantic
models (e.g. w2v add - [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]) and neural architectures (e.g. MatchPyramid - [15]). We find that
semantic models, based on biased WEs, are most prone to reinforcing gender stereotypes, while
neural systems based on the same word representations can mitigate this efect. Indeed neural
models exhibit low GSR, comparable to that of lexical systems such as BM25. Moreover, we
test the reliability of these conclusions by measuring GSR according to two diferent sets of
WEs (Word2Vec [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and fastText [16]), finding strong agreement between the two. Finally,
we assessed the impact of debiasing WEs [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] on downstream IR tasks. By measuring system
performance and GSR both before and after debiasing, we find this approach to be superficial
and insuficient to reduce the tendency of an IR system to reinforce gender stereotypes.
      </p>
      <p>
        3To be precise, (ℒ) should be query-dependent [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Here we neglect this aspect.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>Part of this work was supported by MIUR (Italian Minister for Education) under the initiative
"Departments of Excellence" (Law 232/2016).
[15] L. Pang, Y. Lan, J. Guo, J. Xu, S. Wan, X. Cheng, Text matching as image recognition, in:</p>
      <p>Proc. of AAAI 2016, AAAI Press, 2016, p. 2793–2799.
[16] A. Joulin, E. Grave, P. Bojanowski, T. Mikolov, Bag of tricks for eficient text classification,
in: Proceedings of the 15th Conference of the European Chapter of the Association for
Computational Linguistics: Volume 2, Short Papers, ACL, 2017, pp. 427–431.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Purcell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Brenner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rainie</surname>
          </string-name>
          , Search engine use
          <year>2012</year>
          ,
          <year>2012</year>
          . URL: https: //www.pewresearch.org/internet/wp-content/uploads/sites/9/media/Files/Reports/ 2012/PIP_Search_Engine_Use_
          <year>2012</year>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. U.</given-names>
            <surname>Noble</surname>
          </string-name>
          ,
          <article-title>Algorithms of oppression: How search engines reinforce racism</article-title>
          , NYU Press,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Greenwald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>McGhee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          ,
          <article-title>Measuring individual diferences in implicit cognition: the implicit association test</article-title>
          .,
          <source>Journal of personality and social psychology 74</source>
          (
          <year>1998</year>
          )
          <fpage>1464</fpage>
          -
          <lpage>1480</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Purpura</surname>
          </string-name>
          , G. Silvello,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Susto</surname>
          </string-name>
          ,
          <article-title>Gender stereotype reinforcement: Measuring the gender bias conveyed by ranking algorithms</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>57</volume>
          (
          <year>2020</year>
          )
          <fpage>102377</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Otterbacher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <article-title>Competent men and warm women: Gender stereotypes and backlash in image search results</article-title>
          ,
          <source>in: Proc. of CHI</source>
          <year>2017</year>
          , ACM,
          <year>2017</year>
          , p.
          <fpage>6620</fpage>
          -
          <lpage>6631</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Matuszek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Munson</surname>
          </string-name>
          ,
          <article-title>Unequal representation and gender stereotypes in image search results for occupations</article-title>
          ,
          <source>in: Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, ACM</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>3819</fpage>
          -
          <lpage>3828</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Rekabsaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schedl</surname>
          </string-name>
          ,
          <article-title>Do neural ranking models intensify gender bias?</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR '20,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2020</year>
          , p.
          <fpage>2065</fpage>
          -
          <lpage>2068</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Harman</surname>
          </string-name>
          ,
          <article-title>The darpa tipster project</article-title>
          ,
          <source>SIGIR Forum 26</source>
          (
          <year>1992</year>
          )
          <fpage>26</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bolukbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.-W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Saligrama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Kalai</surname>
          </string-name>
          ,
          <article-title>Man is to computer programmer as woman is to homemaker? debiasing word embeddings</article-title>
          ,
          <source>in: Advances in neural information processing systems</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>4349</fpage>
          -
          <lpage>4357</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Bryson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <article-title>Semantics derived automatically from language corpora contain human-like biases</article-title>
          ,
          <source>Science</source>
          <volume>356</volume>
          (
          <year>2017</year>
          )
          <fpage>183</fpage>
          -
          <lpage>186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          ,
          <source>in: Advances in neural information processing systems</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Järvelin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kekäläinen</surname>
          </string-name>
          ,
          <article-title>Cumulated gain-based evaluation of ir techniques</article-title>
          ,
          <source>ACM Transactions on Information Systems</source>
          <volume>20</volume>
          (
          <year>2002</year>
          )
          <fpage>422</fpage>
          -
          <lpage>446</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Robertson</surname>
          </string-name>
          , U. Zaragoza,
          <article-title>The Probabilistic Relevance Framework: BM25 and Beyond, Foundations and Trends in Information Retrieval (FnTIR) 3 (</article-title>
          <year>2009</year>
          )
          <fpage>333</fpage>
          -
          <lpage>389</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>I.</given-names>
            <surname>Vulić</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-F. Moens</surname>
          </string-name>
          ,
          <article-title>Monolingual and cross-lingual information retrieval models based on (bilingual) word embeddings</article-title>
          ,
          <source>in: Proc. of SIGIR</source>
          <year>2015</year>
          , ACM,
          <year>2015</year>
          , p.
          <fpage>363</fpage>
          -
          <lpage>372</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>