<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unveiling the blindspots: Examining availability and usage of protected attributes in fairness datasets⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Simson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Fabris</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Kern</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LMU Munich</institution>
          ,
          <addr-line>Ludwigstr. 33, 80809 München</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Max Planck Institute for Security and Privacy</institution>
          ,
          <addr-line>Universitätsstraße 140, 44799 Bochum</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Munich Center for Machine Learning (MCML)</institution>
          ,
          <addr-line>Oettingenstraße 67, 80538 München</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This work examines the representation of protected attributes across tabular datasets used in algorithmic fairness research. Drawing from international human rights and anti-discrimination laws, we compile a set of protected attributes and investigate both their availability and usage in the literature. Our analysis reveals a significant underrepresentation of certain attributes in datasets that is exacerbated by a strong focus on race and sex in dataset usage. We identify a geographical bias towards the Global North, particularly North America, potentially limiting the applicability of fairness detection and mitigation strategies in less-represented regions. The study exposes critical blindspots in fairness research, highlighting the need for a more inclusive and representative approach to data collection and usage in the field. We propose a shift away from a narrow focus on a small number of datasets and advocate for initiatives aimed at sourcing more diverse and representative data.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;critical data studies</kwd>
        <kwd>dataset usage</kwd>
        <kwd>protected groups</kwd>
        <kwd>generalization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Algorithmic fairness has become a significant area of research in recent years, with a growing
body of work aimed at addressing bias and discrimination in machine learning systems.
Identifying and mitigating harmful practices against vulnerable individuals and groups in prediction
algorithms lies at the core of this field and to study these issues adequate and nuanced data
sources are needed.</p>
      <p>In this work, we examine datasets and how they are used within the fairness literature. We
present an overview of attributes which are protected by anti-discrimination legislation across
multiple continents and study their availability in datasets and usage in fairness research. We
identify issues regarding the diversity of protected attributes represented in datasets and their
geographic representativeness, highlighting how populations are neglected in the literature.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <p>
        For this work, we collected and manually annotated usage of tabular datasets in fair classification
tasks. We built on top of a comprehensive survey of fairness datasets by Fabris et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
leveraging the same inclusion criteria. We focus on tabular datasets used for fair classification in
this work, due to their important role in the fairness literature [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. We study the use of tabular
datasets ( = 36) across 142 articles. Since datasets appear in multiple publications and most
publications use multiple datasets, the total number of dataset and publication combinations
examined was  = 280, with  = 233 instances of suficient information to reconstruct (or
reasonably guess) protected attribute usage.
      </p>
      <p>
        To define protected attributes, we draw from domain-specific legislation and human rights
law. We define as protected all attributes which are explicitly mentioned as prohibited drivers of
discrimination and inequality. For example, Article 21 of the Charter of Fundamental Rights of
the European Union states “Any discrimination based on any ground such as sex, race, colour,
ethnic or social origin, genetic features, language, religion or belief, political or any other
opinion, membership of a national minority, property, birth, disability, age or sexual orientation
shall be prohibited” [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        We try to adress the Global North and especially U.S. focus in AI ethics and fairness research
[
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ] by including works from the European Union (Charter of Fundamental Rights of the
European Union [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], EU legislation on fair hiring [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]), from other continents (African Charter on
Human and Peoples’ Rights [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the Arab Charter on Human Rights [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the ASEAN Declaration
of Human Rights [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) and global works (Universal Declaration of Human Rights [12]) besides
works from North America (the American Declaration of the Rights and Duties of Man [13],
US fair lending legislation [14]). However, we acknowledge that we fail to mitigate the Global
North bias completely, given the strong presence of said regions in research.
      </p>
      <p>Drawing from this literature, we provide a shallow categorization of protected attributes,
identifying seven main categories (Table 1). It is worth noting that this is not a complete
categorization of all protected attributes around the globe and across sectors. Rather, our
categorization aims to guide an inclusive discussion of algorithmic fairness research through
the lens of protected attributes.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>
        The geographical provenance of datasets used in the examined literature is clearly skewed.
Among datasets from a single continent, 21 come from North America, 5 from Europe, 2 from
Asia, and 2 from South America, confirming a Global North and especially North American
dominance in AI ethics research [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ]. Since fairness is highly contextual, there is a risk
that the fairness detection and mitigation strategies built by this research community will not
transfer and, therefore, underserve neglected geographical areas [15].
      </p>
      <p>We further notice a highly uneven distribution of both the availability and usage of protected
attributes. The left bar chart in Figure 1 depicts protected attributes available in fairness datasets
and the right chart their usage in the examined literature. There is a particular focus in both
availability (n=17) and usage (n=167) on race as a protected attribute. On the other hand,
attributes about religion, belief and opinion are entirely missing on both sides. Information on
disability and health conditions is also infrequently available ( = 3) and never used in the
surveyed literature. Socioeconomic status descriptors are more commonly available yet often
neglected. This threatens the applicability of research findings across contexts, as information
on race for example is hardly available in EU data [16].1</p>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
      <p>We unveil blindspots in fairness research, demonstrating a neglect of vulnerable subpopulations
in the literature. We will further present additional results from our data collection at the
conference, indicating other troubling practices in the field, such as a lack of suficient reporting
1There was also a small number of protected attributes used in the literature but not referenced in legislation, such
as employment status, alcohol consumption, neighborhood, body-mass index, and profession.</p>
      <p>n=14
n=17
n=8
n=8
n=8
n=8
n=7
of dataset usage impacting reproducibility and potentially harmful practices in the processing
of protected attributes leading to a neglect of minorities.</p>
      <p>While valid reasons exist against the collection of protected data [17], motivating e.g. the line
of work on fairness under unawareness [14, 18], we believe they are not suficient to explain
the observed lack in usage of particular attributes. We observe a clear trend towards certain
protected attributes being more readily available in datasets which, however, is amplified by
a strong tendency of papers to (1) repeatedly focus on the same small number of datasets
and (2) especially rely on race and sex as protected attributes. It is worth noting that this
trend extends to fairness research more broadly, including qualitative studies. These practices
also have a tendency to self-reinforce, increasing the likelihood of future research to conform.
Recent articles published at fairness conferences, such as FAccT (the ACM Conference on
Fairness, Accountability, and Transparency) and AIES (the AAAI/ACM Conference on Artificial
Intelligence, Ethics and Society), for example, mention race and gender by an order of magnitude
more frequently than religion, disability, socioeconomic status, and sexual orientation [19].</p>
      <p>
        We argue for a move towards a research roadmap to tackle these issues within the complex
social, legal and technical landscape they reside in (as advocated, for example, in Guo et al.
[20]). In particular, we propose a move away from focusing exclusively on a small number
of datasets[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], such as Adult, German Credit or COMPAS. Instead we suggest an increased
focus on using a diverse set of datasets and sourcing more representative data to fill in the gaps
of available datasets. We call for dedicated initiatives, including for example data donation
campaigns and citizen science initiatives, capable of filling this gap and responsibly handling
the collected data. We refer readers to the full paper [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for a more nuanced discussion. A list of
datasets and their protected attributes, as well as further analyses are available on Github.
[12] United Nations, Universal declaration of human rights, 1948. https://www.un.org/en/
about-us/universal-declaration-of-human-rights.
[13] Organization of American States, American declaration of the rights
and duties of man, 1948. https://www.oas.org/en/iachr/mandate/Basics/
american-declaration-rights-duties-of-man.pdf.
[14] J. Chen, N. Kallus, X. Mao, G. Svacha, M. Udell, Fairness under unawareness: Assessing
disparity when protected class is unobserved, in: danah boyd, J. H. Morgenstern (Eds.),
Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* 2019,
Atlanta, GA, USA, January 29-31, 2019, ACM, 2019, pp. 339–348. URL: https://doi.org/10.
1145/3287560.3287594. doi:10.1145/3287560.3287594.
[15] N. Sambasivan, E. Arnesen, B. Hutchinson, T. Doshi, V. Prabhakaran, Re-imagining
algorithmic fairness in india and beyond, in: M. C. Elish, W. Isaac, R. S. Zemel (Eds.), FAccT
’21: 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event /
Toronto, Canada, March 3-10, 2021, ACM, 2021, pp. 315–328. URL: https://doi.org/10.1145/
3442188.3445896. doi:10.1145/3442188.3445896.
[16] S. Jaime, C. Kern, Ethnic classifications in algorithmic fairness: Concepts, measures and
implications in practice, in: The 2024 ACM Conference on Fairness, Accountability, and
Transparency, FAccT ’24, Association for Computing Machinery, New York, NY, USA,
2024, p. 237–253. URL: https://doi.org/10.1145/3630106.3658902. doi:10.1145/3630106.
3658902.
[17] M. Andrus, E. Spitzer, J. Brown, A. Xiang, What we can’t measure, we can’t understand:
Challenges to demographic data procurement in the pursuit of fairness, in: M. C. Elish,
W. Isaac, R. S. Zemel (Eds.), FAccT ’21: 2021 ACM Conference on Fairness, Accountability,
and Transparency, Virtual Event / Toronto, Canada, March 3-10, 2021, ACM, 2021, pp. 249–
260. URL: https://doi.org/10.1145/3442188.3445888. doi:10.1145/3442188.3445888.
[18] A. Fabris, A. Esuli, A. Moreo, F. Sebastiani, Measuring fairness under unawareness of
sensitive attributes: A quantification-based approach, J. Artif. Intell. Res. 76 (2023) 1117–
1180. URL: https://doi.org/10.1613/jair.1.14033. doi:10.1613/JAIR.1.14033.
[19] A. Birhane, E. Ruane, T. Laurent, M. S. Brown, J. Flowers, A. Ventresque, C. L. Dancy,
The forgotten margins of AI ethics, in: FAccT ’22: 2022 ACM Conference on Fairness,
Accountability, and Transparency, Seoul, Republic of Korea, June 21 - 24, 2022, ACM,
2022, pp. 948–958. URL: https://doi.org/10.1145/3531146.3533157. doi:10.1145/3531146.
3533157.
[20] A. Guo, E. Kamar, J. W. Vaughan, H. M. Wallach, M. R. Morris, Toward fairness in AI for
people with disabilities sbg@a research roadmap, ACM SIGACCESS Access. Comput. 125
(2020) 2. URL: https://doi.org/10.1145/3386296.3386298. doi:10.1145/3386296.3386298.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Simson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kern</surname>
          </string-name>
          ,
          <article-title>Lazy data practices harm fairness research</article-title>
          , in: FAccT '24: 2021 ACM Conference on Fairness, Accountability, and
          <string-name>
            <surname>Transparency</surname>
          </string-name>
          , Rio de Janeiro, Brasil, June 3-6,
          <year>2024</year>
          , ACM,
          <year>2024</year>
          . URL: https://doi.org/10.1145/3630106.3658931. doi:
          <volume>10</volume>
          . 1145/3630106.3658931.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Messina</surname>
          </string-name>
          , G. Silvello,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Susto</surname>
          </string-name>
          ,
          <article-title>Algorithmic fairness datasets: the story so far</article-title>
          ,
          <source>Data Min. Knowl. Discov</source>
          .
          <volume>36</volume>
          (
          <year>2022</year>
          )
          <fpage>2074</fpage>
          -
          <lpage>2152</lpage>
          . URL: https://doi.org/10.1007/ s10618-022-00854-z. doi:
          <volume>10</volume>
          .1007/S10618-022-00854-Z.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mehrabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Morstatter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Saxena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Galstyan</surname>
          </string-name>
          ,
          <article-title>A survey on bias and fairness in machine learning</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>54</volume>
          (
          <year>2022</year>
          )
          <volume>115</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>115</lpage>
          :
          <fpage>35</fpage>
          . URL: https: //doi.org/10.1145/3457607. doi:
          <volume>10</volume>
          .1145/3457607.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Euopeam</given-names>
            <surname>Union</surname>
          </string-name>
          ,
          <article-title>Charter of fundamental rights of the european</article-title>
          union c-
          <volume>364</volume>
          /01,
          <year>2000</year>
          . https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32000X1218%
          <fpage>2801</fpage>
          %
          <fpage>29</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C. T.</given-names>
            <surname>Okolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Dell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vashistha</surname>
          </string-name>
          ,
          <article-title>Making ai explainable in the global south: A systematic review</article-title>
          ,
          <source>in: ACM SIGCAS/SIGCHI Conf. on Computing and Sustainable Societies (COMPASS)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>439</fpage>
          -
          <lpage>452</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Roche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wall</surname>
          </string-name>
          ,
          <source>Artificial intelligence ethics: An inclusive global discourse?</source>
          ,
          <source>arXiv preprint arXiv:2108.09959</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Septiandri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Constantinides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tahaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Quercia</surname>
          </string-name>
          , WEIRD faccts:
          <article-title>How western</article-title>
          , educated, industrialized, rich, and democratic is facct?,
          <source>in: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency</source>
          ,
          <source>FAccT</source>
          <year>2023</year>
          , Chicago, IL, USA, June 12-15,
          <year>2023</year>
          , ACM,
          <year>2023</year>
          , pp.
          <fpage>160</fpage>
          -
          <lpage>171</lpage>
          . URL: https://doi.org/10.1145/3593013.3593985. doi:
          <volume>10</volume>
          .1145/3593013.3593985.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Baranowska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Dennis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Graus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hacker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Saldivar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. Z.</given-names>
            <surname>Borgesius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Biega</surname>
          </string-name>
          ,
          <article-title>Fairness and bias in algorithmic hiring: a multidisciplinary survey (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>[9] Organisation of African Unity, African charter on human and peoples</article-title>
          ' rights,
          <year>1981</year>
          . https://au.int/sites/default/files/treaties/36390-treaty-0011_
          <article-title>-_african_charter_on_ human_and_peoples_rights_e</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>Council of the League of Arab States</article-title>
          ,
          <source>Arab charter on human rights</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <article-title>Association of Southeast Asian Nations, Asean declaration of human rights</article-title>
          ,
          <year>2012</year>
          . https: //asean.org
          <article-title>/asean-human-rights-declaration/.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>