<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Trust in Fair Algorithms: Pilot Experiment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mattia Cerrato</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marius Köppel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kiara Stempel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alesia Vallenas Coronel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Niklas Witzig</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ETH Zürich (IPA)</institution>
          ,
          <addr-line>Otto-Stern-Weg 5, 8093 Zürich CH</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Johannes Gutenberg-University Mainz</institution>
          ,
          <addr-line>Saarstraße 21, 55122 Mainz DE</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We study human reliance on (un)fair algorithmic recommendations. We train two separate models to predict the employment status of (anonymized) persons in the US-Census 2018. The two models difer in their fairness, with their overall accuracy being the same for (difering by) binary 1 gender ("overall accuracy equality") [2]. We use the predictions by the two models in a pilot experiment, in which participants performed the same prediction task while receiving assistance from either the fair or unfair model. We find that people rely more on the predictions by the fair algorithm with unknown gender, yet stronger on unfair models once the gender is revealed. The present data remains limited in size, but we aim to identify these efects and potential sources of individual treatment heterogeneity using a much larger experimental sample in the future.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Algorithmic Fairness</kwd>
        <kwd>Trust</kwd>
        <kwd>User Study</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        1We fully acknowledge that this dichotomy is far from ideal (see also the discussions in Pinney et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]), but the
data availability does not allow the use of non-binary gender attributes so far.
      </p>
      <p>EWAF’24: European Workshop on Algorithmic Fairness, July 01–03, 2024, Mainz, Germany
* Corresponding author.
insignificant) evidence that prior beliefs, i.e., the general and potentially biased expectation about
the employment status of men and women are a potential source of treatment heterogeneity.</p>
    </sec>
    <sec id="sec-2">
      <title>2. (Un-)fairness in Predicting Employment Status</title>
      <p>The main goal is to deploy two models with difering fairness levels, measured by accuracy,
to ofer predictions to participants. Both models are trained on the same 2018 U.S. Census
data, predicting employment status using variables like age, education, disability, marital status,
and other sociodemographic factors. We exclude the sensitive attribute, gender, from the
model’s input. We use BinaryMI [? ] for both models, which aims to make a fair prediction
by learning a representation of the data that is invariant to the sensitive attribute, by using a
stochastically quantized binary neural layer. All parameters for training are set to the same
values for both models, except for  , which is the tradeof parameter in the loss function
weighting the importance of representation invariance and accuracy on the target variable, i.e.,
the employment status. For the fair model, we set  to 0.05, whereas for the unfair one, we set
it to 0. Both models are trained for 50 epochs. Before model training, the dataset is subsampled
to retain only 10% of the instances of employed women, reducing their representation in the
learning process. While the fair model reaches an accuracy of 79.37% for male and 79.18% for
female, the unfair version reaches an accuracy of 81.98% for male, but only 79.24% for female. In
the experiment, we select 10 random profiles from the validation dataset, ensuring a balance of
gender and employment status, and representing two profiles each from the prediction quintiles
of the fair, unfair models, and their diference.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Online Experiment</title>
      <p>
        The experiment, centering on a prediction task and supplemented by auxiliary measures, was
conducted on Prolific (www.prolific.com) in December 2023 with 103 participants. It lasted an
average of 27.23 minutes, with 77% of participants receiving a 2-pound bonus. Main Task: For
participants in the online experiment, the main task is to give an accurate probabilistic statement
if a given (anonymized) person is employed or not based on shown personal characteristics.
The statement ranges from 0 to 100 and the closer the statement is to the true employment
status (either 0 or 1 (=100)), the higher the chances that a participant obtains a bonus prize. We
apply the randomized quadratic scoring rule [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which is proper even for risk-averse subjects
and incentivizes participants to report their beliefs truthfully. The main task consists of three
phases, which entail 10 statements each (s.t. participants give 30 statements in total). We
randomly draw one statement to determine participants’ payofs. Panels (a) - (c) of Figure
1 show screenshots of the three phases. In phase 1 of the main task, participants obtain
information about six personal characteristics of a given anonymized person, i.e., about their
age, education level, marital status, whether they have a disability, if they moved in the previous
year, and if they have single or multiple ancestries. Information about the gender of the person
remains undisclosed. Based on this information, participants need to form and enter a belief
about the employment status of the person. In phase 2, participants receive the prediction
by either the fair or unfair model, which is randomly determined at the beginning of the
(a) Maintask Phase 1
(b) Maintask Phase 2 (unfair)
(c) Maintask Phase 3 (unfair)
Weight on Advice Phases 2−1 and 3−2
3-2, separately for Fair and unfair. (e) Intercept and remaining interactions remain omitted.
experiment (and remains constant throughout, "between-subject" treatment). In addition, they
are reminded of the (un)fairness by the global accuracy plot next to the personal characteristics.
Participants now may update their belief estimates. Phases 1 and 2 are carried out sequentially
for each anonymized person, s.t. participants first enter their initial estimate and then may
revise their estimate. After phases 1 and 2 (for each of the 10 profiles), they are introduced to
phase 3, which discloses the gender of the profile and participants may revise their (possibly
updated) estimate a final time. This task allows us to construct the central measure of trust
on algorithmic predictions, namely the weight of advice (WOA). This measure stems from
the advice-taking literature and we use the censored measure by Greiner et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]: WOA =
, 1 . This measure is 0 if participants — in their
min
︁( max 0, || umpoddaetlepdreesdtiicmtiaotne −− initial estimate | )︁
︁(
initial estimate |
︁)
updated estimate — ignore the model prediction, and 1 if they react as much (or more strongly)
as suggested by the model prediction. Crucially, we estimate the WOA separately for phases
2 and 1, i.e., based on the first update with a model estimate, and for phases 3 and 2, the final
estimate once the profiles’ gender was revealed.
      </p>
      <p>
        Additional tasks and measures: We carry
out several additional tasks and surveys. We gather data on prior beliefs about the employment
status of men and women, a survey with attitudes towards the statistical model by Zhou et al.
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and a general trust in various institutions survey [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. We use this data to investigate
heterogeneous treatment efects, i.e., if our chosen treatment variation invoked systematically
diferent behavior across participants depending on these characteristics. For brevity, we focus
on the role of prior beliefs in the later analyses.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>Main Treatment Diferences: Panel (d) in Figure 1 plots the average WOA for the two groups.
On average, the WOA is 0.48 in the unfair group and 0.52 in the fair group. This diference is
statistically significant on the 10% level (  = 0.061) in a two-sided t-test. Participants showed a
stronger reaction to predictions from the fair model than those from the unfair model. Despite
the small diference, it’s significant considering the low "treatment intensity" between the
models, making a 4 percentage point diference noteworthy. Contrary, the mean WOA values
between phases 3 and 2 (i.e., once the gender of the candidate has been revealed) show a lower
reliance on the predictions by the fair model (panel (d)). On average, the WOA on the fair model
is 0.41, yet 0.45 on the unfair model ( = 0.052). This suggests that participants react more
strongly to fair predictions if the sensitive attribute remains undisclosed but more strongly to
unfair predictions once the sensitive attribute is known. Regressions and Heterogeneous
Treatment Efects: The Table in panel (e) in Figure 1 regresses the (final) revised belief
statement on the initial (revised) statement, the model predictions from the main task between
phase 2 and 1 (3 and 2) and a dummy for the treatment group. Columns 1 and 3 include data
from phases 2-1, whereas column 2 and 4 from phases 3-2. Focussing on the first two columns,
we observe a stronger reliance on the participant’s own initial (revised) estimates compared
to the model predictions, while this asymmetry is much more pronounced between phases
3 and 2. In line with the evidence from panel (d), the interaction efect of the fair treatment
with the model prediction has a positive (negative) coeficient, implying that participants in
the fair treatment rely more (less) strongly on the model predictions during phases 2-1 (3-2).
The interaction efects were not statistically significant, likely due to the low power of the pilot
experiment. The analysis includes a treatment heterogeneity segment, examining an interaction
between fair treatment, model predictions, and "Belief Delta" – a measure of the diference in
participants guesses regarding the employment rates of men versus women. The coeficient
of this interaction is positive (negative) for phases 2-1 (3-2), implying that the more "biased" a
person the stronger (weaker) their reliance on the fair predictions. Again, these interactions are,
however statistically insignificant.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Next Steps</title>
      <p>The key finding of this study is that participants in our pilot experiment reacted more strongly to
the predictions by a fair model compared to an unfair one in the presence of unknown sensitive
attributes, yet less once the sensitive attribute was revealed. This novel insight indicates that
the efectiveness of fair models depends on the decision-makers using them. We recognize
the limitations of our pilot study, notably its low power, leading to provisional conclusions.
Methodologically, we didn’t consider the statistically optimal posterior belief (assuming Bayesian
participants), potentially explaining some treatment diferences. We plan to overcome these
limitations with a larger online experiment and by theoretically modeling task behavior.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Pinney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Raj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Ekstrand</surname>
          </string-name>
          ,
          <article-title>Much ado about gender: Current practices and future recommendations for appropriate gender-aware information access</article-title>
          ,
          <source>in: Proceedings of the 2023 Conference on Human Information Interaction and Retrieval</source>
          ,
          <source>CHIIR '23</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2023</year>
          . URL: http://dx.doi.org/10.1145/3576840.3578316. doi:
          <volume>10</volume>
          .1145/3576840.3578316.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rubin</surname>
          </string-name>
          ,
          <article-title>Fairness definitions explained</article-title>
          ,
          <source>in: Proceedings of the International Workshop on Software Fairness, ICSE '18</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2018</year>
          . URL: http://dx.doi.org/10.1145/ 3194770.3194776. doi:
          <volume>10</volume>
          .1145/3194770.3194776.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hossain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Okui</surname>
          </string-name>
          ,
          <article-title>The binarized scoring rule</article-title>
          ,
          <source>The Review of Economic Studies</source>
          <volume>80</volume>
          (
          <year>2013</year>
          )
          <fpage>984</fpage>
          -
          <lpage>1001</lpage>
          . URL: http://dx.doi.org/10.1093/restud/rdt006. doi:
          <volume>10</volume>
          .1093/restud/rdt006.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Greiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Grünwald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lindner</surname>
          </string-name>
          , G. Lintner,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiernsperger</surname>
          </string-name>
          , Incentives, Framing, and
          <article-title>Reliance on Algorithmic Advice: An Experimental Study</article-title>
          ,
          <source>WorkingPaper</source>
          <volume>01</volume>
          /
          <year>2024</year>
          , WU Vienna University of Economics and Business,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .57938/ 137467e7-9d47
          <string-name>
            <surname>-</surname>
          </string-name>
          4a6c
          <string-name>
            <surname>-</surname>
          </string-name>
          87dc-673ff4f58009.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mittal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Understanding relations between perception of fairness and trust in algorithmic decision making (</article-title>
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>[6] Oecd, OECD guidelines on measuring trust, Organization for Economic Co-operation and Development (OECD</article-title>
          ), Paris Cedex, France,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>