<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Analysis for Interpreting Large Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elisabetta Rocchetti</string-name>
          <email>elisabetta.rocchetti@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alfio</string-name>
          <email>alfio.ferrara@unimi.it</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Università degli Studi di Milano, Department of Computer Science</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>32</volume>
      <fpage>23</fpage>
      <lpage>26</lpage>
      <abstract>
        <p>Being able to understand the inner workings of Large Language Models (LLMs) is crucial for ensuring safer development practices and fostering trust in their predictions, particularly in sensitive applications. Causal Mediation Analysis (CMA) is a causality framework which fits perfectly for this scenario, providing a mechanistic interpretation of the behaviour of LLM components and assessing a specific type of knowledge in the model (e.g. presence of gender bias). This study discusses the challenges and potential pathways in applying CMA to open LLMs' black boxes. Through three exemplary case studies from the literature, we show the unique insights CMA can provide. We elaborate on the inherent challenges and opportunities this approach presents. These challenges range from the influence of model architecture on prompt viability to the complexities of ensuring metric comparability across studies. Conversely, the opportunities lie in the dissection of LLMs' knowledge through the extraction of the specific domains of knowledge activated during processing. Our discussion aims to provide a comprehensive insight into CMA, focusing on essential aspects to equip researchers with the knowledge necessary for crafting efective CMA experiments tailored towards interpretability objectives.</p>
      </abstract>
      <kwd-group>
        <kwd>LLM</kwd>
        <kwd>interpretability</kwd>
        <kwd>causality</kwd>
        <kwd>causal mediation analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Large Language Models (LLMs) have gained a great amount of success and have become
ubiquitous in many research and application areas. Understanding their behaviour is of central interest
to correct them at inference-time [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and to guarantee safer development [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In the area of XAI,
mechanistic interpretability techniques involve deconstructing the computational processes
of a model into its elements, with the aim of uncovering, understanding, and confirming the
algorithms (referred to as circuits in some studies) that are executed by the model’s weights [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Among these techniques, Causal Mediation Analysis (CMA) provides a causal approach which
aims at extracting reliable cause-efect relations between inputs and outputs, contrarily to those
XAI techniques relying merely on simple and correlations. The architecture of a LLM can be
interpreted as a structural causal model: within this framework, CMA enables the isolation of
independent contributions from individual neural network components.
      </p>
      <p>In this study, we aim to elucidate the CMA technique and its application in probing the
inner mechanisms of LLMs. Through a series of case studies drawn from existing literature, we
highlight the potential benefits and opportunities that CMA ofers. Furthermore, we delve into
the primary challenges encountered when applying CMA, as well as the pressing issues that
must be addressed to enhance its robustness and flexibility. This paper is structured as follows:
Section 2 introduces the CMA formulation; Section 3 shows some of the works in the literature
applying CMA to three diferent case studies; Section 4 shows how to apply interventions for
CMA including an illustrative example; Section 5 discusses the limitations and issues of CMA,
alongside its potential and challenges; Section 6 concludes.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Causal Mediation Analysis</title>
      <p>
        Consider a causal model including three variables,  ,  and  , representing an intervention,
an outcome and a mediator respectively. In particular, the intervention  afects the outcome
variable  , and  is placed between these two and modifies some intermediate process between 
and  . How can we measure the separate efects of  and  on  ? Linear regression paradigms [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
rely on the “no interaction” property, thus they cannot work in nonlinear systems where editing
 could change the efect of  on  . Causal mediation analysis [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] aims at measuring the efects
of an intervention  on an outcome variable  when an intermediate variable  is standing
between the two, modifying some intermediate process between  and  . This method can
provide an answer to our question, since it removes these nonlinear barriers using causal
assumptions1. Let our system be the one depicted in Figure 1, where  =  1( 1), z =  2(,  2),
y =  3(, z,  3),  ,  ,  are discrete or continuous random variables,  1,  2,  3 are arbitrary
functions and  1,  2,  3 represent omitted factors which are assumed to be mutually independent
yet arbitrarily distributed [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This technique allows the separation of the total efect of  on 
in indirect efect and direct efects . The total efect (   ) measures the change in  produced by a
change in  , for example from  =  to  =  ′ [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We can express   at the population level
using the formula
      </p>
      <p>TE, ′ = ( | = 
′) − ( | = )
(1)
1However, one assumption is still made: the error terms must be mutually independent.</p>
      <p>
        The notion we introduce here about direct efects (DE) refers to what is technically called
”Natural Direct Efects ”2. As defined in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], DE is the expected change in  induced by changing
 from  to  ′ while keeping all mediating factors constant at whatever value they would have
obtained under  =  , before the transition from  to  ′. Estimating DE from the population
data is formalised by
      </p>
      <p>DE, ′( ) =
∑[( |
′, z) − ( |,</p>
      <p>
        z)] ( z|)


 = 
estimated by
where the condition probabilities use short-hand notations for  =  ,  =  ′, and  =
Contrarily to the DE, the indirect efect (IE) is defined as the expected change in
 while
keeping  constant and changing  to the value it would have attained has  been set to
′ (according to each individual) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This requires, indeed, a counterfactual representation
IE, ′( ) =
∑ ( |,
z)[ ( z| ′) −  ( z|)]
Equation 3 is a general formula for estimating the mediating efects, and can be also applied to
(2)
(3)
any nonlinear system.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Literature Review</title>
      <p>
        CMA has been recently applied in the field of XAI for LLM [
        <xref ref-type="bibr" rid="ref10 ref2 ref6 ref7 ref8 ref9">2, 6, 7, 8, 9, 10, 11, 12</xref>
        ]. Indeed,
the neural network architecture of a LLM can be viewed as a structural causal model. We
can view a subset of a language model’s internal components as an instance of the mediator
variable  . Suppose we select a specific neuron to be our z: then, z’s output is influenced by the
model’s input, and it afects the model output [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. To show the real efectiveness of CMA, we
cover three diferent case studies from the literature: gender bias detection [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7, 12</xref>
        ], syntactic
agreement [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and arithmetic reasoning [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <sec id="sec-4-1">
        <title>3.1. Gender bias detection</title>
        <p>
          The first attempt towards the causal mediation formula application is shown in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and extended
in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. These studies introduce a novel approach for probing both the structural dynamics and
predictive behaviors of LLMs, with a particular focus on uncovering and quantifying gender
bias. The methodology centers around inputting specifically crafted prompts, such as “
The
nurse said that [blank]”, into LLMs to observe the predictive preference between gendered
pronouns “he” and “she” filling the blank. This setup allows for an examination of bias: a model
demonstrating a consistent higher likelihood for “she” in contexts traditionally stereotyped
towards women is flagged as exhibiting gender bias. To quantify this bias, the authors define a
grammatical gender bias measure that compares the prediction probabilities of anti-stereotypical
and stereotypical pronouns. Through designed interventions—manipulating the input sentence
to replace profession nouns with their anti-stereotypical pronouns—the authors calculate TE,
2The term natural here refers to the fact that we want to observe the change in  after a change in  while holding
 at a constant value, and the level at which this constant value is set can vary based on the individual we are
considering.
        </p>
        <p>DE, and IE. These metrics illuminate the separate and combined influences of the intervention
and the mediator variable (a specific neuron or set of neurons within the LLM) on the model’s
output.</p>
        <p>This comprehensive methodology has yielded insightful findings: larger models are
disproportionately afected by gender bias, and the manifestation of bias significantly varies across
diferent datasets. Moreover, certain biases were found to align with crowdsourced gender
perceptions. Importantly, the study also pinpoints the localization of gender bias within the
model, identifying middle network layers and specific attention heads as primary contributors.
These findings not only enhance our understanding of gender bias within LLMs but also guide
targeted interventions for mitigating such biases, thereby paving the way for more equitable AI
systems [12].</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Syntactic agreement</title>
        <p>
          The study by [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] explores the application of CMA to probe models’ sensitivity to syntactic
agreement, assessing how diferent syntactic structures influence a model’s preference for
verb inflections. The evaluated structures range from simple agreements to complex scenarios
involving object relative clauses and distractors, aiming to understand the model’s grammatical
preferences. The intervention swap-number is introduced to challenge the model with
counterfactual prompts, altering the number feature of subjects to examine the model’s inflection choice
(e.g. “The friend (that) the lawyers *likes/like” becomes “The friend (that) the lawyer likes/*like”,
with the asterisk denoting the erroneous inflection). This approach helps identify if the model
favors the correct grammatical form, with expectations set for the model’s preference metrics
in response to these interventions.
        </p>
        <p>Their findings reveal nuanced insights into model behavior: contrary to previous gender bias
studies, model size does not linearly correlate with the magnitude of syntactic preference. The
presence of adverbial distractors increases total efects, suggesting improved accuracy, while
attractors decrease accuracy. The study also highlights the distribution of syntactic knowledge
across model layers and the impact of structural separations on subject-verb agreement, ofering
a comprehensive view of how LLMs process syntactic information.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Arithmetic reasoning</title>
        <p>
          In exploring arithmetic reasoning within LLMs, the study by [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] applies CMA to dissect the
internal mechanics of LLMs as they process mathematical concepts. The authors hypothesize a
network subset specialized in arithmetic reasoning, tested through task-specific prompts that
blend operands and operations into arithmetic problems of varying complexity. Modifications
for the study include altering operands and operations to gauge the model’s computational
accuracy and mediator contribution. This involves generating problems like ”How much is  1
plus  2? ” and assessing outcomes against counterfactual scenarios, quantified through an IE
formulation. Key activation sites identified include: the Multi Layer Perceptron (MLP) modules
at initial layers for operand tokens, intermediate attention blocks for sequence ends, and
laterlayer MLP modules for final token processing [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. This suggests attention mechanisms channel
necessary information for MLPs to execute computations. Further analysis on number retrieval
and factual knowledge, using randomized templates, indicates the last-token MLP’s broad role
in information processing, not strictly limited to arithmetic. This contrasts with early MLP
involvement in factual retrieval, highlighting arithmetic specificity in late MLP activations.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Applying the Causal Mediation Formula in LLMs</title>
      <p>
        Manipulating the internal representations of LLMs enables the generation of genuine
counterfactual outputs. This is achieved by transferring internal representations between model
executions that use both original and modified utterances. Here, we detail how this process is
used to compute total, direct, and indirect efects, leveraging the gender bias case study from [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Given the problem of bias detection, we need to engineer prompts so that we induce the model
to lean towards expressing its bias, if it does have any. For instance, the authors in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] feed a
LLM with prompts  like “The accountant said that [blank]”, where the profession “accountant”
is interpreted as stereotypically female, as result of a crowdsourced stereotypicality metric [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Then, the evaluation consists in testing which among the tokens “he” and “she” has the highest
probability of being predicted instead of the [blank] space. Given this example, if the model
consistently shows a higher likelihood for the stereotypical pronoun “she”   (she|) than the
anti-stereotypical pronoun “he”, then the LM is said to be biased ( are the model parameters)
y() =
  ( anti-stereotypical ∣ )
  ( stereotypical ∣ )
.
      </p>
      <p>
        (4)
If y() &lt; 1 , the prediction is stereotypical; if y() &gt; 1 , the prediction is anti-stereotypical;
and if y() = 1 , the prediction is unbiased [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We illustrate a set-gender-male neuron
intervention as described by [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. Under null intervention,  = The accountant said that. The
set-gender-male intervention exchanges the word “accountant” with its antistereotypical
counterpart, which is “man”. The two variants are processed by the same network and the
probabilities of the two candidates “she” and “he” are evaluated. Figure 2 depicts the procedure
for obtaining the candidates’ probabilities. TE can thus be computed as3,
Calculating DE and IE requires capturing intermediate representations from the mediator z,
potentially involving multiple MLPs or attention layers. In our example, z is an MLP. To compute
the DE, we need to take the “accountant ” representation resulting from z when it processes the
original sentence. Then, this representation replaces the one resulting from z when it processes
“man” in the alternate sentence. Figure 3 shows this procedure. DE is computed as
DE( set-gender, null ; y, ) =
yset-gender ,znull () ()
ynull ()
Concerning the IE computation, we need to extract the “man” representation resulting from z
when it processes the alternate sentence, and this then replaces the “accountant ” representation
from z when it processes the original sentence. Figure 4 shows this procedure. IE can be
computed as
IE( set-gender, null ; y, The accountant said that) =
ynull ,zset-gender () ()
ynull ()
− 1 = 0.5 / 0.25 = 1.2
0.2 0.12
Results obtained with this analysis can be then verified employing a LLM initialised with
random weights, and comparing results with the ones coming from the original execution.
      </p>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion</title>
      <p>As demonstrated in previous sections, implementing CMA is relatively straightforward.
However, there are critical points to address to design an efective CMA experiment. In this section,
3The probability values shown in these examples are not coming from a real experiment.
A
MLP
A
+ ..
+ ..
MLP</p>
      <p>MLP
A
A
+
+</p>
      <p>A
A
+
+
we discuss important limitations and issues to consider when applying this technique. In
particular, we cover challenges about interventions, prompts and metrics engineering.</p>
      <p>Intervention engineering-related challenges. Designing interventions in a prudent
manner is a key factor to experiment success. The aim here is to thinking which syntax could
trigger the desired LLM knowledge the most. For example, if gender bias is the object of
inspection, words having a strong stereotypical connotations are better suited for replacing
neutral expressions. One could also design diferent interventions to trigger the model at
diferent levels, and then compare results for these experiments. Another challenge is to
produce alternative sentences from which to extract alternate representation for interventions.
These sentences may vary syntactically from the originals, yet their semantics are required to
convey a concept that is diametrically opposed to that of the original sentences, contingent
upon the chosen intervention. For instance, let’s take the concept of “leadership” and explore
how we might intervene in a sentence to shift the perception from a traditional to a more
inclusive understanding, while ensuring grammatical correctness. If the original sentence was
“The successful leader commanded his team with firmness and ensured compliance through strict
policies.”, the intervention process would include multiple modifications, for example using
gender-neutral pronouns, a softer tone, and democratic policies. The intervened sentence could
be: “The successful leader guided their team with understanding and fostered collaboration through
lfexible policies. ”.</p>
      <p>Prompt engineering-related challenges. Model selection is pivotal in CMA prompt
generation, demanding tailored strategies based on the chosen model. The prompt “The nurse
said that [blank]” exemplifies how decoder-only and auto-regressive models handle candidate
probabilities diferently compared to masked models that employ a [MASK] token. For
non-endof-sentence evaluation prompts, like “[blank] dream is to become a doctor ”, modifications are
necessary to maintain evaluation consistency across models. Indeed, decoder-only model could
not be tested using this framework. A trivial solution to this issue is to manipulate the prompt so
that the candidates are placed last, but in this case we cannot assure consistency of magnitudes
for the computed causal efects without deeper investigations. Does the model behave diferently
when choosing diferent formulations of utterances? Do the estimated probabilities shift due to
the varying structures? Does the model produce more uncertain results after modification? This
highlights the intricacies of prompt engineering in CMA, requiring not only specific adaptations,
such as rephrasing to fit model requirements but also a deeper linguistic analysis to isolate the
intended efect from potential confounders. For instance, addressing issues like coreference and
complex sentence structures ensures the reliability of results by minimizing the influence of
unrelated variables. Relevant linguistic features to inspect prior to CMA experiments must be
selected according to what type of linguistic analysis has been found to be relevant for LLMs.</p>
      <p>
        Metrics engineering-related challenges. Diferent works in the literature use diverse
metrics to compute the desired efect, and these metrics are usually employed to perform
comparisons across models and datasets. However, these metrics lack of an absolute scale
suggesting how to interpret the diferent magnitudes in the same and across diferent case
studies [
        <xref ref-type="bibr" rid="ref2 ref7">2, 7</xref>
        ]. We argue that this limitation restricts analysis to merely ranking efects rather
than quantifying them relative to each other, complicating comparisons such as determining
whether BERT exhibits more bias than GPT2, or identifying which templates trigger the most
bias. There are many variables afecting the magnitude of the output probabilities employed in
the computation of metrics, including the presence of a more dificult computation involved in
the process (e.g. coreference resolution in Winograd-style datasets), or even relative position on
the candidate tokens in the sentence. It’s imperative that these metrics are formulated to ensure
comparability in scale and consistency across diferent experimental designs. For instance,
future research could focus on developing a normalized bias index designed to measure and
compare biases across models, taking into account factors such as coreference dificulty and
token positioning.
      </p>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusions</title>
      <p>We have shown which types of insights CMA can extract from Transformer-based LLMs through
three diferent exemplar case studies from the literature. Moreover, we detail the application of
this analysis to equip readers with a practical understanding of the technique, thereby enabling
them to more efectively engage with both the challenges and opportunities CMA presents.
Challenges include the impact of model architecture on prompt viability and the intricacies of
ensuring metric comparability across studies. On the opportunity side, it includes the ability to
dissect the knowledge within LLMs, ofering insights into the knowledge domains activated
during processing. All the case studies presented in this work share something, which we
argue to be rather important: they all sought to uncover which knowledge a Transformer has
learned during its training process. CMA gives us the capability to examine the activation
within a LLM’s neural architecture, thereby discerning the specific domains of knowledge
engaged during processing. For example, it could be feasible to delineate the global and human
values encapsulated within the documents constituting the training dataset. This concept is
particularly compelling as it afords an objective representation of contemporary societal values,
the educational paradigms imparted to a generation, or the extraction of characteristic human
values from historical contexts. The nature of the insights gleaned is inherently dependent on
the composition of the training data supplied to the model.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This work was supported in part by project SERICS (PE00000014) under the NRRP MUR program
funded by the EU - NGEU. Views and opinions expressed are however those of the authors
only and do not necessarily reflect those of the European Union or the Italian MUR. Neither the
European Union nor the Italian MUR can be held responsible for them.
Wild: A Circuit for Indirect Object Identification in GPT-2 small, in: NeurIPS ML Safety
Workshop, 2022.
[11] K. Meng, D. Bau, A. Andonian, Y. Belinkov, Locating and Editing Factual Associations in</p>
      <p>GPT, 2023. doi:10.48550/arXiv.2202.05262. arXiv:2202.05262.
[12] Y. Da, M. N. Bossa, A. D. Berenguer, H. Sahli, Reducing Bias in Sentiment Analysis Models
Through Causal Mediation Analysis and Targeted Counterfactual Training, IEEE Access
(2024) 1–1. doi:10.1109/ACCESS.2024.3353056.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Viégas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Pfister</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wattenberg</surname>
          </string-name>
          ,
          <article-title>Inference-time intervention: Eliciting truthful answers from a language model</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>36</volume>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Stolfo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Belinkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sachan</surname>
          </string-name>
          ,
          <article-title>A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis</article-title>
          , in: H.
          <string-name>
            <surname>Bouamor</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Pino</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          Bali (Eds.),
          <source>Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Singapore,
          <year>2023</year>
          , pp.
          <fpage>7035</fpage>
          -
          <lpage>7052</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .emnlp- main.435.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Räuker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Casper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hadfield-Menell</surname>
          </string-name>
          ,
          <article-title>Toward transparent ai: A survey on interpreting the inner structures of deep neural networks</article-title>
          ,
          <source>in: 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML)</source>
          , IEEE,
          <year>2023</year>
          , pp.
          <fpage>464</fpage>
          -
          <lpage>483</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Baron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Kenny</surname>
          </string-name>
          ,
          <article-title>The moderator-mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations</article-title>
          .,
          <source>Journal of personality and social psychology 51</source>
          (
          <year>1986</year>
          )
          <fpage>1173</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          ,
          <article-title>The Causal Mediation Formula-A Guide to the Assessment of Pathways and Mechanisms</article-title>
          ,
          <source>Prevention Science</source>
          <volume>13</volume>
          (
          <year>2012</year>
          )
          <fpage>426</fpage>
          -
          <lpage>436</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11121- 011- 0270- 1.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Vig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gehrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Belinkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nevo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shieber</surname>
          </string-name>
          ,
          <article-title>Investigating Gender Bias in Language Models Using Causal Mediation Analysis</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2020</year>
          , pp.
          <fpage>12388</fpage>
          -
          <lpage>12401</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Vig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gehrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Belinkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nevo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sakenis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shieber</surname>
          </string-name>
          ,
          <source>Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias</source>
          ,
          <year>2020</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.
          <year>2004</year>
          .
          <volume>12265</volume>
          . arXiv:
          <year>2004</year>
          .12265.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Finlayson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gehrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shieber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Linzen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Belinkov</surname>
          </string-name>
          ,
          <article-title>Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th</article-title>
          <source>International Joint Conference on Natural Language Processing</source>
          (Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>1828</fpage>
          -
          <lpage>1843</lpage>
          . doi:
          <volume>10</volume>
          .18653/ v1/
          <year>2021</year>
          .acl- long.144.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Geiger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Icard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Potts</surname>
          </string-name>
          ,
          <source>Causal Abstractions of Neural Networks, in: Advances in Neural Information Processing Systems</source>
          , volume
          <volume>34</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2021</year>
          , pp.
          <fpage>9574</fpage>
          -
          <lpage>9586</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Variengien</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Conmy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shlegeris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Steinhardt</surname>
          </string-name>
          , Interpretability in the
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>