<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Type-II Errors of Independence Tests Can Lead to Arbitrarily Large Errors in Estimated Causal E ects: An Illustrative Example</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicholas Cornia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joris M. Mooij</string-name>
          <email>j.m.mooijg@uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Informatics Institute University of Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Estimating the strength of causal e ects from
observational data is a common problem in
scienti c research. A popular approach is
based on exploiting observed conditional
independences between variables. It is
wellknown that this approach relies on the
assumption of faithfulness. In our opinion, a
more important practical limitation of this
approach is that it relies on the ability to
distinguish independences from (arbitrarily
weak) dependences. We present a simple
analysis, based on purely algebraic and
geometrical arguments, of how the estimation
of the causal e ect strength, based on
conditional independence tests and background
knowledge, can have an arbitrarily large
error due to the uncontrollable type II error of
a single conditional independence test. The
scenario we are studying here is related to
the LCD algorithm by Cooper [1] and to the
instrumental variable setting that is popular
in epidemiology and econometry. It is one
of the simplest settings in which causal
discovery and prediction methods based on
conditional independences arrive at non-trivial
conclusions, yet for which the lack of uniform
consistency can result in arbitrarily large
prediction errors.</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>Inferring causation from observational data is a
common problem in several elds, such as biology and
economics. To deal with the presence of unmeasured
confounders of observed random variables the so-called
instrumental variable technique [2] has found
applications in genetics [3], epidemiology [4, 5] and economics
[6]. Given two observable random variables possibly
in uenced by a hidden confounder, an instrumental
variable is a third observed variable which is assumed
to be independent of the confounder. In practice it
is di cult to decide whether the instrumental
variable de nition is satis ed, and the method has aroused
some skepticism [7]. In this paper, we study a
setting that is similar in spirit to the instrumental
variable model, but where all conditional independence
assumptions are directly testable on the observed data.</p>
      <sec id="sec-2-1">
        <title>A similar scenario was rst studied by Cooper [1] and independently rediscovered in the context of genome biology by Chen et al. [8].</title>
      </sec>
      <sec id="sec-2-2">
        <title>An important assumption in causal discovery meth</title>
        <p>ods based on conditional independences is
faithfulness, which means that the observed joint distribution
does not contain any additional (conditional)
independences beyond those induced by the causal structure.
Usually, faithfulness is justi ed by the assumption that
unfaithful distributions are a set of Lebesgue measure
zero in the set of the model parameters. By showing
that one can create a sequence of faithful distributions
which converges to an unfaithful one, Robins et al.
proved the lack of uniform consistency of causal
discovery algorithms [9]. Zhang and Spirtes [10] then
introduced the \Strong Faithfulness" assumption to recover
the uniform consistency of causal discovery. Using
geometric and combinatorial arguments, Uhler et al. [11]
addressed the question of how restrictive the Strong
Faithfulness assumption is in terms of the volume of
distributions that do not satisfy this assumption. Even
for a modest number of nodes and for sparse graphs,
the \not strongly faithful" regions can be surprisingly
large, and Uhler et al. argue that this result should
discourage the use of large scale causal algorithms based
on conditional independence tests, such as the PC and</p>
      </sec>
      <sec id="sec-2-3">
        <title>FCI algorithms [12].</title>
        <p>In this work, we analyse in the context of the LCD
setting how an error in a single conditional
independence test may already lead to arbitrarily large
errors in predicted causal e ect strengths, even when
the faithfulness assumption is not violated. Our
results may not be surprising for those familiar with the
work of [9], but we believe that the analysis we present
here may be easier to understand to those without a
background in statistics, as we separate statistical
issues (the possibility of type II errors in the conditional
independence test from a nite sample) from a rather
straightforward analysis of the problem in the
population setting. We use an algebraic approach, showing
how causal prediction may lead to wrong predictions
already in the simple context of linear structural
equation models with a multivariate Gaussian distribution.
In Section 1, we begin with a brief description of the
problem setting in a formal way, giving the de nitions
of the causal e ect, instrumental variable, LCD
algorithm and the toy model we present. We consider
three observed random variables (X1; X2; X3), which
is the minimal number such that a non-trivial
conditional independence test can be obtained. In Section 2,
we show how an (arbitrarily weak) conditional
dependence that goes undetected can in uence our
estimation of the causal e ect of X2 on X3 from the observed
covariance matrix, when a confounder between X2 and</p>
      </sec>
      <sec id="sec-2-4">
        <title>X3 is almost o -set by a direct e ect from X1 to X3.</title>
      </sec>
      <sec id="sec-2-5">
        <title>In fact, we show that this phenomenon can lead to an</title>
        <p>arbitrarily large error in the estimated causal e ect as
the noise variance of X2 approaches zero. We nish
with conclusions in Section 3.
1
1.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Problem setting</title>
      <sec id="sec-3-1">
        <title>LCD algorithm</title>
        <p>The model we are interested in arises from the work
of Cooper [1], who proposed the \LCD" algorithm for
causal discovery in observational databases and the
more recent paper of Chen et al.[8], who proposed the
\Trigger" algorithm to infer transcriptional regulatory
networks among genes. Throughout this section we
will assume:</p>
        <sec id="sec-3-1-1">
          <title>Acyclicity;</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>No Selection Bias.</title>
          <p>De nition 1.1. (LCD setting) Given three
random variables X1; X2; X3 such that the following
statistical properties and prior assumptions are satis ed:
Statistical dependences:</p>
          <p>X1 6?? X2
X2 6?? X3</p>
          <p>X1 ?? X3jX2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Prior assumptions:</title>
        <p>An(X1) \ fX2; X3g = ;</p>
        <sec id="sec-3-2-1">
          <title>Faithfulness</title>
          <p>where An(X) is the set of the causal ancestors of X
(which includes X itself), so this condition means that
we assume that X1 is not caused by the other observed
variables X2; X3.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Cooper [1] proved that:</title>
          <p>Theorem 1.1. Under the assumptions in De nition
1.1, the causal structure must be a subgraph of:
X1</p>
          <p>X2</p>
          <p>X3
Here, the directed arrows indicate a direct causal
relationship and the bidirected edge denotes an unobserved
confounder.</p>
          <p>Our primary interest is to predict p(X3jdo(X2)), the
distribution of X3 after an intervention on X2. In
general, this quantity may di er from p(X3jX2), the
conditional distribution of X3 given X2 [13]. In the
linear-Gaussian case, the quantity
@E(X3jdo(X2))
measures the causal e ect of X2 on X3.</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>It is easy to show that in the LCD setting, these quantities are equal:</title>
          <p>Corollary 1.1. Under the LCD assumptions in
Definition 1.1,</p>
          <p>p(X3jdo(X2)) = p(X3jX2):
Therefore, in the linear-Gaussian case, the quantity
@E(X3jdo(X2))
1.2</p>
          <p>Relationship with instrumental variables</p>
        </sec>
        <sec id="sec-3-2-4">
          <title>The other model relevant for our discussion is the so called instrumental variable model. Following Pearl [13], we de ne:</title>
          <p>De nition 1.2. (Instrumental Variable setting)</p>
        </sec>
        <sec id="sec-3-2-5">
          <title>Given three random variables X1; X2; X3, we call X1</title>
          <p>an instrumental variable if the following conditions are
satis ed:
Statistical dependences:</p>
          <p>X1 6?? X2</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Prior assumptions:</title>
        <p>X1 ?? X3jdo(X2)</p>
        <sec id="sec-3-3-1">
          <title>Faithfulness</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>The second assumption says that X1 and X3 are inde</title>
          <p>pendent after an intervention on the variable X2. In
terms of the causal graph, this means that all the
unblocked paths between X1 and X3 contain an arrow
that points to X2.</p>
        </sec>
        <sec id="sec-3-3-3">
          <title>Unfortunately the instrumental variable property cannot be directly tested from observed data. The causal graph for the IV setting is a subgraph of:</title>
          <p>X1
X2
X3</p>
        </sec>
        <sec id="sec-3-3-4">
          <title>So, a possible confounder between X2 and X3 is al</title>
          <p>lowed, in contrast with the LCD setting. Note that
the LCD setting is a special case of the IV model.
Lemma 1.1. Under the IV assumptions in De nition
1.2 and for the linear-Gaussian case, the quantity
Cov(X1; X3)</p>
          <p>Cov(X1; X2)
is a valid estimator for the causal e ect of X2 on X3.
1.3</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Type II errors in LCD</title>
        <sec id="sec-3-4-1">
          <title>In practice, the con dence on the result of the con</title>
          <p>ditional independence test X1 ?? X3jX2 in the LCD
setting depends on the sample size. Indeed, it could
be hard to distinguish a weak conditional dependence</p>
          <p>X1 6?? X3jX2
from a conditional independence using a sample of
nite size. Here we study the question of what happens
to our prediction of the causal e ect of X2 on X3 if
the conditional independence test encounters a type II
error (i.e., erroneously accepts the null hypothesis of
independence).</p>
          <p>Note that a type I error (i.e., erroneously rejecting the
null hypothesis of independence) in the tests X1 6?? X2
and X2 6?? X3 will not be as dangerous as a type II
error in the conditional independence test. Indeed, the
probability of a type I error can be made arbitrarily
small by tuning the signi cance level appropriately. In
addition, a type I error would let the LCD algorithm
reject a valid triple, i.e., lower the recall instead of
leading to wrong predictions.</p>
          <p>For these reasons we study the model described in the
following de nition, which allows the presence of a
hidden confounder X4, and a direct e ect from X1 on X3
(not mediated via X2). We assume that these
additional features result in a possible weak conditional
dependence between X1 and X3 given X2. For
simplicity we consider only the linear-Gaussian case. We
also assume no confounders between X1 and X2, or
between X1 and X3, or between X1; X2; X3. This
simpli cation will not in uence the nal result of the
paper, because we will prove how unboundedness of the
causal e ect estimation error is already achieved for
this special case.</p>
        </sec>
        <sec id="sec-3-4-2">
          <title>De nition 1.3. We assume that the \true" causal model has the following causal graph:</title>
          <p>X4</p>
          <p>X2
X1</p>
          <p>X3
which is one of the possible causal structures that is
compatible with the following conditions:
Statistical dependences:</p>
          <p>X1 6?? X2
X2 6?? X3
A weak conditional dependence</p>
          <p>X1 6?? X3jX2</p>
        </sec>
      </sec>
      <sec id="sec-3-5">
        <title>Prior assumptions:</title>
        <sec id="sec-3-5-1">
          <title>Faithfulness</title>
          <p>An(X1) \ fX2; X3g = ;</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>The observed random variables are X1; X2; X3 while</title>
        </sec>
        <sec id="sec-3-5-3">
          <title>X4 is a hidden confounder, assumed to be independent from X1.</title>
          <p>The joint distribution of the observed variables is
assumed to be a multivariate Gaussian distribution with
covariance matrix and zero mean vector. We also
assume that the structural equations of the model are
linear. Then
where
is the vector of the extended system,</p>
          <p>X = AX + E;
X =</p>
          <p>X1; : : : ; X4
E = E1; : : : ; E4</p>
          <p>
            T
T
(
            <xref ref-type="bibr" rid="ref2">2</xref>
            )
is the vector of the independent noise terms, such that
E
          </p>
          <p>N 0;
= diag i2 ;
and A = ( ij ) 2 M4 R is (up to a permutation of
indices) a real upper triangular matrix in the space
M4(R) of real 4 4 matrices that de nes the causal
strengths between the random variables of the system.</p>
        </sec>
        <sec id="sec-3-5-4">
          <title>Remark 1.1. In [14], an implicit representation for</title>
          <p>the confounder X4 is used, by using non-zero
covariance between the noise variables E2; E3. It can be
shown that for our purposes, the two representations
are equivalent and yield the same conclusions.</p>
        </sec>
        <sec id="sec-3-5-5">
          <title>In the Gaussian case, a conditional independence is equivalent to a vanishing partial correlation:</title>
          <p>Lemma 1.2. Given a set of three random variables
(X1; X2; X3) with a multivariate Gaussian distribution
the conditional independence</p>
          <p>
            X1 ?? X3 j X2
is equivalent to a vanishing partial correlation
13 2 = q 1
13
23:
In contrast with the LCD model in De nition 1.1, the
equality (
            <xref ref-type="bibr" rid="ref1">1</xref>
            ) no longer holds. We are interested in the
error in the estimation of the e ect of X2 on X3 that
would be due to a type II error of the conditional
independence test in the LCD algorithm. The next section
is dedicated to the analysis of the di erence between
the true value (
            <xref ref-type="bibr" rid="ref4">4</xref>
            ) and the estimated one in (
            <xref ref-type="bibr" rid="ref1">1</xref>
            ):
jE X3jX2
          </p>
          <p>
            E X3jdo(X2) j = jg A;
jjX2j;
where the \causal e ect estimation error" is given by:
(
            <xref ref-type="bibr" rid="ref3">3</xref>
            )
(
            <xref ref-type="bibr" rid="ref4">4</xref>
            )
32
22
g A;
=
23:
(
            <xref ref-type="bibr" rid="ref5">5</xref>
            )
2
          </p>
          <p>Estimation of the causal e ect error
from the observed covariance matrix</p>
        </sec>
        <sec id="sec-3-5-6">
          <title>The following proposition gives a set of equations for the observed covariance matrix , given the model parameters A; and the linear structural equation model (2).</title>
          <p>Proposition 2.1. The mapping : (A; ) 7! that
maps model parameters (A; ) to the observed
covariance matrix according to the model in De nition 1.3
is given by:
11 = 12
12 =
13 =
11 23 =
11 22 =
11 33 =</p>
          <p>2
12 1</p>
        </sec>
        <sec id="sec-3-5-7">
          <title>The individual components in (6){(11) can now be obtained by straightforward algebraic calculations.</title>
          <p>
            Remark 2.1. (Instrumental variable estimator)
From equation (
            <xref ref-type="bibr" rid="ref8">8</xref>
            ) it follows immediately that for
13 = 0, we have
which corresponds to the usual causal e ect estimator
in the instrumental variable setting [3].
          </p>
        </sec>
        <sec id="sec-3-5-8">
          <title>The lemma we present now re ects the fact that we</title>
          <p>are always free to choose the scale for the unobserved
confounder X4:
Lemma 2.1. The equations of proposition 2.1 are
invariant under the following transformation
4j =</p>
          <p>Remark 2.2. (Geometrical Interpretation) From
a geometrical point of view the joint system of
equations for the observed covariance matrix de nes a
manifold M in the space of the model parameters
M4(R) D 2 , where M4(R) is the space of the
possible causal strengths ij and</p>
          <p>
            D 2 =
3
Y[0; ii]
i=1
is the compact hypercube of the noise variances. Note
that we have used the symmetry 44 = 42 = 1 and
that
2
i
ii
from equations (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ), (
            <xref ref-type="bibr" rid="ref10">10</xref>
            ) and (
            <xref ref-type="bibr" rid="ref11">11</xref>
            ). Note that the map
: (A; ) 7! is not injective. This means that
given an observed covariance matrix , it is not
possible to identify the model parameters in a unique way.
Indeed, the number of equations is six, while the
number of model parameters is eight. Geometrically, this
means that the manifold M does not reduce to a
single point in the space of model parameters.
Nevertheless it is still an interesting question whether the
function g is a bounded function on M or not, i.e.,
whether we can give any guarantees on the estimated
causal e ect. Indeed, for the instrumental variable
case with binary variables, such bounds can be derived
(see, e.g., [13]).
          </p>
        </sec>
        <sec id="sec-3-5-9">
          <title>The following Theorem and its Corollary are the main</title>
          <p>results of this paper. We will prove that there still
remain degrees of freedom in the noise variances 22; 32
and the signs s1; s2, given the observed covariance
matrix , that will lead to an unbouded causal e ect
estimation error g(A; ).</p>
          <p>Theorem 2.1. Given the causal model in De nition
1.3, there exists a map
: M3(R)</p>
          <p>D( )
f 1; +1g2 !</p>
          <p>
            M4(R) (
            <xref ref-type="bibr" rid="ref12">12</xref>
            )
such that for all (A; ):
( (A; ); 22; 32; s1; s2) = A:
(
            <xref ref-type="bibr" rid="ref13">13</xref>
            )
Here D( ) = [0; m= 11] [0; det =m] R2 is the
rectangle where the noise variances of X2 and X3 live,
with m de ned below in (19). The map gives explicit
solutions for the causal strengths ij , given the
observed covariance matrix , the noise variances 22; 32
and signs si = 1. The components of are given by:
12 =
42 = s1
43 = s2
12
11
r m
          </p>
        </sec>
        <sec id="sec-3-5-10">
          <title>The signs in the equations are a consequence of the second degree polynomial equations.</title>
          <p>Corollary 2.1. It is possible to express the error in
the estimated causal e ect as
g</p>
          <p>( ; 22; 32; s1; s2);
s1s2
pdet
m 32pm
mp 2
2
= # 12 +
m 22
11 22 :
By optimizing over 32 we get:
with
b ( 22) =</p>
          <p>m
The length of the interval [b ; b+] is a function of
( ; 22) and satis es
11 22 :
(21)
23 2 [b ; b+]</p>
          <p>R;
pdet
pm
mp 2</p>
          <p>2
23 =
22
m
+ # 12 :
m 22</p>
        </sec>
        <sec id="sec-3-5-11">
          <title>Proof. Equation (20) follows from (18) and:</title>
          <p>
            From equation (
            <xref ref-type="bibr" rid="ref11">11</xref>
            ), combined with the results of
Theorem 2.1, we can obtain the following inequality, using
also the fact that 32 11 &gt; 0:
Proof. The proof proceeds by explicitly solving the
system of equations (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ){(
            <xref ref-type="bibr" rid="ref11">11</xref>
            ). Some useful identities
are:
(
            <xref ref-type="bibr" rid="ref14">14</xref>
            )
(15)
(16)
m 32pm
and the most important one for our purpose:
          </p>
        </sec>
        <sec id="sec-3-5-12">
          <title>The two solutions of the inequality de ne the interval</title>
          <p>[b ; b+]. Its length is a decreasing function of 22.</p>
        </sec>
        <sec id="sec-3-5-13">
          <title>Unfortunately, the causal e ect strength 23 in equa</title>
          <p>tion (18) is unbounded. This means that for all the
choices of the observed covariance matrix that are
in accordance with the model assumptions in De
nition 1.3, the set of model parameters (A; ) 2 M
that would explain leads to an unbounded error g.</p>
        </sec>
        <sec id="sec-3-5-14">
          <title>Indeed, a singularity is reached in the hyperplane</title>
          <p>2 = 0, which corresponds to making the random
2
variable X2 deterministic with respect to its parents
X1; X4. Figure 1 shows the singularity of the function
jg( ; 22; 32)j in the limit 22 ! 0. The rate of growth is
proportional to the inverse of the standard deviation
of the noise variable E2:</p>
          <p>1
jgj /
2
as 2 ! 0.</p>
          <p>(22)</p>
          <p>Remark 2.3. (Lower bound for 22) Corollary 2.1
is the main result of our analysis. The right hand term
in (20) consists of two terms: the rst one, through #,
represents the contribution of the partial correlation,
and is small if 13 2 is small. The second term is a
fundamental, intrinsic quantity not controllable from
the conditional independence test and the sample size.
However, in situations where one is willing to assume
a lower bound on 22:
it is possible to give a con dence interval [b+; b ] for
the function g, depending on the choice of the lower
bound ^22.</p>
          <p>Remark 2.4. (IV estimation error)
In the instrumental variable literature the IV
estimator is used, presented in Lemma 1.1. Unfortunately,
this estimator and its error function
13
(23)
is proportional to 13 and from (17) one can deduce a
similar growing rate of the function h in terms of the
variance of the noise term E2:
jhj /
1
2
as 2 ! 0.</p>
          <p>(24)
Remark 2.5. (Singularity analysis)
Figure 2 shows a contour plot of jgj on the rectangle
D( ) 3 ( 22; 32). The singularity in the causal e ect
function g is reached in the degenerate case, when the
conditional distribution of X2 given X1 and X4
approaches a Dirac delta function. This cannot be
detected empirically, as we can still have well-de ned
covariance matrices of the observed system even if the
covariance matrix of the extended one is degenerate.
Let us investigate in detail the limit for 22 ! 0 from
the point of view of the causal model. This
proposition will show a simple example of how the causal
strengths can be arbitrarily large, keeping the entries
of the observed covariance matrix ij nite.
Proposition 2.2. Assume that the observed
covariance matrix is positive-de nite. Then, for the limit
2
2 ! 0 we have the following scenario for the causal
strength parameters:
8
&gt; 23
&lt;</p>
          <p>43
&gt;
: 13</p>
          <p>
            1
2
sgn( 42) 2 1
sgn(
            <xref ref-type="bibr" rid="ref12">12</xref>
            ) 2 1:
This limit, in which our error in the estimated causal
e ect strength of X2 on X3 diverges, is illustrated in
Figure 3.
          </p>
          <p>X1
12
X2
1</p>
          <p>1
1</p>
          <p>X3
Corollary 2.1 shows how the causal e ect estimation
error can be extremely sensitive to small perturbations
of our model assumptions. Equation (20) holds for any
value of # (which is proportional to the partial
correlation 13 2) and the second term vanishes when the
confounder is not present. This shows that with a
nite sample, a type II error in the conditional
independence test may lead to an arbitrarily large error in
the estimated causal e ect. Even in the in nite sample
limit, this error could be arbitrarily large if faithfulness
is violated. The result is in agreement with the results
in [9], and it shows in a clear algebraic way how type</p>
        </sec>
        <sec id="sec-3-5-15">
          <title>II errors of conditional independence tests can lead to</title>
          <p>wrong conclusions.</p>
          <p>We believe that this conclusion holds more generally:
even when we increase the complexity and the number
of observed variables, the in uence of confounders will
still remain hidden, mixing their contribution with the
visible parameters, thereby potentially leading to
arbitrarily large errors. This means that for individual
cases, we cannot give any guarantees on the error in
the estimation without making further assumptions.</p>
        </sec>
        <sec id="sec-3-5-16">
          <title>An interesting question for future research is whether</title>
          <p>this negative worst-case analysis can be supplemented
with more positive average-case analysis of the
estimation error. Indeed, this is what one would hope if</p>
        </sec>
        <sec id="sec-3-5-17">
          <title>Occam's razor can be of any use for causal inference problems.</title>
        </sec>
        <sec id="sec-3-5-18">
          <title>Other possible directions for future work are:</title>
          <p>Study more complex models, in terms of
the number of nodes, edges and cycles.
Bayesian model selection: We hope that the</p>
        </sec>
        <sec id="sec-3-5-19">
          <title>Bayesian approach will automatically prefer a simpler model that excludes a possible weak conditional dependence even though the partial correlation from the data is not exactly zero.</title>
          <p>Bayesian Information Criterion: We could
directly assign a score based on the likelihood
function of the data given the model parameters
(A; ) and the model complexity, without
assuming any prior distribution for the model
parameters.</p>
          <p>Nonlinear structural causal equations: To
deal with nonlinearity it is possible to consider
Spearman's correlation instead of the usual one,
using the following relationships:
m =
=
=
! =
# =
\Environment" variable: In many
applications in biology, for example where X1 is
genotype, X2 gene expression and X3 phenotype,
the observed random variables X2 and X3 are
strongly dependent on the environmental
conditions of the experiment. It might be reasonable
to assume that most of the external variability
is carried by the covariance between the
environment variable W and the other measured ones,
including possible confounders. This leads to the
following graphical model, which could be useful
in deriving some type of guarantees for this
scenario:</p>
          <p>W
X4</p>
          <p>X2
X1</p>
          <p>X3</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>Acknowledgements</title>
        <sec id="sec-3-6-1">
          <title>We thank Tom Heskes for posing the problem, and</title>
        </sec>
        <sec id="sec-3-6-2">
          <title>Jonas Peters for inspiring discussions. We thank the reviewers for their comments that helped us improve the manuscript.</title>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G. F.</given-names>
            <surname>Cooper</surname>
          </string-name>
          .
          <article-title>A simple constraint-based algorithm for e ciently mining observational databases for causal relationships</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          ,
          <volume>1</volume>
          :
          <fpage>203</fpage>
          {
          <fpage>224</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.J.</given-names>
            <surname>Bowden</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.A.</given-names>
            <surname>Turkington</surname>
          </string-name>
          . Instrumental Variables. Cambridge University Press,
          <year>1984</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Didelez</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Sheehan</surname>
          </string-name>
          .
          <article-title>Mendelian randomization as an instrumental variable approach to causal inference</article-title>
          .
          <source>Statistical Methods in Medical Research</source>
          ,
          <volume>16</volume>
          :
          <fpage>309</fpage>
          {
          <fpage>330</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Greenland</surname>
          </string-name>
          .
          <article-title>An introduction to instrumental variables for epidemiologists</article-title>
          .
          <source>International Journal of Epidemiology</source>
          ,
          <volume>29</volume>
          :
          <fpage>722</fpage>
          {
          <fpage>729</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Lawlor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Harbord</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A. C.</given-names>
            <surname>Sterne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Timpson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. D.</given-names>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Mendelian randomization: Using genes as instruments for making causal inferences in epidemiology</article-title>
          . Statistics in Medicine,
          <volume>27</volume>
          :
          <fpage>1133</fpage>
          {
          <fpage>1163</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.D.</given-names>
            <surname>Angrista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. G.</given-names>
            <surname>Imbens</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.B.</given-names>
            <surname>Rubinc</surname>
          </string-name>
          .
          <article-title>Identi cation of causal e ects using instrumental variables</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          ,
          <volume>91</volume>
          :
          <fpage>444</fpage>
          {
          <fpage>455</fpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. A. Jaeger J.</given-names>
            <surname>Bound</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Baker</surname>
          </string-name>
          .
          <article-title>Problems with instrumental variables estimation when the correlation between the instruments and the endogeneous explanatory variable is weak</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          ,
          <volume>90</volume>
          :
          <fpage>443</fpage>
          {
          <fpage>450</fpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Emmert-Streib</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Storey</surname>
          </string-name>
          .
          <article-title>Harnessing naturally randomized transcription to infer regulatory relationships among genes</article-title>
          .
          <source>Genome Biology</source>
          ,
          <volume>8</volume>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Scheines J. M. Robins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Spirtes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wasserman</surname>
          </string-name>
          .
          <article-title>Uniform consistency in causal inference</article-title>
          .
          <source>Biometrika</source>
          ,
          <volume>90</volume>
          :
          <fpage>491</fpage>
          {
          <fpage>515</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Spirtes</surname>
          </string-name>
          .
          <article-title>Strong faithfulness and uniform consistency in causal inference</article-title>
          .
          <source>In Proceedings of the Nineteenth Conference on Uncertainty in Arti cial Intelligence (UAI</source>
          <year>2003</year>
          ), pages
          <fpage>632</fpage>
          {
          <fpage>639</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Uhler</surname>
          </string-name>
          , G. Raskutti, P. Buhlmann, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Geometry of the faithfulness assumption in causal inference</article-title>
          .
          <source>The Annals of Statistics</source>
          ,
          <volume>41</volume>
          :
          <fpage>436</fpage>
          {
          <fpage>463</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Spirtes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. N.</given-names>
            <surname>Glymour</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Scheines</surname>
          </string-name>
          .
          <article-title>Causation, prediction, and search</article-title>
          . The MIT Press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          .
          <article-title>Causality: models, reasoning and inference</article-title>
          . Cambridge University Press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Drton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Foygel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Sullivan</surname>
          </string-name>
          .
          <article-title>Global identi ability of linear structural equation models</article-title>
          .
          <source>The Annals of Statistics</source>
          ,
          <volume>39</volume>
          :
          <fpage>865</fpage>
          {
          <fpage>886</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>