<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recovering from Selection Bias using Marginal Structure in Discrete Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Robin J. Evans</string-name>
          <email>evans@stats.ox.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vanessa Didelez</string-name>
          <email>Vanessa.Didelez@bristol.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Statistics, University of Oxford</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Mathematics, University of Bristol</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper considers the problem of inferring a discrete joint distribution from a sample subject to selection. Abstractly, we want to identify a distribution p(x; w) from its conditional p(x j w). We introduce new assumptions on the marginal model for p(x), under which generic identi cation is possible. These assumptions are quite general and can easily be tested; they do not require precise background knowledge of p(x) or p(w), such as proportions estimated from previous studies. We particularly consider conditional independence constraints, which often arise from graphical and causal models, although other constraints can also be used. We show that generic identi ability of causal e ects is possible in a much wider class of causal models than had previously been known.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Selection bias occurs when samples are obtained from
a population in a manner which depends upon the
attributes of the samples themselves. This can happen
by accident|such as survey participants self-selecting
in a way which is correlated with their responses|or
by design, for example in a case-control study.
In this paper we show that under plausible and testable
structural hypotheses about the relationships between
the variables being measured, we can recover from
selection bias, even without having external information
about the distribution of variables in the population.
Example 1.1. Consider a case-control study in which
participants are selected according to a disease status
W with dw = 2 levels. We are interested in the e ect
of a discrete measured treatment X on W which,
assuming no confounding is the same as estimating the</p>
      <p>X</p>
      <p>Y</p>
      <p>W
conditional distribution p(w j x). Since our data are
selected according to disease status, what we observe,
ignoring sampling error, is the conditional distribution
p(x j w). It is well known that we can use p(x j w) to
obtain the causal odds-ratio for X on W , which is one
measure of the causal e ect, but that the full
conditional distribution p(w j x) is generally not identi able.
However, suppose that we also measure a covariate
Y , which is known to be marginally independent of
X (see Figure 1). The marginal independence means
that p(x; y) = p(x) p(y), so</p>
      <p>X p(w)p(x; y j w)
w</p>
      <p>!
=</p>
      <p>X p(w)p(x j w)
w</p>
      <p>X p(w)p(y j w)
w
!
(1)
for each value of x; y; letting dx; dy be the number
of levels of X and Y respectively, this gives at most
(dx 1)(dy 1) non-redundant equations. Since W
is binary, each equation is quadratic in the single
unknown p(w = 0) and can have at most two solutions.
The true marginal distribution of W must be a
solution to each equation, so the distribution is identi able
up to at most two solutions1.</p>
      <p>To be concrete, suppose we were given the following
conditional distributions p(x; y j w).</p>
      <p>1Provided that Y 6?? W j X; see the discussion in the
next section.</p>
      <p>W = 0
0
1
Since X and Y are positively correlated given W = 0,
and negatively correlated given W = 1, it is clear that
by taking some appropriate convex combination, we
should nd a table for independence. In fact, choosing
p(w = 0) = 14 gives the marginal table:
Having recovered p(w = 0) we obtain the entire joint
distribution p(x; y; w) = p(w) p(x; y j w) and
therefore any causal e ects identi able from it, including
p(w j x).</p>
      <p>Note that no information external to the study was
required (e.g. background knowledge of the distribution
of p(x)) other than the marginal independence X ?? Y .
In addition, we will see that this condition is in
principle testable by equality and inequality constraints. In
particular (1) may have no solution, so the model can
be refuted.</p>
      <p>Intuitively the above method works because the
independence X ?? Y is `destroyed' if the marginal
distribution of W is speci ed incorrectly (except in certain
degenerate cases). Hence a correct speci cation of the
marginal p(w) results in a distribution in the marginal
model for (X; Y ), and an incorrect speci cation
usually does not.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Previous Work and Set-Up</title>
      <p>
        Selection bias has received much attention in the
statistical literature for many decades dating back at least
to [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The term is now used to refer to a very wide
range of problems: bias can occur due to selection into
the sample, drop-out, missingness or due to possibly
inadvertently and inappropriately stratifying the
statistical analysis. When using DAGs to represent the
situation, selection bias can often be illustrated as the
result of conditioning on a variable that is a common
child of other variables, sometimes called
\colliderstrati cation bias" [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In some of these cases, the
bias can be avoided e.g. by inverse-probability
weighting given a su ciently informative set of covariates
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Further, in the context of a case-control study
it is well known that the odds-ratio can be recovered
under very mild assumptions, see [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for a recent
application relying on an argument based on conditional
independences in DAGs. However, many popular
statistical methods aiming to attenuate selection bias rely
on speci c assumptions about the selection mechanism
or external (quantitative) information, see e.g. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] in
the context of meta-analyses. Here, we aim to exploit
conditional or marginal independences and the
constraints they impose on the model. Our approach is
similar in spirit to [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], who use `distortions' in
multivariate Gaussian distributions caused by selection to
identify the parameters of the original model.
Recently, a number of papers have considered the
speci c question of identifying causal e ects under
selection. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] give conditions for identi cation of causal
odds-ratios as well as the testability of the causal null
hypothesis based on assumptions that can be read o
DAGs with unobservable variables. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] extend these
results and additionally propose a method of controlling
selection bias using instrumental variables to enable
recovery of e ect measures other than odds-ratios.
Regarding general causal inference under selection, [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
provide results showing the impossibility of full
identi cation of causal e ects under selection in a large
class of Markovian models. In contrast, we achieve
a much wider range of identi cation results by using
the weaker notion of `generic identi ability' (de ned
in Section 2). We exploit equality and inequality
constraints derived from known conditional or marginal
independences. This is closely related to ancestral
graph models, which fully characterize the conditional
independence constraints implied by selection in DAG
models [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Similarly, [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] noted such constraints and
showed that selection in DAGs could be used to
generate arbitrary hierarchical models.
      </p>
      <p>Several of the above authors formalise the problem of
selection by including a binary random indicator S of
whether or not an individual is selected in the sample.
The available data then comes from the conditional
distribution p(x; wjs = 1), where X; W may be vectors
of random variables. There is little we can do without
further assumptions, but suppose that X ?? S j W , i.e.
the selection only depends upon the variables in W ;
in this case, provided p(w j s = 1) &gt; 0, we can recover
p(x j w) = p(x j w; s = 1). If the distribution of W in
the population is known, say from a previous study,
then we can obtain p(x; w) and in e ect `recover' from
the selection bias. In this paper we will assume that
no such information is available, and ask what
assumptions allow p(x; w) to be recovered from p(x j w). We
assume throughout that p(x; w) &gt; 0.</p>
      <p>The rest of the paper is structured as follows. After
formally de ning generic identi ability in Section 2,
we address in Section 3 the role of constraints on the
marginal model p(x) for recovering p(x; w). In Section
4 we then consider special cases where the marginal
constraints arise from various independence
assumptions. While we use directed acyclic graphs (DAGs)
throughout to depict marginal or conditional
independences, Secion 5 formally addresses situations where
the model is explicitly de ned by a DAG, such as is
typical for causal models. It turns out that it is an
important prerequisite for identi cation that the model
be non-decomposable. Further implications and
examples for causal inference are discussed in Section 6,
and practical considerations given in Section 7.
2</p>
      <sec id="sec-2-1">
        <title>Identi ability</title>
        <p>We distinguish `strict' (sometimes called `global') and
`generic' identi ability. The latter allows identi
ability to fail on a lower dimensional subset O of the model
M. It can be argued that, under certain assumptions,
observations are `unlikely' to lie exactly in such a
subset.</p>
        <p>Consider random variables (X; W ) taking values in a
nite discrete product space X W. Let
=</p>
        <p>X W
fp &gt; 0 : X p(x; w) = 1g
x;w
be the strictly positive probability simplex of
distributions over X W. An algebraic model M
is a collection of probability distributions satisfying a
nite set of polynomial constraints in the
probabilities p(x; w). For example, the model M under which
X ?? W is the set of p 2 such that
p(x; w)</p>
        <p>!
X p(x; w0)
w0</p>
        <p>X p(x0; w)
x0
!
= 0;
8w; x:
Consider the act of conditioning on W , de ned by the
rational map : p(x; w) 7! p(x; w)=p(w). Let N =
(M) be the image of the model under this operation,
which is also an irreducible algebraic variety in the
invariants p(x j w) [6, Proposition 4.5.6].</p>
        <p>De ne the bre of</p>
        <p>at p 2 M as the set</p>
        <p>F (p) = fq 2 M : (q) = (p)g:
An injective map is one for which all bres have
cardinality one.</p>
        <p>De nition 2.1. Let k 2 N. We say that M is
generically k-identi able if the bres F (p) have cardinality
at most k for all p 2 M n O, where O is some proper
(i.e. lower dimensional) algebraic subset of M.
In the case k = 1 then is just generically identi able.
If a model is not generically k-identi ed for any k 2 N
then it is unidenti able.</p>
        <p>In other words, generic k-identi ability requires the
map to be at most k-to-one on a set that contains
`almost all' the distributions. Generic identi ability
corresponds to the map being injective on M n O.
Throughout this article we use `generically' to mean
`at all but a strict sub-variety of the parameter values
within the model (M)'.</p>
        <p>If O = ; then this is sometimes referred to as strict
(k-)identi ability. Strict identi ability of p(x; z) is a
very stringent condition and essentially never occurs
in our framework. If W ?? X, for example, it is clear
that p(x j w) = p(x) does nothing to help us identify
p(w), nor therefore p(x; w). We consider the weaker
notion of generic identi ability so that identi ability
may fail on a small (i.e. measure zero) subset O.
We say a model on X; W is testable if it places
nontrivial equality constraints on p(x j w), i.e. N is a strict
sub-variety of the set of all conditional distributions
p(x j w). We will work throughout the paper as though
we can observe p(x j w) itself; in practice we would
instead have a consistent estimator with asymptotically
correct standard errors.
3</p>
      </sec>
      <sec id="sec-2-2">
        <title>General Marginal Models</title>
        <p>In this section we consider the possibility of using
constraints on the marginal model of X to recover p(x; w)
from p(x j w). The rst result gives us a necessary
condition for identi ability.</p>
        <p>Lemma 3.1. Suppose that MX be a marginal model
for X, and let A be the (dx dw)-matrix with entries
ax;w = p(x j w). Then p(x; w) is k-identi able from
p(x j w) only if A has rank dw.</p>
        <p>If W ?? X then, as previously noted, the rank
condition will not be satis ed. It may fail in a more subtle
way if, for example, there is a lower dimensional
variable that mediates the relationship between X and W ;
see Example 6.3.</p>
        <p>
          In general it is di cult to characterize exactly when
a model will be generically identi able, but an
important special case comes when W 's dependence on X
is completely unrestricted: that is, our model makes
no restriction on p(w j x) no matter what the value of
p(x). This variation independence is known as a
`parameter cut', or simply a cut [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>Theorem 3.2. Let M be a model for (X; W ) with a
cut between X and W jX, such that p(w j x) is
unrestricted and the marginal model for p(x) is a variety
of dimension dx 1 l.</p>
        <p>Then given p(xjw), the variety de ned by the bre
fq(x; w) 2 M : q(x j w) = p(x j w)g
generically has dimension max(0; dw 1 l). In
particular, p(x; w) is generically k-identi able if and only
if dw l + 1.</p>
        <p>The quantity l is the number of independent
constraints on the marginal model for X. We note that
Theorem 3.2 implies that some constraints on p(x) are
needed; when l = 0 the inequality in Theorem 3.2 is
only satis ed if W is constant, meaning that there is
in fact no selection.</p>
        <p>Proof. Since p(x j w) is known, we just need (w) such
that Pw (w)p(x j w) is contained in the marginal
model for X, i.e. nd a point which is in the marginal
model for X as well as the column span of the matrix
P with (x; w)th entry p(x j w).</p>
        <p>Let q(x) be a point in both the marginal model for X
and the column span of P . Let the kernel of the
Jacobian of constraints which de ne the marginal model
for X be K, which is generically of rank dx 1 l by
assumption.</p>
        <p>The tangent space at q of the variety of distributions
that satisfy both the marginal model and lie within P
must lie within the intersection of the column spans of
P and K. By Proposition A.1 any set of min(l + 1; dw)
columns of P are generically linearly independent of
K. The dimension of the tangent space of the
intersection is then the number of remaining columns, which
is max(0; dw l 1). Generically we may choose this
point so that the dimensions are equal (in the language
of algebraic geometry, the two manifolds are
generically transverse).</p>
        <p>It is important to keep in mind that generic identi
ability allows for identi ability to fail on interesting
sub-models. It may, therefore, be a serious problem if
W is (conditionally) independent of some part of X.
Example 3.3. Consider the graphical model in Figure
2, which implies</p>
        <p>p(x; y; z; w) = p(y j x)p(z j x)p(x j w)p(w):
In this case Y ?? Z j X is a marginal constraint,
however W does not depend arbitrarily on X; Y; Z so we
cannot apply Theorem 3.2. In fact W ?? Y; Z j X so
p(x; y; z) = p(y j x)p(z j x) X p(w)p(x j w);
w
which satis es the required marginal constraint for any
p(w). We cannot, therefore, use this constraint to
identify p(w).</p>
        <p>This phenomenon is generalized in the following
lemma.</p>
        <p>Lemma 3.4. Suppose distributions in M are of the
form p(x; y; w) = p(x; w)p(y j x; w), where p(y j x; w)
is variation independent of p(x; w). Then p(x; y; w)
is identi able from p(x; y j w) if and only if p(x; w) is
identi able from p(x j w).</p>
        <p>Proof. We have</p>
        <p>p(x; y j w) = p(y j x; w) p(x j w);
so the two factors are recoverable from p(x; y j w).
Since p(y j x; w) is variation independent of p(x; w), it
is also variation independent of p(w) given p(x j w),
which are functions of p(x; w). It follows that no
restriction is placed on p(w) by p(y j x; w).</p>
        <p>
          Our method relies strongly on variation dependence
between p(w) and p(x j w); the kind of variation
independence in the Lemma and in Example 3.3 is not
helpful for identifying p(x; w). Note that the special
case in which X is a trivial random variable shows
that we must not have a parameter cut between the
marginal model for W and the distribution of XjW ;
contrast this with the conditions of Theorem 3.2. This
is related to the observations of [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] on semi-supervised
learning.
4
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Independences</title>
        <p>Marginal and conditional independence constraints
often arise in statistical models, commonly from Markov
assumptions such as those implied by causal models.
In this section we shall focus on when p(x; w) is
identi able in this context.
4.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Marginal Independence</title>
      <p>Consider again Example 1.1 in which X ?? Y but we
can only directly observe p(x; y j w). Consider each
p(x; y j w) for w 2 W as a single distribution for X; Y ;
we need to nd a convex combination of these
distributions which satis es the independence constraint.
If all the variables are binary we have two 2 2-tables
in the 3-dimensional probability simplex, and want to
nd a point on the line segment between these tables
which lies on the surface of marginal independence.
This is illustrated in Figure 3. Assuming that neither
1. the two conditional tables have opposite signed
correlations (or log-odds-ratios), in which case
there is exactly one convex combination that
satis es marginal independence;
2. the two tables have the same sign, and there is no
solution which satis es the independence;
3. the two tables have the same sign, but there are
one or two convex combinations which satisfy the
independence.</p>
      <p>Situation 2 allows the assumption of X ?? Y to be
refuted, and corresponds to an inequality constraint
on p(x; y j w). This is the only constraint on this model
in the binary case.</p>
      <p>
        The possibility of two solutions is due to the quadratic
nature of the independence constraint. A line segment
which intersects the surface of independence twice is
shown in Figure 3. If there are two distinct solutions
then there is no way to identify the true distribution
uniquely without further assumptions. However, this
situation can only occur if the joint distribution
exhibits a form of Simpson's paradox: X and Y are
positively (or negatively) correlated within each level of W ,
but are independent after collapsing across the levels
of W . It has been argued that this happens relatively
rarely in practice [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], and one can see from Figure 4
that it requires the conditional tables to lie in speci c
corners of the simplex.
      </p>
      <p>The method extends readily to non-binary X and Y ;
we obtain (dx 1)(dy 1) separate quadratics in the
unknown p(w = 0), and they will generically have
distinct solutions; it follows that p(w = 0) is the common
solution to these quadratics, and that the model is
generically identi able. If W is non-binary then we
obtain quadratic equations in multiple variables, but
analogous results are available.</p>
      <p>Lemma 4.1. Consider the model M on discrete
variables X; Y; W such that X ?? Y . Then p(w; x; y) is
generically k-identi able from p(x; y j w) if and only if
dw 1 (dx 1)(dy 1). Identi ability fails if either
X ?? Y; W or Y ?? X; W .</p>
      <p>If dw = 2 then p(w; x; y) is 2-identi able from
p(x; w j y) if X 6?? Y j W , and unidenti able otherwise.
The proof is in the appendix.</p>
      <p>Note that, under the causal graphical model in
Figure 1, the causal e ect of X on W is given by
p(w j do(x)) = p(w j x). This quantity is never strictly
identi able from p(x j w), but it is generically identi
able from p(x; y j w) under the marginal independence
model.</p>
      <p>Note that, although we lose identi ability of p(w) if
(for example) X ?? W j Y , we can still test this as a
causal null hypothesis because it is observable directly
from p(x; y j w).</p>
      <p>In the all-binary case, the marginal independence
model is not `testable' in the sense that it places no
equality constraints on p(x; y j w): thus for some
distributions p(x; y; w) which do not satisfy the marginal
independence, we can follow the procedure given
above and fallaciously `recover' some other
distribution q(w)p(x; y j w) which does satisfy the model.
However in the case where dx = 3, there are two
independent constraints, so we can use one to recover p(w)
and the other to check the model.
4.2</p>
    </sec>
    <sec id="sec-4">
      <title>Conditional Independence</title>
      <p>Consider the model de ned by the graph in Figure 4,
which implies that X ?? Y j Z. Under selection bias
on W we can only observe the conditional
distribution p(x; y; z j w). Following the same approach as in
Example 1.1 we obtain
p(z) p(x; y; z)
p(x; z) p(y; z) = 0;
8x; y; z;
if W is binary we obtain p(x; y; z) = p(x; y; zjw =
0) + (1 )p(x; y; zjw = 1), so each factor is linear in
the unknown = p(w = 0). This gives us quadratic
equations in , which are generically distinct for each
level of Z.</p>
      <p>Lemma 4.2. Let M be the model on X; Y; Z; W
dened by X ?? Y j Z. Then p(x; y; z; w) is
generically k-identi able from p(x; y; z j w) if and only if
dw 1 (dx 1)(dy 1)dz.</p>
      <p>If dw = 2 the model is unidenti able if and only if
X ?? Y j Z; W .</p>
      <p>The result is really just an application of Lemma 4.1
to separate models for each level of Z. If Z is trivial it
reduces to the marginal independence case. Note that
the condition X ?? Y j W; Z for non-identi ability is
testable directly from the observed conditional
distribution. The unidenti ability arises because any choice
of p(w) will give the required conditional
independence.</p>
      <p>
        Particular cases of unidenti ability arise when either
X ?? Y; W j Z or Y ?? X; W j Z (see Figure 5(a),
Example 3.3); in fact if W is binary then unidenti ability
is equivalent to at least one of these constraints holding
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In these instances, selecting on W does not destroy
the conditional independence we are using to recover
the joint distribution: we can observe it directly in the
conditional distribution. In a sense, because no
structure is lost, we cannot try to recover it by choosing the
correct margin p(w).
      </p>
      <p>As the next example illustrates, we can nd that extra
independences give `reduced' identi ability by
making some constraints redundant, without leading to
full unidenti ability. In particular, although the
allbinary case of a conditional independence under
selection generically gives two non-redundant quadratic
equations, in degenerate cases these equations may
have the same roots.</p>
      <p>Example 4.3. Consider the model in Figure 5(b), in
which Z ?? W j X; Y as well as X ?? Y j Z. In this case
p(x; y; z j w) = p(z j x; y) p(x; y j w):
In order to have X ?? Y j Z we need p(x; y; z) =
p(x; z)p(y j z). So since p(x; y; z) = p(z j x; y)p(x; y) =
p(z j x; y) Pw p(x; y j w)p(w), we need to pick p(w) so
that the odds-ratio between X and Y exactly cancels
out the x; y factor of p(z j x; y). This gives a single
quadratic equation; and solution to this gives p(w = 0)
such that p(w)p(x; y; z j w) is inside the model, so we
have generic 2-identi ability. Note that although there
is only one equation to solve in one unknown, unlike
the marginal independence model this one is testable,
since Z ?? W j X; Y can be checked directly from the
observed conditional p(x; y; z j w).</p>
      <p>The example shows that no three-way interaction
is present between X; Y; Z in p(x; y; z j w), and that
this remains true for any weighted combination
Pw (w)p(x; y; z j w). If the three-way interaction is
zero then the two constraints X ?? Y j Z = 0 and
X ?? Y j Z = 1 become equivalent, thus we have a
partial degeneracy: there is only one non-redundant
constraint to ful l. This issue is explored more
generally in Section 5.1.
5</p>
      <sec id="sec-4-1">
        <title>Directed Acyclic Graph Models</title>
        <p>
          In the previous sections, we used directed acyclic
graphs (DAGs), also known as Bayesian networks,
to represent marginal and conditional independences.
Here we give explicit results for situations where the
model is de ned by a DAG, such as is typical for causal
models. A DAG model is de ned by a recursive
factorization according to the structure of a graph, or
equivalently a collection of conditional independence
constraints. Let G be a directed acyclic graph with
disjoint sets of vertices V [_W representing vectors of
random variables XV ; XW . We make use of the fairly
standard terminology of parents (paG ), children,
ancestors (anG ), etc. See, for example, [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>We will assume that our distribution obeys the Markov
property with respect to a DAG G, so that
p(xV ; xW ) =</p>
        <p>p(xv j xpa(v)):</p>
        <p>Y</p>
        <p>Lemma 5.1. Let p(xV j xW ) be a conditional
distribution from a DAG. Then p(xV ; xW ) is identi able from
p(xV j xW ) if and only if p(xan(W )nW ; xW ) is identi
able from p(xan(W )nW j xW ).</p>
        <p>Proof. If V 6= anG (W ) n W then there exists some
childless v 2 V in G. Hence there is a
parameter cut between W [ V n fvg and fvgjW [ V n fvg,
so p(xV ; xW ) = p(xv j xV nv; xW ) p(xV nv; xW ) where
p(v j xV nv; xW ) = p(xv j xpa(v)) is variation
independent of p(xV nv; xW ). The result then follows from
Lemma 3.4 and repeating over all v 62 anG (W ).
In light of this result, we henceforth assume that all
variables are ancestors of the selection variables W .
5.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Hierarchical Models</title>
      <p>A hierarchical model over XV is the set of distributions
p(xV ) which factorize as
p(xV ) = Y</p>
      <p>C (xC )</p>
      <p>C2C
for some collection of inclusion maximal sets C;
undirected graphical models and decomposable models
are special cases. Note that if p(xV ; xW ) is
hierarchical with maximal sets C, then p(xV j xW ) =
p(xV ; xW )=p(xW ) also factorizes into functions over
the maximal sets C [ fW g. Any DAG model is
contained in the hierarchical model with maximal sets
fvg [ paG (v), which has consequences for our
conditional distributions.</p>
      <p>Lemma 5.2. Let p(xV ; xW ) obey the global Markov
property for a DAG G on V [ W . Then p(xV j xW ) is
of the form
p(xV j xW ) =</p>
      <p>W (xW )
v(xv; xpa(v)):</p>
      <p>(2)</p>
      <p>Y
That is, a hierarchical function with maximal sets W
and fvg [ paG (v); v 2 V [ W .</p>
      <p>This follows directly from writing out the factorization
for DAGs.</p>
      <p>Note that (2) yields testable constraints on p(xV j xW ),
but these constraints do not give any way to identify
p(xW ). Clearly the true p(xV ; xW ) must satisfy the
same hierarchical model, but so will any distribution
of the form q(xW )p(xV j xW ).</p>
      <p>Proposition 5.3. Let M be the model implied by a
DAG G, and M the hierarchical model with maximal
sets</p>
      <p>C =
fvg [ paG (v) : v 2 V [ W
:
Then p(xV ; xW ) is generically k-identi able from
p(xV j xW ) only if dw 1 d(M) d(M).
In particular p(xV ; xW ) is generically k-identi able
from p(xV j xW ) only if G is not decomposable.
The proof is given in the appendix. Proposition 5.3 is
illustrated by Example 3.3 where the model is indeed
decomposable, while in Example 4.3 it is not
decomposable. We conjecture that the converse of
Proposition 5.3 also holds, so that if dw 1 d(M) d(M)
then generic k-identi ability follows. We can obtain
the following weaker result from Theorem 3.2,
however.</p>
      <p>Proposition 5.4. Let G be a DAG with vertices V [
fwg such that paG (w) = V , and such that G imposes
l independent constraints on XV . Then p(xV ; xw) is
identi able from p(xV j xw) if and only if dw l + 1.
6</p>
      <sec id="sec-5-1">
        <title>Causal Models and Examples</title>
        <p>
          In the context of causal inference, our results will be
useful when it is known a priori that the causal e ect
of interest is identi ed from p(x; w), p(x), or p(wjx).
Before quantifying a causal e ect, one may rst want
to test the null hypothesis of no causal e ect. As
noticed before, when the null translates into a conditional
independence (or absence of a directed edge in a DAG),
this may actually hamper identi cation, but we
conjecture that the null can often be see from p(xjw) itself,
using e.g. result from [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>Example 6.1. It may be plausible that the causal
e ect of X on W , p(wjdo(x)), is identi ed given a
su cient set of covariates C, e.g. by the back-door
formula. In the context of a case control study,
however, one would then often just estimate the
(conditional) causal odds-ratio between X and W given C
as the most straightforward causal quantity. Our
results would be useful if data on additional covariates
Z is available, where it is known that, for example,
X ?? Z j C because subject matter tells us that X
is only determined by C. The conditional
independence then imposes the sort of constraints which may
enable identi cation of p(x; c; z; w). Hence we can
obtain other causal e ect measures, such as risk di
erences or risk ratios. In particular we can obtain the
intervention distribution p(yjdo(x)) by integrating out
C, so that unconditional causal parameters can be
reported which are more easily compared with results
from randomized controlled trials. However, problems
may occur: if the association between Z and W is
weak, identi cation becomes unstable, as with a weak
instrument in instrumental variable methods.
Additionally, if W represents a very rare disease, then the
marginal distribution of X is approximately the same
as that for the controls, p(x) p(xjw = 0), so
recovering p(x) exactly will typically not lead to interesting
new insights. However, our method could still be used
for sensitivity analysis in such situations.</p>
        <p>Example 6.2. Consider the causal model represented
by the graph in Figure 6 under selection on X4. This
1
2</p>
        <p>4
model arises in the context of dynamic treatment
regimes, where X1; X3 are treatments, X2; X4
outcomes, and X0 covariates. A typical quantity of
interest is p(x4 j do(x1; x3)), which can be expressed as
(for example) either of
p(x4 j do(x1; x3)) =</p>
        <p>X p(x4 j x1; x3; x0)p(x0)
x0
= X p(x4 j x1; x2; x3)p(x2 j x1):</p>
        <p>x2
However neither of these quantities is a function of
p(x0123 j x4), so we need to recover the distribution
p(x4) before proceeding. The graph exhibits the
conditional independence constraints X0 ?? X1 and
X0; X1 ?? X3 j X2 and these independences are not
visible after selection on X4, so they provide a method
for recovering p(x4).</p>
        <p>The conditional distribution has the form of a
hierarchical model
p(x0123 j x4) = (x0; x1; x2) (x2; x3) (x0; x1; x3; x4);
and so q(x4)p(x0123 j x4) factorizes in the same way for
any margin q(x4). Notice that the constraints which
are lost are precisely that X0 ?? X1 and X0 ?? X3 j X2.
These constitute (d0 1)(d1 1) and (d0 1)(d3 1)d2
constraints respectively, but the three-way interaction
X0; X2; X3 is not present in the hierarchical model,
so in fact there are only (d0 1)(d3 1) additional
constraints from the conditional independence (this is
essentially the same as what happens in Example 4.3).
Thus we can generically identify the model only if
(d0
1)(d1 + d3
2)
d4
1:
In the all-binary case this corresponds to two
constraints to resolve one parameter.</p>
        <p>Example 6.3. The ability to identify distributions in
DAG models may su er from `choke points' caused by
lower dimensional variables. Consider the model in
Figure 7 with selection on W . Then p(x; y; z; u j w) =
p(x; y; z j u)p(u j w), so that provided du 1 (dx
1)(dy 1)dz we can generically identify p(u) from
the rst factor. However, knowing p(u) and p(u j w)
will not allow us to generically identify p(w) unless
dw du. Such a lack of identi cation is because
p(w j x; y; z) is not unrestricted if dw &gt; du, leading to a
rank de ciency of the matrix with entries p(x; y; z j w)
and violating the conditions of Lemma 3.1. Note,
however, that this rank problem is detectable even if U is
unobserved.
7</p>
      </sec>
      <sec id="sec-5-2">
        <title>Discussion</title>
        <p>We envisage that the practical application of our
results will work as follows. In a given data situation,
we would need to know that sampling was subject to
selection and which variable(s) this a ected. Then, a
set of marginal or conditional independences the true
distribution should obey needs to be postulated.
Identi ability can then be checked using our results and if
it holds the full model can be t with appropriate
numerical procedures. Note that if these postulated
constraints are just enough to enable identi ability they
cannot empirically be tested and must be based on
subject matter knowledge which will typically be
informed by causal assumptions.</p>
        <p>In principle it is a relatively straightforward matter
to t these models; to perform maximum likelihood
estimation we need to maximize the conditional
loglikelihood
lXjW (p)</p>
        <p>X nxw log p(x j w)
x;w
= X nxw log p(x; w)</p>
        <p>x;w
= lXW (p)
lW (p):</p>
        <p>
          X nxw log p(w)
x;w
The TM algorithm of [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] deals with this computation
by linearizing the marginal log-likelihood lW (p) at each
iteration; this may be useful if the marginal likelihood
is hard to calculate but nding its gradient at a single
point is not. Nave numerical methods can also work.
The likelihood behaves much like that of a latent
variable model, with some parameter values leading to
poor identi ability in nite samples. The stronger
the dependence between W and the variable(s) in the
marginal model, the better. If identi cation fails, this
will be manifested as slow convergence and a at
loglikelihood. The likelihood may be multi-modal, so
trying di erent starting points would be advisable.
The all-binary conditional independence model from
Section 4.2 requires the existence of a common root to
two separate quadratic equations. This corresponds
to a single polynomial constraint on the conditional
probabilities p(x; y; z j w), which can be used to test
the model. Using the computational algebra software
Singular we can compute this polynomial, and
determine that it is homogeneous of degree 8, but it does
not obviously admit a simple interpretation.
Such constraints may be considered analogous to
the Verma constraints of [
          <xref ref-type="bibr" rid="ref16 ref19">16, 19</xref>
          ], which arise from
marginal distributions of models de ned by
conditional independence constraints.
7.2
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Extensions</title>
      <p>
        All the variables in our causal models were assumed
to be observed, but in principle an extension could be
made to conditional independence or other constraints
which arise in the presence of latent variables. We
conjecture that distributions arising from the graph in
Figure 6 are recoverable even if X0 is unobserved, for
example (this is the Verma model [
        <xref ref-type="bibr" rid="ref16 ref19">16, 19</xref>
        ]).
Although this paper only considers discrete variables,
some of the results here could be extended to the
case where the marginal model involves continuous
variables, provided the selection variable W is still
discrete. In the marginal independence case
(Example 1.1), for example, the problem then becomes one
of mixing a nite number of densities fw(x; y) with
weights (w) such that Pw (w)fw(x; y) factorizes
over x and y. This will not be possible for general
fw so, as in the discrete case, we obtain a testable
constraint on the model.
Proof of Lemma 4.1. The rst part follows from
Theorem 3.2 with l = (dx 1)(dy 1), and the failure
under the additional independence from Lemma 3.4.
Suppose dw = 2, write = p(w = 0). Marginal
independence entails that p(x; y) p(x)p(y) = 0 for each
x; y, so using
p(x; y) =
p(x; y j w) + (1
)p(x; y j w)
(here w is used as an abbreviation for w = 0, and w for
w = 1) leads to a quadratic equation axy 2 + bxy +
cxy = 0 with coe cients
axy
bxy
cxy
(p(x j w)
p(x j w))(p(y j w)
p(y j w))
axy
cxy
      </p>
      <p>
        p(x; yjw) + p(xjw)p(yjw)
p(x; yjw)
We have complete unidenti ability if and only if axy =
bxy = cxy = 0 for all x; y. From the forms of cxy and
bxy it is clear that this occurs only if X ?? Y j W .
Conversely, if X ?? Y j W and X ?? Y then for binary
W this implies that either X ?? W; Y or Y ?? W; X [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
so identi ability fails if and only if X ?? Y j W .
Proof of Proposition 5.3. We prove the case W =
fwg, from which the main result is a fairly easy
extension.
      </p>
      <p>The map : p(xV ; xw) 7! p(xV j xw) maps M into the
d(M) dw +1 dimensional space of hierarchical models
after conditioning on Xw; therefore the image of M
M also lies in a d(M) dw + 1 dimensional space. It
follows that generic bres of when applied only to M
must have dimension at least d(M) d(M) + dw 1,
and in particular have positive dimension if d(M)
d(M) &lt; dw 1; it follows that d(M) d(M) dw 1
is necessary for generic identi ability.</p>
      <p>In particular, if G is decomposable then M = M, so
there is no identi ability for any dw 2.</p>
      <p>A.1</p>
    </sec>
    <sec id="sec-7">
      <title>Generic Identi ability</title>
      <p>Proof of Lemma 3.1. Suppose we nd p(w) such that
p(x) = Pw p(w) p(x j w) is contained in MX . This is
simply a sum over the columns of A, so if it is not of full
column rank then we can clearly add any vector in the
kernel of A to p(w) and obtain the same distribution.
Hence p(w) is not identi able.</p>
      <p>Proposition A.1. Let M be a model for (X; W ) in
which there is a parameter cut between X and W jX,
and such that p(w j x) is unrestricted. The matrix C
with entries cxw = p(x j w) generically has rank r =
min(dx; dw). Further, if we x dw 1 columns of C
and p(x), the remaining column of C is generically
not contained within the span of any r 1 dimensional
vector space.</p>
      <p>Proof. Since p(w j x) is unrestricted, the matrix A
with entries axw = p(w j x) is generically of full rank
min(dx; dw). Now, since cxw = axw p(x)=p(w) is just
a rescaling of the rows and columns of A, the rank of
C is the same.</p>
      <p>For the last part, the column p(x j w) of C is obtained
from A by multiplying it by the non-singular
diagonal matrix with entries p(x)=p(w). Any restriction of
p(x j w) to a subspace would imply a similar restriction
on p(w j x), so by contradiction no such restriction
exists.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bareinboim</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          .
          <article-title>Controlling selection bias in causal inference</article-title>
          .
          <source>In Proceedings of AAAI-12</source>
          , pages
          <fpage>100</fpage>
          {
          <fpage>108</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bareinboim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          .
          <article-title>Recovering from selection bias in causal and statistical inference</article-title>
          .
          <source>In Proceedings of AAAI-14</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>O.</given-names>
            <surname>Barndor -Nielsen</surname>
          </string-name>
          .
          <article-title>Information and exponential families in statistical theory</article-title>
          . John Wiley &amp; Sons,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Berkson</surname>
          </string-name>
          .
          <article-title>Limitations of the application of fourfold table analysis to hospital data</article-title>
          .
          <source>Biometrics Bulletin</source>
          , pages
          <volume>47</volume>
          {
          <fpage>53</fpage>
          ,
          <year>1946</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Copas</surname>
          </string-name>
          .
          <article-title>What works?: selectivity models and metaanalysis</article-title>
          .
          <source>J. Roy. Statist. Soc. Ser. A</source>
          ,
          <volume>162</volume>
          (
          <issue>1</issue>
          ):
          <volume>95</volume>
          {
          <fpage>109</fpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Cox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Little</surname>
          </string-name>
          , and
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>O'Shea. Ideals, varieties, and algorithms: an introduction to computational algebraic geometry and commutative algebra</article-title>
          . Springer,
          <year>2007</year>
          .
          <string-name>
            <given-names>Third</given-names>
            <surname>Edition</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Didelez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kreiner</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Keiding</surname>
          </string-name>
          .
          <article-title>Graphical models for inference under outcome-dependent sampling</article-title>
          .
          <source>Statistical Science</source>
          , pages
          <volume>368</volume>
          {
          <fpage>387</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Drton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sturmfels</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Sullivant</surname>
          </string-name>
          . Lectures on algebraic statistics. Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Edwards</surname>
          </string-name>
          and
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Lauritzen</surname>
          </string-name>
          .
          <article-title>The TM algorithm for maximising a conditional likelihood function</article-title>
          .
          <source>Biometrika</source>
          ,
          <volume>88</volume>
          (
          <issue>4</issue>
          ):
          <volume>961</volume>
          {
          <fpage>972</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Geneletti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Richardson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Best</surname>
          </string-name>
          .
          <article-title>Adjusting for selection bias in retrospective, case{control studies</article-title>
          .
          <source>Biostatistics</source>
          ,
          <volume>10</volume>
          (
          <issue>1</issue>
          ):
          <volume>17</volume>
          {
          <fpage>31</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Greenland</surname>
          </string-name>
          .
          <article-title>Quantifying biases in causal models: classical confounding vs collider-strati cation bias</article-title>
          .
          <source>Epidemiology</source>
          ,
          <volume>14</volume>
          (
          <issue>3</issue>
          ):
          <volume>300</volume>
          {
          <fpage>306</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hernan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hernandez-D az</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Robins</surname>
          </string-name>
          .
          <article-title>A structural approach to selection bias</article-title>
          .
          <source>Epidemiology</source>
          ,
          <volume>15</volume>
          (
          <issue>5</issue>
          ):
          <volume>615</volume>
          {
          <fpage>625</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Lauritzen</surname>
          </string-name>
          .
          <article-title>Generating mixed hierarchical interaction models by selection</article-title>
          .
          <source>Unpublished tech report.</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Pavlides and M. D. Perlman</surname>
          </string-name>
          .
          <article-title>How likely is Simpson's paradox? The American Statistician</article-title>
          ,
          <volume>63</volume>
          (
          <issue>3</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Richardson</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Spirtes</surname>
          </string-name>
          .
          <article-title>Ancestral graph Markov models</article-title>
          .
          <source>Annals of Statistics</source>
          ,
          <volume>30</volume>
          (
          <issue>4</issue>
          ):
          <volume>962</volume>
          {
          <fpage>1030</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Robins</surname>
          </string-name>
          .
          <article-title>A new approach to causal inference in mortality studies with a sustained exposure period| application to control of the healthy worker survivor e ect</article-title>
          .
          <source>Mathematical Modelling</source>
          ,
          <volume>7</volume>
          (
          <issue>9</issue>
          ):
          <volume>1393</volume>
          {
          <fpage>1512</fpage>
          ,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>B.</given-names>
            <surname>Scho</surname>
          </string-name>
          lkopf,
          <string-name>
            <given-names>D.</given-names>
            <surname>Janzing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sgouritsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Mooij</surname>
          </string-name>
          .
          <article-title>Semi-supervised learning in causal</article-title>
          and anticausal settings,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>E.</given-names>
            <surname>Stanghellini</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Wermuth</surname>
          </string-name>
          .
          <article-title>On the identi cation of path analysis models with one hidden variable</article-title>
          .
          <source>Biometrika</source>
          ,
          <volume>92</volume>
          (
          <issue>2</issue>
          ):
          <volume>337</volume>
          {
          <fpage>350</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Verma</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          .
          <article-title>Equivalence and synthesis of causal models</article-title>
          .
          <source>In Proceedings of the 6th Conference on Uncertainty in Arti cial Intelligence (UAI-90)</source>
          , pages
          <fpage>255</fpage>
          {
          <fpage>270</fpage>
          ,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>