<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Causality on Longitudinal Data: Stable Speci cation Search in Constrained Structural Equation Modeling</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ridho Rahmadi</string-name>
          <email>r.rahmadi@cs.ru.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Perry Groot</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marianne Heins</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hans Knoop</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tom Heskes</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics</institution>
          ,
          <addr-line>Universitas Islam</addr-line>
          <country country="ID">Indonesia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Expert Centre for Chronic Fatigue, Radboud University Medical Centre</institution>
          ,
          <addr-line>Nijmegen</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Computing and Information Sciences, Radboud University Nijmegen</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Netherlands Institute for Health Services Research</institution>
          ,
          <addr-line>Utrecht</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>Developing causal models from observational longitudinal studies is an important, ubiquitous problem in many disciplines. A disadvantage of current causal discover algorithms, however, is the inherent instability in structure estimation. With nite data samples small changes in the data can lead to completely di erent optimal structures. The present work presents a new causal discovery algorithm for longitudinal data that is robust for nite data samples. We validate our approach on a simulated data set and real-world data on Chronic Fatigue Syndrome patients.</p>
      </abstract>
      <kwd-group>
        <kwd>Longitudinal data</kwd>
        <kwd>Causal modeling</kwd>
        <kwd>Structural equation model</kwd>
        <kwd>Stability selection</kwd>
        <kwd>Multi-objective evolutionary algorithm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Developing causal models from observational longitudinal studies is an
important, ubiquitous problem in many disciplines, which has led to the development
of a variety of causal discovery algorithms in the literature [1{5]. A disadvantage
of current causal discovery algorithms, however, is the inherent instability in
structure learning. With nite data samples small changes in the data can lead
to completely di erent optimal structures, since errors made by the discovery
algorithm may be propagated and lead to further errors [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] we developed
a robust causal discovery algorithm for cross-sectional data. The method
performs structure search over Structural Equation Models (SEMs) by maximizing
model scores in terms of data t and complexity. The present work extends our
causal discovery algorithm to longitudinal data. We describe how longitudinal
causal relationships can be modelled for an arbitrary number of time slices.
Furthermore, we show how a longitudinal causal model can easily be scored using
standard SEM software by data reshaping. The algorithm produces accurate
structure estimates and is shown to be robust for nite samples. We validate our
approach on one simulated longitudinal data set and one real-world longitudinal
data set for Chronic Fatigue Syndrome.
      </p>
      <p>Copyright c 2015 for this paper by its authors. Copying permitted for private and academic
purposes.</p>
    </sec>
    <sec id="sec-2">
      <title>Proposed method</title>
      <p>
        We use a SEM for causal modeling. The general form of the equations is
xi = fi(pai; "i); i = 1; : : : ; n:
(1)
where pai denotes the parents which represent the set of variables considered to
be direct causes of Xi and "i represents errors on account of omitted factors that
are assumed to be mutually independent [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In this study, we focus on causal
models with no reciprocal relationships, and no latent variables. Thus the causal
model can also be represented by a Directed Acyclic Graph (DAG). We score
models using both the chi-square 2 (measuring the data t) and the model
complexity (measuring the number of parameters).
      </p>
      <p>
        We use the method we developed in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to perform exploratory search over
SEM models. Based on the idea of stability selection [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the method subsamples
the data D with size bjDj=2c without replacement and generates Pareto optimal
models for each subset. After that, all Pareto optimal models are transformed
into their corresponding model equivalent classes, called Completed Partially
Directed Acyclic Graph (CPDAG) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. From these CPDAGs we compute the edge
and causal path stability graph, such as Figure 3a, by grouping them according
to model complexity and computing their selection probability, i.e., the number
of occurrences divided by the total number of models for a certain level of model
complexity. Stability selection is then performed by specifying two thresholds,
sel (boundary of selection probability) and bic (boundary of complexity). For
example, setting sel = 0:6 means that all causal relationships with edge
stability or causal path stability (Figure 3) above this threshold are considered
stable. The second threshold bic is used to control over tting. We set bic to
the level of model complexity at which the minimum average Bayesian
Information Criterion (BIC) score is found. For example, bic = 7 means that all
causal relationships with an edge stability or a causal path stability lower than
this threshold (Figure 3) are considered parsimonious. Causal relationships that
intersect with the top-left region are considered both stable and parsimonious
and called relevant, from which we can derive a causal model.
      </p>
      <p>
        The method in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] only handles cross-sectional data. Based on the idea of
\unrolling" the network in Dynamic Bayesian Networks [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], we extended the
method to handle longitudinal data. We model longitudinal causal relationships
with a SEM model consisting of two time slices (Figure 1a) that can be
\unrolled" into a network with an arbitrary number of time slices (Figure 1b). Time
slice ti represents the relationships within a time slice (intra-slice causal
relationships, solid arcs in Figure 1a). Causal relationships between time slices (inter-slice
causal relationships, dashed arcs in Figure 1a) always go forward in time, i.e.,
from time slice ti 1 to time slice ti.
      </p>
      <p>To score our models on longitudinal data we use data reshaping. In the
reshaped data, the rst n data points contain the relations that occur in the
rst two time slices t0 and t1. The next n data points contain the relations
that occur in time slices t1 and t2. The i-th subset of n data points contain
used to generate longitudinal data. It contains four continuous variables
(X1; : : : ; X4) in three di erent time slices t0; : : : ; t2.
the relations in time slices ti 1 and ti. The reshaped data then allows us to use
standard SEM software to compute the scores.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Application to Simulated Data</title>
      <p>For this experiment, we generated a longitudinal data set with 400 instances
from a causal graph as depicted in Figure 1b. The data set consists of three time
slices with four continuous variables for each time slice.1 When we searched over
SEM models we added prior knowledge that variables X1 and X2 do not cause
variable X3 directly. We performed the search over 200 subsets.</p>
      <p>
        As the true model is known, we measure the performance of our method by
means of the Receiver Operating Characteristic (ROC) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for both edges and
causal paths. The threshold
sel is
      </p>
      <p>xed to a value ( sel 2 f0:3; 0:6; 0:8; 0:9g)
while</p>
      <p>bic is varied. We compute the True Positive Rate (TPR) and the False
Positive Rate (FPR) from the CPDAG of the true model. As for an example,
in the case of edge stability, a true positive means that an edge that appears
within the top-left region bounded by sel and
bic also exists in the CPDAG of
the true model. Figure 2 portrays the ROC curves for both edge and causal path
stability. Generally we can see that higher values of sel tend to give better ROC
curves. This suggests that our approach is able to
nd the underlying structure
with high reliability scores. A notable point is that the ROC curves stop at a
TPR and/or FPR value lower than 1. Since some of the edges and paths are
disallowed (i.e., no edges in time-slice ti 1 and no paths from ti to ti 1) some
of the edges and causal paths in the stability graphs end up with a selection
probability of 0 and the result is that the ROC curves cannot reach the upper
right corner with TPR = FPR = 1.</p>
      <p>1Available at http://bit.ly/1L6dBOo
0.3
0.6
0.8
0.9
0
.
1
8
.
0
6
R .0
P
T .4
0
2
.
0
0
.
0
0.3
0.6
0.8
0.9
0.0
0.1
0.2
0.4
0.5
0.6</p>
      <p>0.7
0.3</p>
      <p>
        FPR
(a)
For an application to real-world data, we consider a data set about Chronic
Fatigue Syndrome (CFS) which consists of 183 subjects and ve time slices
with six discrete variables [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The variables are, fatigue severity, the sense
of control over fatigue, focusing on the symptoms, the objective activity of the
patient (oActivity ), the subject's perceived activity (pActivity ), and the physical
functioning. We use Expectation Maximization implemented in SPSS [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] to
impute the missing values. As all of the variables have large scales, e.g., in the
range between 0 to 155, we treat them as continuous variables. We added prior
knowledge that the variable fatigue does not cause any of the other variables
directly. We performed the search over 200 subsets.
      </p>
      <p>Figure 3a shows that nineteen relevant edges were found, consisting of eleven
intra-slice and eight inter-slice relationships which among of these, six are between
the same variables and two are between di erent variables. Figure 3b shows that
thirty-two relevant causal paths were found, consisting of twelve intra-slice and
twenty inter-slice relationships which among of these, six are between the same
variables and fourteen are between di erent variables. For a more intuitive
representation, we combine the stability graphs into a model using the following
procedure. First, the nodes are linked according to the nineteen relevant edges.
Second, edges are oriented according to our background knowledge. Eight of the
inter-slice relationships are oriented from time slice ti 1 to ti and ve of the
intra-slice edges can be oriented since it is known that the variable fatigue does
not directly cause any other variable. Third, the edges are oriented according to
the relevant causal paths, which results in another twenty-eight directed edges.
The inferred model is shown in Figure 4. Each edge is annotated with a
reliability score which is the maximum score obtained in the top-left region of the edge
stability graph.</p>
      <p>pbic
0 2 4 6 8 11 14 17 20 23 26 29 32 35 38 41 44 47 50</p>
      <p>Model complexity
psel
psel
0 2 4 6 8 11 14 17 20 23 26 29 32 35 38 41 44 47 50</p>
      <p>Model complexity</p>
      <p>(b)</p>
      <p>
        From the stability graphs we can see that the most stable causal relations are
the inter-slice relations between the same variables followed by some of the
intraslice causal relations. Almost all of the inter-slice relations between di erent
variables are not considered relevant. A directed edge X ! Y in Figure 4 indicates
that a change in variable X causes a change in variable Y . In the intra-slice causal
relationships, we found that all variables are direct causes for fatigue severity.
We also found all variables, except fatigue, to be direct causes for the perceived
activity. Furthermore, the variable control is a direct cause for both focusing
on the symptoms and physical functioning. Generally the inter-slice
relationships show direct causes between the same variables. In addition, the variables
pActivity and control indicate a stronger direct cause for fatigue severity and
focusing on symptoms, respectively, as they contribute a direct cause in both
time slices. The inferred model is consistent with results reported in the medical
literature [
        <xref ref-type="bibr" rid="ref12 ref14 ref15">12, 14, 15</xref>
        ].
      </p>
      <p>fatigue( i−1)</p>
      <p>1
pActivity( i−1) 1
oActivity( i−1) 1
focusing( i−1) 1
functioning( i−1) 1
control( i−1) 1
0.93
0.64
fatigue( i)</p>
      <p>1
pActivity( i)
0.71
oActivity( i)</p>
      <p>0.91
focusing( i)</p>
      <p>0.99
functioning( i)</p>
      <p>1
control( i)
0.96
1
1
0.61
0.99
1
Causal discovery from longitudinal data is an important, ubiquitous problem
in science. Current causal discovery algorithms, however, have di culty dealing
with the inherent instability in structure estimation. The present work
introduces a new discovery algorithm for longitudinal data that is robust for nite
samples. Experimental results on both arti cial and real-world data sets show
that the method results in reliable structure estimates. Future research will aim
to estimate the size of causal e ects.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>The research leading to these results has received funding from the DGHE
of Indonesia and the European Community's Seventh Framework Programme
(FP7/2007-2013) under grant agreement n 305697.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Riva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bellazzi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Learning temporal probabilistic causal models from longitudinal data</article-title>
          .
          <source>Arti cial Intelligence in Medicine</source>
          <volume>8</volume>
          (
          <issue>3</issue>
          ) (
          <year>1996</year>
          )
          <volume>217</volume>
          {
          <fpage>234</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Parner</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arjas</surname>
          </string-name>
          , E.:
          <article-title>Causal reasoning from longitudinal data</article-title>
          .
          <source>Rolf Nevanlinna Inst</source>
          ., University of Helsinki (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Marsh</surname>
            ,
            <given-names>H.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yeung</surname>
            ,
            <given-names>A.S.:</given-names>
          </string-name>
          <article-title>Causal e ects of academic self-concept on academic achievement: Structural equation models of longitudinal data</article-title>
          .
          <source>Journal of educational psychology 89(1)</source>
          (
          <year>1997</year>
          )
          <volume>41</volume>
          {
          <fpage>54</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Friedman</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Russell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Learning the structure of dynamic probabilistic networks</article-title>
          .
          <source>In: Proceedings of the Fourteenth conference on Uncertainty in arti cial intelligence</source>
          , Morgan Kaufmann Publishers Inc. (
          <year>1998</year>
          )
          <volume>139</volume>
          {
          <fpage>147</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Modelling gene expression data using dynamic Bayesian networks</article-title>
          .
          <source>Technical report</source>
          , Computer Science Division, University of California, Berkeley, CA (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Spirtes</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Introduction to causal inference</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          <volume>11</volume>
          (
          <year>2010</year>
          )
          <volume>1643</volume>
          {
          <fpage>1662</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Rahmadi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heins</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoop</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heskes</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <article-title>The OPTIMISTIC consortium: Causality on cross-sectional data: Stable speci cation search in constrained structural equation modeling</article-title>
          .
          <source>arXiv:1506</source>
          .
          <article-title>05600 [stat</article-title>
          .ML] (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pearl</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Causality: models, reasoning and inference</article-title>
          . Cambridge Univ Press (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Meinshausen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , Buhlmann, P.:
          <article-title>Stability selection</article-title>
          .
          <source>Journal of the Royal Statistical Society: Series B (Statistical Methodology)</source>
          <volume>72</volume>
          (
          <issue>4</issue>
          ) (
          <year>2010</year>
          )
          <volume>417</volume>
          {
          <fpage>473</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Chickering</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          :
          <article-title>Learning equivalence classes of Bayesian-network structures</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          <volume>2</volume>
          (
          <year>2002</year>
          )
          <volume>445</volume>
          {
          <fpage>498</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Fawcett</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>ROC graphs: Notes and practical considerations for researchers</article-title>
          .
          <source>Machine learning 31</source>
          (
          <year>2004</year>
          )
          <volume>1</volume>
          {
          <fpage>38</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Heins</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoop</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burk</surname>
            ,
            <given-names>W.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bleijenberg</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The process of cognitive behaviour therapy for chronic fatigue syndrome: Which changes in perpetuating cognitions and behaviour are related to a reduction in fatigue?</article-title>
          <source>Journal of psychosomatic research 75(3)</source>
          (
          <year>2013</year>
          )
          <volume>235</volume>
          {
          <fpage>241</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>IBM</given-names>
            <surname>Corp</surname>
          </string-name>
          . Armonk, NY: IBM SPSS Statistics for Windows,
          <source>Version</source>
          <volume>19</volume>
          .
          <fpage>0</fpage>
          . (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Vercoulen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swanink</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galama</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fennis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jongen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hommes</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van der Meer</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Bleijenberg</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The persistence of fatigue in chronic fatigue syndrome and multiple sclerosis: development of a model</article-title>
          .
          <source>Journal of psychosomatic research 45(6)</source>
          (
          <year>1998</year>
          )
          <volume>507</volume>
          {
          <fpage>517</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Wiborg</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoop</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>L.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bleijenberg</surname>
          </string-name>
          , G.:
          <article-title>Towards an evidence-based treatment model for cognitive behavioral interventions focusing on chronic fatigue syndrome</article-title>
          .
          <source>Journal of psychosomatic research 72(5)</source>
          (
          <year>2012</year>
          )
          <volume>399</volume>
          {
          <fpage>404</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>