<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a sharp estimation of transfer entropy for identifying causality in nancial time series</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alex Seres</string-name>
          <email>alejandro.seres@e-campus.uab.cat</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alejandra Caban~a</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Argimiro Arratia</string-name>
          <email>argimiro@cs.upc.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Autonoma de Barcelona</institution>
          ,
          <addr-line>Mathematics, Barcelona</addr-line>
          ,
          <country country="ES">SPAIN</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universitat Politecnica de Catalunya</institution>
          ,
          <addr-line>Computer Science, Barcelona</addr-line>
          ,
          <country country="ES">SPAIN</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present an improvement of an estimator of causality in nancial time series via transfer entropy, which includes the side information that may a ect the cause-e ect relation in the system, i.e. a conditional information-transfer based causality. We show that for weakly stationary time series the conditional transfer entropy measure is nonnegative and bounded below by the Geweke's measure of Granger causality. We use k-nearest neighbor distances to estimate entropy and approximate the distribution of the estimator with bootstrap techniques. We give examples of the application of the estimator in detecting causal effects in a simulated autoregressive stationary system in three random variables with linear and non-linear couplings; in a system of non stationary variables; and with real nancial data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The determination of cause-e ect relation among nancial time series poses
several research challenges: besides the proper detection of the causality, it is
important to quantify its strength and the e ects of side information that may be
present in the system. Also, considering well known observed stylized facts about
time series, the measure of causality should be sensible to detection of possible
non linear dependence, and adaptable to non stationary time series. In
econometrics the standard tool for testing statistical causality is the one proposed by
Granger [
        <xref ref-type="bibr" rid="ref5 ref6">6, 5</xref>
        ], for the bivariate case, and extended by Geweke [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to a
conditional Granger causality test. These tests assume a linear relationship among
the causes and e ects since it is implemented by tting autoregressive models.
More recently there have been several approaches to testing causality based on
non parametric methods, kernel methods and information theory, among others,
in order to cope with non linearity and non stationarity, e.g., [
        <xref ref-type="bibr" rid="ref1 ref11 ref16 ref17 ref2">2, 17, 11, 1, 16</xref>
        ].
      </p>
      <p>
        The purpose of this work is to contribute to the enhancement of an
estimator of bivariate causality proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], and implemented as software tool
TRENTOOL, which has been shown to be robust in detecting the direction and
measuring the strength of causality in complex biological systems. Our work
and contributions is structured as follows. In Section 2 we present our modi ed
version of Wibral et al. transfer entropy based causality test by extending the
de nition to a conditional causality test, and thus, accounting for side
information. In Section 3 we show that this conditional transfer entropy is a measure of
statistical causality in the same sense as Granger causality (i.e. that the causes
precede the e ects). Section 4 describes the general steps to estimate transfer
entropy. We use the same technique of k-nearest neighbor distances of [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] for
estimating mutual information, but we make use of a bootstrap technique suited
for stationary time series [
        <xref ref-type="bibr" rid="ref10 ref12">12, 10</xref>
        ], for approximating the distribution of the
estimator. As a proof of concept, in Section 5, we report three applications of the
modi ed causality measure. Two with simulated data, one to assess the
sensitivity of the causality test to linear and non linear couplings in a system of
three variables, the other composed of two non stationary variables. The third
experiment is our reality check where we use real nancial data: we try to
determine the possible in uence of a market index to another by testing for causality
among the German DAX index and the Spanish IBEX.
      </p>
      <p>
        Due to space restrictions we omit several technical details in this report.
These can be found in the extended document [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Granger Causality and Transfer Entropy</title>
      <p>
        Let X = fXtg, Y = fYtg, Z = fZtg be three stochastic processes de ned on
a common probability space, from which we try to infer a causal interaction
between X and Y , and where Z represents the side information to complete
the system. We denote the realization of these random variables at time t as
xt, yt, zt. Further we use Xt; Yt and Zt to denote the state space vectors that
characterize the processes at time t, in this case we choose the whole collection
of random variables up to time t, to follow closely the de nition given in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
Conditional Granger causality. It is said that X does not (Granger) cause
Y , relative to the side information, Z, if for all t 2 Z, P (YtjXt 1; Yt k; Zt 1) =
P (YtjYt k; Zt 1); where k 2 N is the lag, and P ( j ) stands for conditional
probability distribution. In the bivariate case (unconditional causality) the side
information is omitted. To determine conditional Granger causality the following
vector regression models are considered
      </p>
      <p>Yt = LY (Yt 1) + LXY (Xt 1) + LZY (Zt 1) + Y;t</p>
      <p>
        Yt = LeY (Yt 1) + LeZY (Zt 1) + eY;t
where LY ; LXY ; LZY , LeY , LeZY are linear functions and Y;t, eY;t are the residuals
of the regression of Yt on Xt and information Zt, and without Xt. Then one
can quantify the usefulness of including X in explaining Y using Geweke's test
based on the variances of the residuals [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]: FX!Y jZ = log var(eY;t) .
var( Y;t)
      </p>
      <p>
        Note that the residual variance of the second regression will always be larger
or equal to the one from the 0. For its statistical treatment,
rst, so FX!Y jZ
it is known that the corresponding maximum likelihood estimator will have a
2 distribution under the null hypothesis FX!Y jZ = 0, and a non-central 2
distribution under the alternative hypothesis FX!Y jZ &gt; 0 [
        <xref ref-type="bibr" rid="ref3 ref5">5, 3</xref>
        ].
Transfer entropy. The transfer entropy of process X to Y , conditioned to the
side information Z, can be obtained from mutual information as
T EX !Y jZ = I(Y +; X jY
      </p>
      <p>
        Z )
where A denotes all information available of the past of A, A+ its immediate
future, and A B represents the concatenation of random vectors A and B.
Wibral et al. in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] argue extensively on the inadequacy of this formula in the
bivariate case for truly capturing cause-e ect relations, and go on to propose
a transfer entropy with self prediction optimality at lag u, which extended to
account for the side information Z, has the form:
      </p>
      <p>T ESP O:X !Y;ujZ = I(Yt; Xt ujYt 1; Zt 1)
Then they show,
Theorem 1. For two discrete-time random processes X,Y with a state-space
representation X; Y coupled from X to Y via a non-zero delay , T ESP O:X !Y;ujZ
is maximal for u = . This also holds in presence of additional coupling from Y
to X.
tu</p>
      <p>This approach suits every requisite that we need to correctly estimate
information transfer, but in practice the associated estimator resulted in an
unreliable predictor due to the amount of information about Xt present in Xt u,
for 1 &lt; u &lt; , which causes it to assign a larger value to the information
\rebounds" produced by the approximations assumed. (By rebound it is meant the
contribution of information due to indirect causality, e.g. X ! Z and Z ! Y .)
This can be corrected by using the state space representation for the target Yt:
T ESP O:X !Y;ujZ = I(Yt; Xt ujYt 1; Zt 1)
(1)
By the properties of the state-space representation we are not losing any
condition that we imposed to the former estimator, and we obtain an advantage: we
are compensating the rebound e ect of Xt in Xt u, because we swap the
numerical value of Yt for a state vector that contains information of the past which,
supposing that the interaction lag is constant for all t considered, accounts for
a more precise value for the unwanted rebounds, with a clear apex at the true
delay. On the other hand we have a disadvantage, related to the method used
to estimate conditioned densities. As a result of using an element of a larger
dimension we get a bigger error in the estimation.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Transfer Entropy is a measure of Causality</title>
      <p>
        Barnett et al. in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] showed that for random variables with a Gaussian
distribution, measuring transfer entropy is equivalent to measuring Granger causality.
This result easily extends to the conditional causality concept. Formally, it can
be shown that FX!Y jZ = 2T EX !Y jZ .
      </p>
      <p>
        Nevertheless, we want to evaluate our capacity to infer causality against
nancial data, which is widely known to behave neither as stationary nor to be
Gaussian. For this motive it is very important to con rm that when measuring
information transfer between processes we are, indeed, testing for statistical
causality. We have not found any explicit proof of this proposition, but there are
some signs in the literature. One is an argument asserting that non-zero Granger
causality means non-zero transfer entropy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], but it is a bit vague. Next is an
indication for a lower bound for mutual information [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which resulted really
helpful in proving that for weakly stationary variables, not necessarily Gaussian,
the linear in uences detected with Granger causality are no more than the ones
detected for Gaussian variables, and therefore we will detect them at the very
least. We state without proof the following proposition,
Proposition 1. Given two stochastic processes fXtg,fYtg, and supposing that
X causes Y in the Granger sense, then T EX !Y jZ 0 tu
      </p>
      <p>Now, conditional transfer entropy relates to Granger causality as follows:
Theorem 2. Let fXtg, fYtg, fZtg be three jointly distributed, weakly stationary
stochastic processes, de ned in a common probability space. Then the standard
measure of Granger causality and transfer entropy are related by:
2T EX !Y jZ</p>
      <p>FX!Y jZ
Proof. It su ces to prove that conditional mutual information is minimized by
Gaussian distributions. Since T EX !Y jZ is only a speci c form of a conditional
mutual information the result will follow.</p>
      <p>
        As proposed in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] we will set a minimization problem on the mutual
information for continuous variables using Lagrange multipliers. If we use (x; y) =
0(x; yjz), then the formula that we need to minimize does not change:
I(X; Y jZ) =
0(x; yjz)log
      </p>
      <p>dxdy
Z Z</p>
      <p>Z Z
=
@L
Now impose @ (x;y) = 0, since the rst and second order moments are nite, we
can swap the partial derivative with the integral sign, thus obtaining:</p>
      <p>+ 1 (x; y) + 2x2 (x; y)) + 3y2 (x; y)
+ 4yx (x; y) dxdy +</p>
      <p>1 + 2 x + 3 y + 4 xy ;
which is zero if the derivative inside the integral is zero, as it is positive. Then:
0 =
log
and hence (x; y) = x(x) y(y)e (1+ 1)e 2x2 3y2 4yx, where 1, 2( 0),
3( 0), 4 are constants xed by the constraints. Therefore a probability
density that follows this equation is consistent with a Gaussian. To prove that it is
the only one, integrate the expression of (x; y) over y and make the change of
variables x = i =c:</p>
      <p>Z
Z</p>
      <p>Z
Z
(x; y) dy =</p>
      <p>x(x) y(y)e (1+ 1)e
X (x) =</p>
      <p>x(x) y(y)e (1+ 1)e
e(1+ 1)e 2x2 =
e(1+ 1)e
2 2=c2 = Z
y(y)e
3y2</p>
      <p>4yx dy )
ei 4y =c[ Y (y)e
3y2 ] dy
2x2
2x2
3y2
3y2
4yx dy )
4yx dy )
The left hand side of this last equality is a Gaussian density function, and the
right hand side, its Fourier transform, Y (y)e 3y2 . Since the Fourier transform
of a Gaussian is another Gaussian, and Y (y) cannot be a constant if it is a
density function, then Y (y) is also Gaussian. The same argument (integrating
over x) shows that X (x) is Gaussian, thus the joint distribution of the random
variables X; Y that minimizes its mutual information conditioned to Z, (x; y) =
0(x; yjz), is a Gaussian distribution. tu
4</p>
    </sec>
    <sec id="sec-4">
      <title>Practical estimation of transfer entropy</title>
      <p>
        In this section we outline our methodology used to estimate transfer entropy. Our
methodology follows the procedures outlined in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], but we introduce some key
modi cations: aside from our particular estimator which deals with conditional
mutual information, we test for null transfer entropy by obtaining a bootstrap
approximation of the distribution of the estimator, a non parametric procedure.
The state-space representation. First we need to reconstruct the
corresponding state-space of the interacting systems from a scalar time series. For this
purpose we use Takens delay embedding [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] with the parameters suggested by
Ragwitz [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Delay embedding states of the systems under consideration can be
written as delay vectors of the form: X(tm) = (Xt; Xt ; Xt 2 ; : : : ; Xt (m 1) );
where m and denote the embedding dimension and Takens' embedding delay
respectively.
      </p>
      <p>Embedding lag. To enhance the predictability between the state variables it
is key to choose the correct variables. Ragwitz' criterion yields delay embedding
states that provide optimal self prediction for a large class of systems. It is based
in a close study of the properties of an assumed multivariate Langevin process
from which the data proceeds, assuming that the univariate series behaves locally
like a Markov chain.</p>
      <p>As for the delay parameter , in the limit of in nite precision of the data and
in nitely long time series, all values are equivalent. In a practical situation,
however, a good choice of the delay is crucial. If is too large, successive elements of
the embedding vector are almost independent and the vectors ll an amorphous
cloud in the state-space. If is too small, successive elements of the embedding
vector are strongly correlated and all vectors are clustered around the
\diagonal" of the state space. We will need to search for meaningful neighbors of these
vectors in the next two steps, and in both extreme cases they are di cult to
obtain. Then, the simplest reasonable estimate of an optimal delay is the rst zero
of the autocorrelation function (k) = (k)= (0), where (k) = Cov(Xt; Xt+k),
k 2 Z. To discern the rst zero, our program performs the test</p>
      <p>H0 : (k) = 0 versus</p>
      <p>
        H1 : (k) 6= 0
Embedding dimension. To select the dimension of the vectors of the
reconstructed state-space we use the false nearest neighbors method [
        <xref ref-type="bibr" rid="ref7 ref9">9, 7</xref>
        ]. This
method compares distances of m-valued state-space vectors to its k-nearest
neighbor (knn) with distances for (m + 1)-valued state-space vectors from the
same data. These distances are really easy to calculate considering that all
vectors have a similar form once we have the estimated delay. Only the knn search
may prove to be a problem if the dimension becomes too big. If the distances
grow too much with the growth in dimension, the neighbors are deemed as false,
and we try again increasing the dimension by one, until no false neighbor is
found. In this way the data itself dictate the embedding dimension, in a way
that it approximates correctly the dimension of the manifold in which the states
of the system reside.
      </p>
      <p>Mutual Information estimation Once we have the parameters involved in
the state space representation of both uni-valued series of data, we apply the
formula (1) to nd the transfer entropy from X to Y:</p>
      <p>T ESP O:X !Y;u = I(Yt; Xt ujYt 1) = I(Yt; Xt u; Yt 1)
I(Yt; Yt 1)
(We are not considering a third variable Zt symbolizing side information to
simplify notation, the inclusion of which would mean four mutual information
terms in the previous formula instead of two.)</p>
      <p>
        To estimate the mutual information I(A; B) between two multivariate time
series, A and B, we use the k-nearest neighbor estimator suggested in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. When
we use this method the only assumption behind is a certain smoothness of the
underlying probability functions, so it ts in as a non-parametric technique. This
estimator has the form
      </p>
      <p>I^(A; B) =
(k)
h (nA + 1) +
(nB + 1)i +
(N )
where N is the length of the series used, nA is the number of points in A whose
pairwise distance is below a certain given threshold (similarly de ne nB), and
(x) is the digamma function: (x) = (x) 1 d d(xx) : The digamma function
satis es the recursion (x + 1) = (x) + 1=x and (1) = C, with C being the
Euler-Mascheroni constant.</p>
      <p>
        Bootstrap sampling Once we have the estimators of the Transfer Entropy
between our time series as a function of the lag, we approximate their distribution
via a bootstrap experiment. Take bootstrap replications of the series and
recompute the estimators as usual. Bootstrapping stationary time series has been
widely studied in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and we follow their procedures. This allows us
to assess (for each lag) whether there is no causation (transfer entropy equal to
zero) by looking at the bootstrap con dence intervals. We think that this is a
more natural way of approximating the distribution of the statistic of the
transfer entropy estimator, than the test conducted in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] which consist in comparing
the observed TE with that of surrogate data obtained by permutations of the
source.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Simulations and applications in nance</title>
      <p>
        We present several experiments that we have performed to assess the e ciency
of our method to estimate transfer entropy, as well an application to real data.
Example 1. Consider an autoregressive stationary system in three variables
with one linear coupling (Y ! Z), and two non-linear ones (X ! Y and X !
Z), proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. We will use di erent lags in each interaction to be able to
clearly di erentiate them, as well as the \rebounds" produced in an indirect
cause, speci cally X!Y = 10, X!Z = 5, Y !Z = 15. The system satis es:
Xt = 3:4Xt 1(1
Yt = 3:4Yt 1(1
Zt = 3:4Zt 1(1
      </p>
      <p>Xt 1)2eXt2 1 + 0:4 1;t
Yt 1)2eYt2 1 + 0:5Xt2 10 + 0:4 2;t</p>
      <p>Zt 1)2eZt2 1 + 0:3Yt 15 + 0:5Xt 5Zt 1 + 0:4 3;t
where 1;t; 2;t; 3;t are Gaussian white noise terms with identity covariance
matrix. To simulate it, we have taken the 30 initial terms of each series as 0, and
begin applying the formula from there for a total of 104 steps, and the last 103
were used to test the transfer entropy estimator, as the arbitrary initial
conditions should have little e ect then. We have performed the geometric time series
bootstrap 60 times for each lag, with average sizes of length 100. We then have
drawn 90% con dence intervals. We have estimated up to lag 30 to save
computation time, as it should be enough to recover the causal interactions. The
results are presented in Figure 1.</p>
      <p>Observe that we have recovered all the true delays as clearly distinct
maximums of the transfer entropy versus the lag plots. Its important to highlight
that the scale in the y-axis, that is the transfer entropy magnitude, is di erent in
each graph, as our estimator depends on the embedding dimension, so each test
is unique and we must look at the error bars generated in the bootstrap process
to look at it with perspective. That being said, variation in TE depending on
extremely di erent embedding dimension reaches a 30% of the larger value, and
in this graph we can make conclusions about the strength of the interactions
as we see di erences of an order of magnitude. That is why, as expected, we
can consider null the Z ! X and Y ! X interactions, as the measured TE
for the lag values statistically di erent to zero is really small. There is a hint of
a pattern, but it should be no more than a residual in uence in the estimator,
motivated by the real in uence in the opposite direction. The Z ! Y interaction
is also fairly small, but really interesting: Z does not in uence directly Y , but X
in uences both, rst Z at u = 5 and Y at u = 10, therefore Zt 5 has important
information on Y via X. Here we have a good example of when we would need
to apply the whole TE formula, conditioning to a known overarching series.</p>
      <p>On the other hand, the autoregressive interaction of all the variables,
represented in the diagonal, has been recovered at a maximum in the u = 1 lag,
with decreasing intensity as the in uence of the past of the series declines and
stabilizing at 0 after. The X ! Y interaction is prominently recovered in a
similar fashion, at u = 10, also showing the decreasing remnants of information still
present at values around the true delay. The linear interaction X ! Y is also
displayed with a maximum at u = 15. And the interaction X ! Z shows two
interesting properties of our estimator, aside of successfully detecting the in
uence at u = 5 with a distinct local maximum. We see how we detect a maximum
at u = 25, which corresponds to the interaction X ! Y ! Z, which means that
we successfully detect indirect interactions, as it displays a strength similar to
Y ! Z. The other important detail is to check how the variances of the indirect
interaction are bigger than those of the direct one.</p>
      <p>We repeat the same analysis using a linear Granger causality estimator,
provided in the R library MSBVAR. We used the regressions up to lag u = 15, as we
did not want the indirect in uences to have enough steps to a ect the test over
the direct interactions that are completed at this lag. The results obtained are
shown in Table 1.</p>
      <p>X−&gt;X</p>
      <p>X−&gt;Y
Y−&gt;X
0 5 10 15 20 25 30</p>
      <p>Lag</p>
      <p>Y−&gt;Y
0 5 10 15 20 25 30</p>
      <p>Lag
Z−&gt;X
0 5 10 15 20 25 30</p>
      <p>Lag</p>
      <p>Z−&gt;Y</p>
      <p>The test recovers swiftly the Y ! Z linear interaction, and also the X ! Y
quadratic interaction, but fails at detecting any interaction in X ! Z, which is
highly non-linear at true lag u = 5. To better check the responsiveness of the
test we have repeated it considering various smaller lag values for X, and we
4
3
1
0
never obtained a favorable p-value. Our TE test shows the various interactions
in the system (direct or indirect), justifying its appropriateness when dealing
with non-linear systems. tu
X−&gt;X</p>
      <p>X−&gt;Y
0
5
10 15 20 25 30</p>
      <p>Lag
0
5</p>
      <p>Example 2. Next we propose a similar (but non stationary) system. Consider
two processes (X ! Y ) with a true delay X!Y = 10 and random variance that
follows an IGARCH(1,1) model:</p>
      <p>Xt = 0:7Xt 1 + 1;t 1;t
Yt = 0:3Yt 1 + 0:5Xt 10Yt 2 + 2;t 2;t
i2;t = 0:2 + 0:9 i2;t 1 + 0:1 i2;t 1
0
5
10 15 20 25 30</p>
      <p>Lag
0
5</p>
      <p>The results are displayed in Figure 2 and the analysis is similar to the previous
one. We recover perfectly the autoregressive tendencies, and we obtain a fairly
clear view of the X ! Y interaction, similar to the X ! Z interaction in
Example 1, but now it takes a larger time to fade to zero, so we can sense the
distortion generated by the variance, although we can still identify the true delay
with precision.</p>
      <p>Example 3. As a real world application of our estimator we take a look at the
relationship between two major European stock indices, the German DAX-30
and the Spanish IBEX-35, between January 1st, 2011 and June 24th, 2016, that
constitute approximately 1200 data points, as we have a daily frequency data
feed. We have tested for transfer entropy the log-returns (log(Pt=Pt 1), where
Pt is the adjusted close price of index at day t) of the DAX to the log-returns of
IBEX, as we suppose that the territorially more in uential stock market a ects
the smaller, especially in the current context where Germany holds a major
portion of Spain's debt. The results are represented in Figure 3. We can see a
strong in uence detected in lag u = 1 that fades quickly in 8 days. It is logical to
think that real stock markets react to their sphere of in uence quickly enough
that daily data does not allow for precise detection of the interaction lag, as
we expect it to be much smaller, but we can de nitely conclude that there is a
causal in uence present.
tu
tu</p>
      <p>DAX−&gt;IBEX
21
.
10
.
08
.</p>
      <p>TE .60</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>A. Caban~a acknowledges support of MINECO projec MTM2015-69493-R.
A. Arratia acknowledges support of MINECO project APCOM
(TIN2014-57226P), and Gen. Cat. SGR2014-890 (MACDA).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barnett</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barnett</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seth</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          :
          <article-title>Granger causality and transfer entropy are equivalent for gaussian variables</article-title>
          .
          <source>Phys. Rev. Lett</source>
          <volume>103</volume>
          ,
          <issue>238701</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Diks</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wolski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Nonlinear granger causality: Guidelines for multivariate analysis</article-title>
          .
          <source>Journal of Applied Econometrics</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Geweke</surname>
          </string-name>
          , J.:
          <article-title>Measurement of linear dependence and feedback between multiple time series</article-title>
          .
          <source>J. J. Amer. Stat. Assoc</source>
          .
          <volume>77</volume>
          ,
          <issue>304</issue>
          {
          <fpage>313</fpage>
          (
          <year>1982</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Gourevitch</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouquin-Jeannes</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faucon</surname>
          </string-name>
          , G.:
          <article-title>Linear and nonlinear causality between signals: methods, examples and neurophysiological applications</article-title>
          .
          <source>Biological Cybernetics</source>
          <volume>95</volume>
          , 349{
          <fpage>369</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Granger</surname>
            ,
            <given-names>C.W.J.</given-names>
          </string-name>
          :
          <article-title>Testing for causality: A personal viewpoint</article-title>
          .
          <source>J. Econ. Dyn. Control</source>
          <volume>2</volume>
          ,
          <issue>329</issue>
          {
          <fpage>352</fpage>
          (
          <year>1980</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Granger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Investigating causal relations by econometric models and crossspectral methods</article-title>
          .
          <source>Econometrica</source>
          <volume>37</volume>
          ,
          <issue>424</issue>
          {
          <fpage>438</fpage>
          (
          <year>1969</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hegger</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kantz</surname>
          </string-name>
          , H.:
          <article-title>Improved false nearest neighbor method to detect determinism in time series data</article-title>
          .
          <source>Physical Review E</source>
          <volume>60</volume>
          (
          <issue>4</issue>
          ),
          <volume>4970</volume>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Karasov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Stogbauer, H.,
          <string-name>
            <surname>Grassberger</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Estimating mutual information</article-title>
          .
          <source>Phy. Rev. E</source>
          <volume>69</volume>
          ,
          <issue>066138</issue>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kennel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brown</surname>
          </string-name>
          , R.,
          <string-name>
            <surname>Abarbanel</surname>
          </string-name>
          , H.:
          <article-title>Determining embedding dimension for phase-space reconstruction using a geometrical construction</article-title>
          .
          <source>Physical Review A</source>
          <volume>45</volume>
          (
          <issue>6</issue>
          ),
          <volume>3403</volume>
          {
          <fpage>3411</fpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kunsch</surname>
            ,
            <given-names>H.R.:</given-names>
          </string-name>
          <article-title>The jackknife and the bootstrap for general stationary observations</article-title>
          .
          <source>Annals of Statistics</source>
          <volume>17</volume>
          ,
          <volume>1217</volume>
          {
          <fpage>1241</fpage>
          (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Marinazzo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pellicoro</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stramaglia</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Kernel method for nonlinear granger causality</article-title>
          .
          <source>Phys. Rev. Lett</source>
          .
          <volume>100</volume>
          ,
          <issue>144103</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Politis</surname>
            ,
            <given-names>D.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romano</surname>
            ,
            <given-names>J.P.:</given-names>
          </string-name>
          <article-title>The stationary bootstrap</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          <volume>89</volume>
          ,
          <issue>1303</issue>
          {
          <fpage>1313</fpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ragwiz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kantz</surname>
          </string-name>
          , H.:
          <article-title>Markov models from data by simple nonlinear time series predictors in delay embedding spaces</article-title>
          .
          <source>Phys. Rev. E</source>
          <volume>65</volume>
          ,
          <issue>056201</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Seres-Cabases</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Causality via transfer entropy</article-title>
          .
          <source>Bachelor Sc. Thesis</source>
          , Universitat Autonoma de Barcelona (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Takens</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Dynamical systems and turbulence</article-title>
          .
          <source>In: Lecture Notes in Mathematics</source>
          , vol.
          <volume>898</volume>
          , Warwick 1980 Symp., pp.
          <volume>366</volume>
          {
          <fpage>381</fpage>
          . Springer (
          <year>1981</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Wibral</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pampu</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Priesemann</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , Siebenhuhner,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Seiwert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Linder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Lizier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Vicente</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          :
          <article-title>Measuring information-transfer delays</article-title>
          .
          <source>PLoS ONE</source>
          <volume>8</volume>
          (
          <issue>2</issue>
          ),
          <year>e55809</year>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Zaremba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aste</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Measures of causality in complex datasets with application to nancial data</article-title>
          .
          <source>Entropy</source>
          <volume>16</volume>
          ,
          <issue>2309</issue>
          {
          <fpage>2349</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>