<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Neural Network for Synthetic Time Series Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Gregnanin</string-name>
          <email>marco.gregnanin@imtlucca.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johannes De Smedt</string-name>
          <email>johannes.desmedt@kuleuven.be</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgio Gnecco</string-name>
          <email>giorgio.gnecco@imtlucca.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maurizio Parton</string-name>
          <email>maurizio.parton@unich.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Graph Neural Networks, Signature Transform, Synthetic Time Series</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Economic Studies, University of Chieti-Pescara</institution>
          ,
          <addr-line>Viale Pindaro 42, Pescara, 65127, Abruzzo</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Laboratory for the Analysis of compleX Economic Systems (AXES), IMT School for Advanced Studies Lucca</institution>
          ,
          <addr-line>Piazza S. Ponziano, 6</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Lucca</institution>
          ,
          <addr-line>55100, Tuscany</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Research Centre for Information Systems Engineering (LIRIS), KU Leuven</institution>
          ,
          <addr-line>Naamsestraat 69, Leuven, 3000, Flemish Region</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Generating synthetic data for financial time series poses challenges, especially taking into account their nonstationary nature. In this work, we introduce the Sig-Graph Generative Adversarial Network (GAN) model, which integrates the following three components: the time series signature, ofering a structured summary of temporal evolution of a times series; a Long Short-Term Memory (LSTM) network, capturing its inherent autoregressive structure; and Graph Neural Networks (GNNs), leveraging geometric patterns within the time series data. Numerical evaluation demonstrates that the Sig-Graph GAN model outperforms several baseline models in replicating the distribution of logarithmic returns over the Standard and Poor's 500 stock exchanges.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Across various domains including, among others, economics and finance, the necessity arises for the
generation of synthetic data. This need is driven by several factors, such as the scarcity of original
data due to privacy concerns and the requirement for data diversity to enhance model generalization.
However, the generation of synthetic time series data poses a considerable challenge due to their
stochastic nature. This holds especially true for the financial case. We show that geometric patterns
play an important role in addressing the generation of synthetic data for a time series, based on a
Generative Adversarial Network (GAN) model. Specifically, transforming time series from a Euclidean
to a non-Euclidean space using a graph-based approach can significantly enhance our understanding
and analysis of complex financial time series behavior. Furthermore, the adoption of a graph-based
representation of a time series liberates one from the constraint of assuming stationarity within the
time series, and facilitates the analysis of geometric patterns (in the now graph-based representation of
the time series) through Graph Neural Networks (GNNs) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Furthermore, we exploit a Long
ShortTerm Memory (LSTM) network to deal with temporal and long-term patterns. Finally, we explore the
application of the time series signature [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which is a concept derived from path theory. The signature
can be viewed as analogous to the Moment Generating Function (MGF), which is useful for comparing
random variable distributions as it encodes all distribution moments into a single function.
∗Corresponding author.
      </p>
      <p>CEUR</p>
      <p>ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>2. Problem Formulation</title>
      <p>′ =  ( , )</p>
      <p>.</p>
      <p>Consider a univariate time series  = { 1,  2, … ,   } extending up to time  . The objective of this research
is to determine a function  (⋅) , given a random variable  and a graph-based representation  of the
original time series  , that is able to generate synthetic data  ′. The synthetic data should closely
resemble the statistical characteristics, temporal dependencies, and geometric patterns observed in the
original time series  . Therefore, we want to find a function  (⋅) such that { 1, … ,   } ≃ { 1′, …  ′}, where</p>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed Sig-Graph GAN Framework</title>
      <p>
        The initialization involves the construction of the time series   = { − , … ,   } observed in  ̃=  +
1 instants, and the random noise matrix   ∈ ℝ ×̃  , where  is the dimension of the noise vector.
Subsequently, the corresponding graph   associated with the time series is formed utilizing the
visibility graph algorithm [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The choice between an undirected or directed graph is considered a
hyperparameter to optimize. Regardless, the adjacency matrix   maintains of dimension  ×̃  ,̃ with
each node corresponding to a specific time observation. The discriminator/generator functions of the
GAN model are structured, respectively, as: Dis( 
, , 
 ), and Gen(  , , 
 ). Here,  and  represent
vectors of learnable parameters for the discriminator and the generator, respectively. We adopt an
identical network configuration for both the generator and the discriminator, and the network structure
combines GNN, LSTM, and Fully Connected (FC) layers.
output.
      </p>
      <p>
        Recurrent Block. The recurrent block processes an input denoted as   for the generator and   for
the discriminator agent. For simplification, we represent this input as   ∈ ℝ ×̃  , where  is set to 1 for
the discriminator. The input goes through LSTM layers, serving to capture temporal and long-term
patterns. Subsequently, a fully connected layer concludes the recurrent block, producing  ̂ 1 ∈ ℝ ×̃  as
Geometric Block. The geometric block accepts input   ∈ ℝ ×̃  along with the adjacency matrix
  ∈ ℝ ×̃  ̃. In our model, the Graph Convolution Network (GCN) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is used as GNN for analyzing
the geometric patterns within the time series data. Subsequently, an LSTM layer processes the data,
facilitating the treatment of temporal and long-term patterns uncovered by the GNN layers. Then, after
the application of a fully connected layer, the geometric block yields an output  ̂ 2 ∈ ℝ ×̃  .
Linear Block. Within the linear block, the outputs from both the recurrent and geometric blocks,
 ̂ 1 ∈ ℝ ×̃  and  ̂ 2 ∈ ℝ ×̃  respectively, are processed. The initial step involves summing the two
outputs, followed by concatenation with the initial input   ∈ ℝ ×̃  . Subsequently, three fully connected
layers with unit counts of 128, 64, and 1 are applied. The final output of each network is denoted as  ̂∗,
wherein ∗ is replaced with “real” for the discriminator and “fake” for the generator.
3.1. Loss Function
Prior to computing the loss function, we subject  ̂∗ to the lead-lag transformation, denoted with
(⋅) . To this end, we apply the truncated signature with a truncation level degree set to 5. For more
comprehensive insights, we also perform a cumulative summation on  ̂∗, subsequently leading to the
computation of a cumulative truncated signature [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The custom loss function adopted is the Mean
as    (⋅)(both having the same dimension  ), the loss function is defined as follows:
Squared Error (MSE). Denoting the truncated signature as   (⋅)and the cumulative truncated signature
MSE( ̂f,  ̂r) =
      </p>
      <p>∑ (  ((  ̂f)) −   ((  ̂r)) )
+
∑ (   ((  ̂f)) −    ((  ̂r)) ) ,


1
1
 =1
 =1
2
2
with “f ” and “r ” denoting respectively the “fake” and “real” data.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Evaluation</title>
      <p>
        We consider, as baseline models, the Quant GAN model, the GARCH(1, 1)model [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and a Monte-Carlo
simulation for the Black and Scholes model.
      </p>
      <p>Dataset and Pre-Processing. For the scope of our analysis, we selected the Standard &amp; Poor’s 500
(S&amp;P 500) stock exchanges. We collected the closing prices of these stock exchanges spanning the
interval from January 4, 2010, to December 30, 2019. Each dataset consists of about 2515 observations.
Before subjecting the dataset to normalization to achieve a mean of zero and a variance of one, a
preliminary step involved computing logarithmic returns denoted as   = log(  ) − log( −1 ). We chose
 =̃ 100 , and  = 4 for the generator.</p>
      <p>
        Evaluation Metrics. To facilitate meaningful comparisons among diferent models, we employ
various evaluation metrics. Notably, we utilize the leverage efect score as defined in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and a
distribution-based metric. In particular, we consider the Earth Mover’s Distance (EMD), also known
as the Wasserstein 1-distance [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This metric quantifies the minimal cost required to transform the
distribution of real data into that of generated data. Then, we compute the Root Mean Squared Error
(RMSE) between the signatures of real and generated data.
      </p>
      <p>Numerical Results. Data generation is conducted at diferent temporal intervals: daily, weekly,
monthly, and long-term, corresponding to 1, 5, 20, and 100 days, respectively. Results for both real and
generated data for the various datasets are presented in Table 1. Optimal results are highlighted in bold.
Our proposed model consistently outperforms the baseline models in terms of the EMD and leverage
efect metric.</p>
      <p>Evaluation metric
EMD(1)
EMD(5)
EMD(20)
EMD(100)
Sig-RMSE(1)
Sig-RMSE(5)
Sig-RMSE(20)
Sig-RMSE(100)
Leverage Efect</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>We introduce a novel approach that combines GNN, LSTM networks, and the Signature transformation
to construct a GAN model for the generation of synthetic stock log-returns. Our methodology leverages
the inherent geometric patterns present within the time series data. We demonstrate that our proposed
model consistently surpasses baseline models. Future works could involve extending the Sig-Graph
GAN model to tackle other time series generation challenges (also in contexts diferent from finance),
as well as assessing its potential in enhancing the performance of trading strategies based on synthetic
data.
Marco Gregnanin and Giorgio Gnecco were partially supported by the PRIN PNRR 2022 project “MOTUS”
(CUP: D53D23017470001), funded by the European Union – Next Generation EU program.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Scarselli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Tsoi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hagenbuchner</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Monfardini,</surname>
          </string-name>
          <article-title>The graph neural network model</article-title>
          ,
          <source>IEEE Transactions on Neural Networks</source>
          <volume>20</volume>
          (
          <year>2008</year>
          )
          <fpage>61</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T. J.</given-names>
            <surname>Lyons</surname>
          </string-name>
          ,
          <article-title>Diferential equations driven by rough signals</article-title>
          ,
          <source>Revista Matemática Iberoamericana</source>
          <volume>14</volume>
          (
          <year>1998</year>
          )
          <fpage>215</fpage>
          -
          <lpage>310</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lacasa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Luque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Nuno</surname>
          </string-name>
          ,
          <article-title>From time series to complex networks: The visibility graph</article-title>
          ,
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>105</volume>
          (
          <year>2008</year>
          )
          <fpage>4972</fpage>
          -
          <lpage>4975</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T. N.</given-names>
            <surname>Kipf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Welling</surname>
          </string-name>
          ,
          <article-title>Semi-supervised classification with graph convolutional networks</article-title>
          ,
          <year>2016</year>
          . URL: https://arxiv.org/pdf/1609.02907. arXiv:
          <volume>1609</volume>
          .
          <fpage>02907</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>I.</given-names>
            <surname>Chevyrev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kormilitzin</surname>
          </string-name>
          ,
          <article-title>A primer on the signature method in machine learning</article-title>
          ,
          <year>2016</year>
          . URL: https://arxiv.org/pdf/1603.03788. arXiv:
          <volume>1603</volume>
          .
          <fpage>03788</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bollerslev</surname>
          </string-name>
          ,
          <article-title>Generalized autoregressive conditional heteroskedasticity</article-title>
          ,
          <source>Journal of Econometrics</source>
          <volume>31</volume>
          (
          <year>1986</year>
          )
          <fpage>307</fpage>
          -
          <lpage>327</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Knobloch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Korn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kretschmer</surname>
          </string-name>
          , Quant GANs:
          <article-title>Deep generation of financial time series</article-title>
          ,
          <source>Quantitative Finance</source>
          <volume>20</volume>
          (
          <year>2020</year>
          )
          <fpage>1419</fpage>
          -
          <lpage>1440</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rubner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tomasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Guibas</surname>
          </string-name>
          ,
          <article-title>The earth mover's distance as a metric for image retrieval</article-title>
          ,
          <source>International Journal of Computer Vision</source>
          <volume>40</volume>
          (
          <year>2000</year>
          )
          <fpage>99</fpage>
          -
          <lpage>121</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>