<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing Algorithm Performance Understanding through tsMorph: Generating Semi-Synthetic Time Series for Robust Forecasting Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Moisés Santos</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>André de Carvalho</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlos Soares</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fraunhofer AICOS Portugal</institution>
          ,
          <addr-line>Porto</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Mathematical and Computer Sciences, University of São Paulo</institution>
          ,
          <addr-line>São Paulo</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LIACC/Faculdade de Engenharia da Universidade do Porto</institution>
          ,
          <addr-line>Porto</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>When never produced as much data as today, and tomorrow will probably produce even more data. The increase is due not only to the larger number of data sources, but also because the source can continuously produce more recent data. The discovery of temporal patterns in continuously generated data is the main goal in many forecasting tasks, such as the average value of a currency or the average temperature in a city, in the next day. In these tasks, it is assumed that the time diference between two consecutive values produced by the same source is constant, and the sequence of values form a time series. The importance, and the very large number, of time series forecasting tasks make them one of the most popular data analysis application, which has been dealt with by a large number of diferent methods. Despite its popularity, there is a dearth of research aimed at comprehending the conditions under which these methods present high or poor forecasting performances. Empirical studies, although common, are challenged by the limited availability of time series datasets, restricting the extraction of reliable insights. To address this limitation, we present tsMorph, a tool for generating semi-synthetic time series through dataset morphing. tsMorph works by creating a sequence of datasets from two original datasets. The characteristics of the generated datasets progressively depart from those of one of the datasets and a convergence toward the attributes of the other dataset. This method provides a valuable alternative for obtaining substantial datasets. In this paper, we show the benefits of tsMorph by assessing the predictive performance of the Long Short-Term Memory Network and DeepAR forecasting algorithms. The time series used for the experiments come from the NN5 Competition. The experimental results provide important insights. Notably, the performances of the two algorithms improve proportionally with the frequency of the time series. These experiments confirm that tsMorph can be an efective tool for better understanding the behaviour of forecasting algorithms, delivering a pathway to overcoming the limitations posed by empirical studies and enabling more extensive and reliable experiments. Furthermore, tsMorph can promote Responsible Artificial Intelligence by emphasising characteristics of time series where forecasting algorithms may not perform well, thereby highlighting potential limitations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;dataset morphing</kwd>
        <kwd>time series</kwd>
        <kwd>synthetic data</kwd>
        <kwd>performance understanding</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Forecasting is one of the main tasks of a decision-making process [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Quantitative approaches
use historical data, such as time series, to make forecasts [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Time series forecasting is an
important tool in several application domains, such as weather, stock markets, and epidemiology.
Several methods for time series forecasting have been proposed in the literature with high
predictive performance on diverse domains [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, according to Wang et al. (2022) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a
limited amount of work tries to understand under which conditions we can expect a forecasting
method to obtain good (and bad) results. According to Baeza-Yates et al. (2024) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], in the context
of Responsible AI, being aware of the limitations of methodologies is crucial for mitigating
them.
      </p>
      <p>
        Empirical analysis can be used to understand algorithm behavior. However, sources of
realistic datasets for time series analysis benchmarking are limited. Some approaches have been
proposed in the literature for generating synthetic time series with realistic characteristics to
generate benchmarking and improve the performance of metalearning, such as Autoregressive
approaches [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Autoregressive approaches generate time series with specific characteristics
with optimization techniques. Furthermore, generative approaches [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ] based on Generative
Adversarial Networks [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] learn the temporal dynamics of realistic time series for the
generation of synthetic series. These frameworks focus on the realism of the generated time series.
However, generating datasets to improve understanding of the performance of algorithms has
two important challenges: (1) the high computational cost of the related approaches; (2) the
absence of mechanisms for gradual variation of data characteristics that lead to variation of the
behavior of algorithms.
      </p>
      <p>
        The work of Correia et al. (2019) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] proposed a simple approach for systematic dataset
generation called dataset morphing. Dataset morphing consists of gradually transforming a
source dataset into a target dataset. Gradual changes in the behavior of learning algorithms in
the sequence of semi-synthetic datasets aim to obtain a better understanding of these algorithms.
This approach was originally proposed for the evaluation of collaborative filtering algorithms. It
is a method with an intuitive implementation and, depending on the transformation function, it
can have a low computational cost. For example, in Correia et al. (2019) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], a dataset morphing
transformation consisted of random rows/columns switching between two collaborative filtering
binary datasets.
      </p>
      <p>In this study, we present tsMorph, a novel approach that extends existing work by introducing
a dataset morphing technique for generating semi-synthetic time series. Our contributions
include:
• Dataset Morphing for Time Series: We propose a method, tsMorph, which adapts the
dataset morphing approach specifically for time series data.
• Empirical Algorithm Evaluation: We demonstrate the application of tsMorph in
empirically evaluating forecasting algorithms, providing a valuable contribution to the
ifeld. Notably, tsMorph is model-agnostic, allowing seamless integration with various
forecasting algorithms.
• Semi-Synthetic Time Series Generation: The use of tsMorph enables the creation of
semi-synthetic time series with gradual variations, enhancing the versatility of generated
datasets for various applications.</p>
      <p>To validate our methodology, we conducted performance analyses on the Long Short-Term
Memory (LSTM) Neural Network and DeepAR algorithms. We applied tsMorph to the NN5
competition dataset, providing valuable insights into the efectiveness of these algorithms in
time series forecasting. Our findings underscore tsMorph’s ability to generate semi-synthetic
time series data that capture transitions between datasets. This functionality establishes our
approach as a valuable tool for gaining insights into the performance of forecasting algorithms.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>Time series data is characterized by its temporal dependency, where the current value is
influenced by past observations. Forecasting models aim to capture and exploit these dependencies
to make accurate predictions. In this section, we provide an overview of time series forecasting
and performance understanding, highlighting key concepts and techniques in these areas.</p>
      <sec id="sec-2-1">
        <title>2.1. Time series forecasting</title>
        <p>
          A time series is a sequence of observations of a variable of interest equally spaced in time [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
Then we can denote a time series by  = {1, 2, . . . ,  } and  ∈ R are the observations.
The number of observations  is the length of the time series. Such a time series can represent
many phenomena in the real world. For example, the demand for beds for patients in a hospital
during a period or the daily closing price of a stock in the stock exchange.
        </p>
        <p>One of the main tasks in time series analysis is forecasting. Time series forecasting consists
of using the available observations of the variable of interest to extrapolate the time series into
the future. Suppose  are the past observations up to period  of a variable of interest such
that  = {}=0. Given , the goal of the time series forecasting task is to obtain a model that
estimates the value of the time series at time  + ℎ. The estimate is usually represented as
ˆ +ℎ =  +ℎ +  , for simplicity, ˆ.</p>
        <p>
          In the forecasting definition, ˆ are forecasts from , ℎ denotes the number of forecast
observations and is called forecast horizon, and  is the forecast error. The mapping  :  → ˆ is
a forecasting model. Forecasting models assume a strong relationship between the available
observations and the future of a variable of interest [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>
          To estimate the forecasting accuracy of a model, a set of predictions for the period  +ℎ to  +
ℎ + , {ˆ +ℎ, . . . , ˆ +ℎ+ }, they are compared to the observed values { +ℎ, . . . ,  +ℎ+ },
where  ∈ N* . There are several measures to do this. We use the Mean Absolute Scaled Error
(MASE) [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] in this work. MASE is a scale-free error measure. Because it is scale-free, this
measure can be applied to analyze time series from diferent scales.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Understanding the behavior of ML forecasting algorithms</title>
        <p>In the literature, some approaches are used to understand ML algorithms. This section discusses
some that serve as a basis for the proposed approach–starting with dataset morphing and the
meta-knowledge analysis for time series forecasting.</p>
        <p>
          Dataset morphing is the process of generating semi-synthetic data from the gradual
transformation of a source dataset into a target dataset. Given a source dataset (), a target dataset
(), and a transformation  , intermediate semi-synthetic datasets  are obtained by the
dataset morphing process. Correia et al. (2019) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] defines a generic dataset morphing process
for various tasks in Equation 1.
        </p>
        <p>
          (ℎ) : { |0 = (),  = (),  =  (− 1)}, 1 ≤  &lt;  (1)
where (ℎ) is the total set of datasets, and  is the number of transformations. In general
terms, the  transformation refers to any function applied to the data capable of gradually
transforming the source dataset into the target dataset. The dataset morphing process is
proposed and used initially to understand the contrasting performance of a pair of algorithms
on pairs of datasets. Correia et al. (2019) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] analyzed the performance curves obtained from the
(ℎ) datasets. Data characteristics related to performance variation, called meta-features,
were also analyzed.
        </p>
        <p>
          Given a set of datasets  = {0, . . . , }, a meta-feature m can be defined as a function
that maps  :  →  [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Function  returns  values that characterize each dataset in .
Several meta-features have been proposed in the literature for diferent specific tasks. According
to Brazdil et al. (2022) [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], meta-features must have three main properties: performance
discrimination power, computationally not very expensive, and suitable dimensionality to the
amount of data available.
        </p>
        <p>
          Another approach that makes use of meta-features is metalearning. Metalearning is the set
of methods that uses knowledge extracted from learning tasks, algorithms, or task performance
evaluation to improve predictive performance, make it faster, or understand how algorithms
work [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. It is also occasionally used to analyze the obtained meta-knowledge, despite being
typically used for algorithm selection. Here we focus on the use of metalearning.
        </p>
        <p>
          The work of Armstrong et al. (2001) [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] analyzes meta-features in the pioneering work
on metalearning for time series forecasting. The visualizations presented demonstrate the
relationships between meta-features and algorithm selection methods. They serve to assess the
benefits derived from automatic selection compared to human expert-based selection.
        </p>
        <p>
          The approach proposed by Lemke and Gabrys (2010) [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] uses decision trees to extract
metaknowledge. The decision trees were trained using metadata, with the forecasting algorithms
serving as the target variable. After inducing a decision tree, rules were extracted on which
methodology to follow according to the value of the meta-features.
        </p>
        <p>
          The work of Talagala et al. (2018) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] presents an extensive meta-knowledge analysis. The
ifrst analysis consists of a probability matrix of the output of a classification model. The rows
represent the time series, the columns are the forecasting models, and the matrix values are
the probabilities of selecting a model for a time series. The authors extract interpretations
based on the hierarchical clustering of the matrix by columns and the characteristics of the time
series in the rows. The second analysis is based on feature importance, measured in individual
conditional expectation score (ICE). This analysis aims to measure the efect of changing the
value of a single time series feature on the probability output of the algorithm.
        </p>
        <p>
          Instance Space Analysis (ISA), as elucidated by Smith-Miles and Munoz (2023) [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], represents
a significant shift in how we evaluate algorithms in machine learning. ISA constructs an instance
space, mapping all potential test cases onto a two-dimensional plane. This approach uncovers
relationships between the structural properties of these cases and algorithm performance. While
previous research by Spiliotis et al. (2020) [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] applied ISA to assess the suitability of time series
data in forecasting competitions, it primarily focused on this aspect and did not explore broader
algorithmic analysis or the creation of synthetic time series.
        </p>
        <p>The following section better explores related work to our proposal. In this work, the focus
will be on the dataset morphing approach. Dataset morphing is the basis for the development
of the tsMorph method.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Synthetic time series</title>
      <p>The need for datasets is a challenge in many applications. A representative diversity of these
datasets is also necessary to reach the performance discrimination power and obtain more
solid conclusions. In the context of time series analysis, some approaches are available in the
literature to generate synthetic data with realistic characteristics.</p>
      <p>
        The autoregressive approach called GRATIS proposed by [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a method that generates
synthetic time series given the desired values of some meta-features. It is an evolutionary
approach that searches for the parameters of a data generation method that minimizes the
distance between the meta-features of the generated time series and values defined by the user.
The data is generated using a Gaussian Mixture Autoregressive (MAR) model. The applications
proposed for GRATIS were generating representative benchmarks and metadata augmentation.
One important characteristic of this method is that the time series generated may be very
diferent from each other. This occurs because the similarity that guides the search is calculated
in the selected meta-features space and not in the time series space. Additionally, the search
optimizes only some meta-features, which means that the values of other meta-features (i.e.,
other characteristics of the time series) can be very diferent.
      </p>
      <p>
        The paper [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] proposed Time Series Generative Adversarial Networks (TimeGAN), an
adaptation of Generative Adversarial Networks (GAN) for time series. GAN is a framework for the
generation of realistic data, and in TimeGAN the focus is on the preservation of the temporal
dynamics in the generated time series. This is achieved by adding to the classic unsupervised
adversarial loss a term representing the similarity to the original data. The work proposes
improving the performance of prediction, forecasting, and classification tasks as direct
applications.
      </p>
      <p>GRATIS and TimeGAN frameworks are methods for generating synthetic time series that
can be used to understand the behavior of algorithms, which is the goal of our work. Both
are committed to the realism and diversity of the data generated. However, they have some
limitations. First, the time series that is generated may be too diferent from each other. By
analyzing the performance of algorithms on such time series, it is possible to obtain an overall
perspective of their behavior (e.g., on what types of time series algorithm A perform better than
algorithm B). However, it may also be important to have a more detailed characterization of
their behavior (e.g., how does the relative performance of algorithms A and B evolve as the
characteristics of the time series gradually change). The time series generated by autoregressive
and the generative processes support this kind of analysis. Secondly, the computational costs of
tsMorph generates semi-synthetic time series for understanding forecasting algorithms. Let a
source time series  () and a target time series  () of the same length. The proposed
transformation function  for the gradual transition between  () and  () is:
 () =   ·  () + (1 −  ) ·  (), 0 ≤  ≤ 
where  is a contribution coeficient equal to
process and  is the index of transformations between  () and  (). The function  is
a linear transformation whose contribution coeficient

spaces the generated semi-synthetic
time series equally in the range of values between  () and  () during the morphing
process. The computational cost of the  transformation is ( × ). For time series with
length  &gt;&gt; , the computational cost of the transformation  is ( ). The dataset morphing
for time series developed in the tsMorph method is defined by:</p>
      <p>−  1 ,  is the number of time series after morphing
 (ℎ) : {| =  ()}, 1 &lt;  &lt; 
where  (ℎ) is the set of datasets and  is the number of time series in the set  (ℎ).
Therefore,  are the semi-synthetic time series gradually generated by the tsMorph method
including source and target time series. Figure 1 illustrates the application of the tsMorph
method on source and target time series with  = 5.</p>
      <p>TimeGAN training and the GRATIS optimization process are high.
4. Dataset morphing for time series
0
10
20
30
40
50</p>
      <p>Given  () and  () and one forecasting algorithm, It is possible to understand the
behavior of algorithms from the perspective of performance variation and meta-features in the
Source
Target
Morphing 1: 75% Source / 25% Target
Morphing 2: 50% Source / 50% Target
Morphing 3: 25% Source / 75% Target
(2)
(3)
set  (ℎ) using tsMorph. We can also control the metadata augmentation process with the
generation of semi-synthetic data based on realistic  () and  () time series.</p>
      <p>The advantage of the tsMorph method is the simplicity of the method, which can be easily
implemented in any programming language from the formulation. It is a computationally cheap
method, as explained earlier. Moreover, it promotes the gradual transformation between the
source and target time series. This last characteristic is of great interest in understanding the
behavior of forecasting algorithms.</p>
      <p>
        Traditional evaluation methods in machine learning often focus on performance metrics
that summarize algorithm performance across specific datasets, lacking detailed insights into
algorithm behavior [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. We propose the tsMorph method as a solution for empirical algorithm
evaluation in response to this limitation. Unlike conventional approaches, tsMorph aims to
understand algorithm performance within a meta-feature space comprehensively. It is crucial to
emphasize that, during the morphing process, the application domain of the time series becomes
irrelevant, as the primary objective is to identify potential limitations in terms of meta-features
inherent to forecasting algorithms when applied to semi-synthetic time series generated from
real-world data.
      </p>
      <p>
        Our method complements other synthetic time series generation methods, such as GRATIS [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
and TimeGAN [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], rather than replacing them. The main focus will be on investigating whether
it is possible to extract interesting meta-knowledge from the time series of the set  (ℎ)
generated with tsMorph to investigate algorithm bias.
      </p>
    </sec>
    <sec id="sec-4">
      <title>5. Empirical validation</title>
      <sec id="sec-4-1">
        <title>5.1. Experimental setup</title>
        <p>We carried out experiments to illustrate the potential of tsMorph to support a better
understanding of the performance of forecasting algorithms. The repository for reproducing the results of
this study is publicly available 1.</p>
        <p>
          We use the 111 time series from the NN5 [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] Competition. The dataset has years of historical
data from cash machines located in diferent UK regions. The goal of the competition is to
forecast 56 days. The main characteristic of choosing this set of this dataset is that it is a
benchmark known in the literature with a time series of the same size. The data have missing
observations that were linearly interpolated since this work objective is not to evaluate this
efect.
        </p>
        <p>
          The forecasting algorithms analyzed in this study include the LSTM and the DeepAR algorithm.
LSTM, proposed by Hochreiter and Schmidhuber (1997) [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], is a well-established ML algorithm
widely employed for time series forecasting tasks. In addition to LSTM, we also investigate the
DeepAR algorithm, a probabilistic forecasting model with autoregressive recurrent networks,
which represents an advancement beyond LSTM. DeepAR was introduced by Salinas et al.
(2020) [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] and ofers promising capabilities for time series forecasting tasks. These algorithms
were chosen for analysis due to their widespread use and the need for a deeper understanding
of their behavior. Both algorithms are included in the Python package neuralforecast [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
1Repository URL: https://github.com/moisesrsantos/tsmorph_aequitas
        </p>
        <p>To illustrate the usefulness of tsMorph, we selected source and target times series from NN5
as follows. We selected ten source time series and one target time series. The ten source time
series are the ones where the algorithm obtained the best predictive performance, and the
target time series is where the algorithm obtained the worst predictive performance according
to MASE. We paired each of the sources with a target and applied tsMorph to generate the
corresponding semi-synthetic data. The goal is to understand the changes in the properties of
the time series as the predictive performance of the algorithm degrades.</p>
        <p>
          The meta-features used in this work are from the Python package catch22 [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. This is a
package designed to identify 22 canonical meta-features from time series data. These 22 time
series characteristics are carefully chosen from a comprehensive pool of 7000 features within
the hctsa [26]. Termed as canonical features, they serve as a condensed representation of the
larger feature set, with empirical emphasis on optimizing predictive accuracy, computational
eficiency, and interpretability. It should be noted that the meta-features were extracted only
from the training data, and the performance is extracted only from the test data, while the
tsMorph is applied to the complete time series. The meta-feature names used in this work were
the short names 2.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>5.2. Understanding the performance of forecasting algorithms</title>
        <p>We delve into the application of tsMorph, a semi-synthetic data generation technique aimed at
enhancing our understanding of algorithm performance, particularly for LSTM and DeepAR
forecasting algorithms. By leveraging tsMorph, we aim to gain insights into the relationship
between algorithm performance and meta-features through augmented time series data.</p>
        <p>To assess the impact of tsMorph on enhancing our understanding of algorithm performance,
we analyze the correlation between the MASE and meta-features. Specifically, we compute the
mean and standard deviation of the correlation across the three meta-features with the lowest
standard deviation of correlation for each algorithm.</p>
        <p>Table 1 presents the results, featuring two sub-tables corresponding to LSTM (a) and DeepAR
(b), respectively. Within each sub-table, the mean and standard deviation of the MASE correlation
for the selected meta-features are displayed.
2Short names and descriptions:
https://time-series-features.gitbook.io/catch22/feature-descriptions/featureoverview-table</p>
        <p>Tables 1 presents the Pearson correlation analysis between the MASE and selected
metafeatures for LSTM and DeepAR algorithms, respectively. Before we interpret the tables, we
explain each of these meta-features. A brief description of each of them follows:
• Forecast Error: this feature provides a measure of the discrepancy arising from employing
the mean of the preceding 3 values in the time series to forecast the subsequent value.
Time series that are straightforward to predict, signifying instances where the mean of
the 3 preceding time steps serves as an accurate prediction for the current value, will
yield low values for this feature.
• Centroid Frequency: this feature calculates the relative power within the lowest 20% of
frequencies. It assigns high values to time series with significant power in low frequencies
and low values to time series that predominantly exhibit power in higher frequencies.
• Low Frequency Power: this feature calculates the frequency at which the amount of
power in frequencies low and higher is the same. It assigns low values to time series with
a concentration power in the low frequencies and the opposite for high values.
• Whiten Timescale: this feature involves computing the ratio of the first zero-crossing
of the autocorrelation function for the residuals to that of the original time series. This
ratio provides insight into the relative predictability and stability of the data.</p>
        <p>For LSTM, the meta-feature "centroid frequency" exhibits a strong positive correlation, while
"forecast error" and "low frequency power" show strong negative correlations. Conversely, for
DeepAR, similar strong correlations are observed for "centroid frequency" and "forecast error",
with an additional negative correlation observed for "whiten timescale". Comparing the two
algorithms, DeepAR demonstrates slightly stronger correlations overall, particularly with the
"forecast error" meta-feature. This suggests that DeepAR may have a more precise forecasting
capability compared to LSTM. The consistent patterns observed in both tables underscore the
significance of these meta-features to understanding algorithm performance, providing valuable
insights into the behavior of LSTM and DeepAR algorithms in time series forecasting tasks.</p>
        <p>To visually represent the efects of the algorithm performance and meta-feature relationship,
we present two sets of figures, each containing three subfigures corresponding to a specific
algorithm. Figure 2 depicts the tsMorph performance understanding plot to LSTM, with each
subfigure showing the dispersion between meta-feature values on the y-axis, the morphing
process step on the x-axis, and a color bar on the right side representing the relative performance
in terms of Mean Absolute Scaled Error (MASE). Similarly, Figure 3 illustrates the tsMorph
performance understanding plot to DeepAR, with each subfigure providing insights into the
dispersion of meta-feature values, morphing process steps, and relative performance in MASE.</p>
        <p>The analysis of the results reveals that tsMorph was successful in generating semi-synthetic
time series with diversity in both the meta-feature space and performance. This achievement
is a significant outcome of the study, as it demonstrates the versatility and efectiveness of
tsMorph in augmenting time series data for machine learning tasks. Across both LSTM and
DeepAR, a strong negative correlation is observed with the "Forecast Error" meta-feature. This
suggests that both algorithms perform better when discrepancies between predicted and actual
values are minimized, highlighting their sensitivity to prediction accuracy. Furthermore, both
algorithms exhibit a positive correlation with the "Centroid Frequency" meta-feature, indicating
(b) Forecast Error.</p>
        <p>(c) Low Frequency Power.
their efectiveness in capturing long-term trends present in low-frequency components of
the data. However, while LSTM shows a negative correlation with "Low Frequency Power,"
suggesting potential challenges with time series exhibiting significant power in low frequencies,
DeepAR demonstrates a negative correlation with the "Whiten Timescale" meta-feature. This
implies that DeepAR may excel in producing stable and predictable forecasts when the data
exhibits lower autocorrelation in the residuals.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <p>In contrast to the extensive work on developing and testing forecasting algorithms, there is
limited work on understanding their behavior. In this work, we propose a simple method of
generating semi-synthetic time series that can be used for that purpose. The tsMorph method
gradually transforms a source time series into a target time series, generating a sequence of
semi-synthetic time series. The data generated by tsMorph supports an empirical and systematic
(b) Forecast Error.</p>
      <p>(c) Whiten Timescale.
approach to understanding the behavior of algorithms.</p>
      <p>The analysis of correlations between algorithm performance and meta-features revealed
valuable insights into the strengths and weaknesses of the LSTM and DeepAR algorithms. Both
algorithms demonstrated sensitivity to prediction accuracy, as indicated by strong negative
correlations with the "Forecast Error" meta-feature. Additionally, their positive correlations
with the "Centroid Frequency" meta-feature suggest proficiency in capturing long-term trends
present in low-frequency components of the data. However, diferences emerged between the
algorithms regarding their responses to other meta-features, such as "Low Frequency Power"
and "Whiten Timescale".</p>
      <p>The visual representation provided by the tsMorph performance understanding plot ofers a
comprehensive depiction of the relationships outlined in the correlation tables. This visualization
highlights how variations in meta-feature values across morphing steps influence algorithm
performance, providing a clearer understanding of the dynamics between meta-features and
forecasting accuracy. By observing dispersion patterns and color gradients, researchers can gain
deeper insights into the impact of meta-features on algorithm performance and identify trends
that may not be immediately apparent from the correlation tables alone. Thus, the tsMorph
performance understanding plot serves as a valuable tool for elucidating the intricate interplay
between meta-features and algorithm performance, enhancing our comprehension of time series
forecasting algorithms.</p>
      <p>One limitation of the work presented here is that the transformation only applies to time
series of the same size. However, we plan to develop other transformation options that can deal
with time series of varying sizes with time series alignment techniques. Additionally, in the
experiments, we focused on analyzing algorithms individually. However, by choosing the target
and source time series diferently, we can carry out other types of analyses (e.g., comparing
the performance of two algorithms). In fact, tsMorph can be used to define and understand the
borders of the meta-data space that delimit the areas of expertise of each algorithm.</p>
      <p>Finally, the semi-synthetic data generated by tsMorph can also be used as training data for
AutoML and meta-learning approaches, addressing a major limitation of the current work in
the area: the limited amount of meta-data.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was partially funded by grants #2019/10012-2 and #2021/13281-4, São Paulo Research
Foundation (FAPESP) and CNPq. This work was partially funded also by AISym4Med (101095387)
through the Horizon Europe Cluster 1: Health, ConnectedHealth (n.º 46858); Competitiveness
and Internationalisation Operational Programme (POCI) and Lisbon Regional Operational
Programme (LISBOA 2020), under the PORTUGAL 2020 Partnership Agreement, through
the European Regional Development Fund (ERDF); NextGenAI - Center for Responsible AI
(2022-C05i0102-02), supported by IAPMEI; FCT plurianual funding for 2020-2023 of LIACC
(UIDB/00027/2020 UIDP/00027/2020). The computational resources of Google Cloud Platform
were provided by the project CPCA-IAC/AF/594904/2023.
[26] B. D. Fulcher, M. A. Little, N. S. Jones, Highly comparative time-series analysis: the
empirical structure of time series and their methods, Journal of the Royal Society Interface
10 (2013) 20130048.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Hyndman</surname>
          </string-name>
          , G. Athanasopoulos,
          <article-title>Forecasting: principles and practice</article-title>
          ,
          <source>OTexts</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Montgomery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Jennings</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kulahci</surname>
          </string-name>
          ,
          <article-title>Introduction to time series analysis and forecasting</article-title>
          , John Wiley &amp; Sons,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Makridakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Spiliotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Assimakopoulos</surname>
          </string-name>
          ,
          <article-title>Statistical and machine learning forecasting methods: Concerns and ways forward</article-title>
          ,
          <source>PloS one 13</source>
          (
          <year>2018</year>
          )
          <article-title>e0194889</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Hyndman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <article-title>Forecast combinations: an over 50-year review</article-title>
          ,
          <source>arXiv preprint arXiv:2205.04216</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U. M.</given-names>
            <surname>Fayyad</surname>
          </string-name>
          ,
          <article-title>Responsible ai: An urgent mandate</article-title>
          ,
          <source>IEEE Intelligent Systems</source>
          <volume>39</volume>
          (
          <year>2024</year>
          )
          <fpage>12</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Hyndman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Gratis: Generating time series with diverse and controllable characteristics</article-title>
          ,
          <source>Statistical Analysis and Data Mining</source>
          <volume>13</volume>
          (
          <year>2020</year>
          )
          <fpage>354</fpage>
          -
          <lpage>376</lpage>
          . doi:
          <volume>10</volume>
          .1002/ sam.11461.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Hyland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Esteban</surname>
          </string-name>
          , G. Rätsch,
          <article-title>Real-valued (medical) time series generation with recurrent conditional gans</article-title>
          , stat
          <volume>1050</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jarrett</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Van der Schaar</surname>
          </string-name>
          ,
          <article-title>Time-series generative adversarial networks</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>I.</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pouget-Abadie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mirza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Warde-Farley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ozair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Generative adversarial networks</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>63</volume>
          (
          <year>2020</year>
          )
          <fpage>139</fpage>
          -
          <lpage>144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Correia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Soares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jorge</surname>
          </string-name>
          ,
          <article-title>Dataset morphing to analyze the performance of collaborative filtering</article-title>
          ,
          <source>in: Discovery Science: 22nd International Conference, DS</source>
          <year>2019</year>
          , Split, Croatia,
          <source>October 28-30</source>
          ,
          <year>2019</year>
          , Proceedings 22, Springer,
          <year>2019</year>
          , pp.
          <fpage>29</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Hyndman</surname>
          </string-name>
          , et al.,
          <article-title>Another look at forecast-accuracy metrics for intermittent demand</article-title>
          ,
          <source>Foresight: The International Journal of Applied Forecasting</source>
          <volume>4</volume>
          (
          <year>2006</year>
          )
          <fpage>43</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Alcobaça</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Siqueira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rivolli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Oliva</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. C. De Carvalho</surname>
          </string-name>
          ,
          <article-title>Mfe: Towards reproducible meta-feature extraction</article-title>
          ,
          <source>The Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>4503</fpage>
          -
          <lpage>4507</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Brazdil</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. N. van Rijn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Soares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanschoren</surname>
          </string-name>
          , Metalearning: Applications to Automated
          <source>Machine Learning and Data Mining</source>
          , Springer Nature,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hutter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kotthof</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Vanschoren,</surname>
          </string-name>
          <article-title>Automatic machine learning: methods, systems</article-title>
          , challenges,
          <source>Challenges in Machine Learning</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Armstrong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Adya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Collopy</surname>
          </string-name>
          ,
          <article-title>Rule-based forecasting: Using judgment in time-series extrapolation, Principles of forecasting: A handbook for researchers and practitioners (</article-title>
          <year>2001</year>
          )
          <fpage>259</fpage>
          -
          <lpage>282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lemke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gabrys</surname>
          </string-name>
          ,
          <article-title>Meta-learning for time series forecasting and forecast combination</article-title>
          ,
          <source>Neurocomputing</source>
          <volume>73</volume>
          (
          <year>2010</year>
          )
          <fpage>2006</fpage>
          -
          <lpage>2016</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Talagala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Hyndman</surname>
          </string-name>
          , G. Athanasopoulos,
          <article-title>Meta-learning how to forecast time series</article-title>
          ,
          <source>Monash Econometrics and Business Statistics Working Papers</source>
          <volume>6</volume>
          (
          <year>2018</year>
          )
          <fpage>16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>K.</given-names>
            <surname>Smith-Miles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Muñoz</surname>
          </string-name>
          ,
          <article-title>Instance space analysis for algorithm testing: Methodology and software tools</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>E.</given-names>
            <surname>Spiliotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kouloumos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Assimakopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Makridakis</surname>
          </string-name>
          ,
          <article-title>Are forecasting competitions data representative of the reality?</article-title>
          ,
          <source>International Journal of Forecasting</source>
          <volume>36</volume>
          (
          <year>2020</year>
          )
          <fpage>37</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>C. C. McGeoch</surname>
          </string-name>
          ,
          <article-title>A guide to experimental algorithmics</article-title>
          , Cambridge University Press,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Crone</surname>
          </string-name>
          , NN5 forecasting competition,
          <year>2008</year>
          . URL: http://www. neural
          <article-title>-forecasting-competition</article-title>
          .com/NN5/index.htm, accessed on 2023-
          <volume>02</volume>
          -08.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <article-title>Long short-term memory</article-title>
          ,
          <source>Neural computation 9</source>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D.</given-names>
            <surname>Salinas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Flunkert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gasthaus</surname>
          </string-name>
          , T. Januschowski, Deepar:
          <article-title>Probabilistic forecasting with autoregressive recurrent networks</article-title>
          ,
          <source>International journal of forecasting 36</source>
          (
          <year>2020</year>
          )
          <fpage>1181</fpage>
          -
          <lpage>1191</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>K. G.</given-names>
            <surname>Olivares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Challú</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Garza</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Canseco</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Dubrawski,
          <article-title>NeuralForecast: User friendly state-of-the-art neural forecasting models</article-title>
          .,
          <source>PyCon Salt Lake City</source>
          , Utah,
          <string-name>
            <surname>US</surname>
          </string-name>
          <year>2022</year>
          ,
          <year>2022</year>
          . URL: https://github.com/Nixtla/neuralforecast.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Lubba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Sethi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Knaute</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Schultz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Fulcher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Jones</surname>
          </string-name>
          , catch22:
          <article-title>Canonical time-series characteristics: Selected through highly comparative time-series analysis</article-title>
          ,
          <source>Data Mining and Knowledge Discovery</source>
          <volume>33</volume>
          (
          <year>2019</year>
          )
          <fpage>1821</fpage>
          -
          <lpage>1852</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>