<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SpoT-Mamba: Learning Long-Range Dependency on Spatio-Temporal Graphs with Selective State Spaces</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jinhyeok Choi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heehyeon Kim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minhyeong An</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joyce Jiyoung Whang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computing, KAIST</institution>
          ,
          <addr-line>Daejeon</addr-line>
          ,
          <country>Republic of Korea</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Spatio-temporal graph (STG) forecasting is a critical task with extensive applications in the real world, including trafic and weather forecasting. Although several recent methods have been proposed to model complex dynamics in STGs, addressing long-range spatiotemporal dependencies remains a significant challenge, leading to limited performance gains. Inspired by a recently proposed state space model named Mamba, which has shown remarkable capability of capturing long-range dependency, we propose a new STG forecasting framework named SpoT-Mamba. SpoT-Mamba generates node embeddings by scanning various node-specific walk sequences. Based on the node embeddings, it conducts temporal scans to capture long-range spatio-temporal dependencies. Experimental results on the real-world trafic forecasting dataset demonstrate the efectiveness of SpoT-Mamba.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Spatio-Temporal Graphs</kwd>
        <kwd>Trafic Forecasting</kwd>
        <kwd>Selective State Spaces</kwd>
        <kwd>Random Walks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The predictive learning methods on time series play a
crucial role in diverse applications, such as trafic and weather
forecasting. The intricate relationships and dynamic
nature of time series are often represented as graphs,
specifically spatio-temporal graphs (STGs), where node attributes
evolve over time. Recently, spatio-temporal graph
neural networks (STGNNs) have emerged as a powerful tool
for capturing both spatial and temporal dependencies in
STGs [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. Many of those methods employ graph
neural networks (GNNs) to exploit spatial dependencies
inherent in the graph structures, integrating them with
recurrent units or convolutions to capture temporal
dependencies [
        <xref ref-type="bibr" rid="ref2 ref4 ref5 ref6 ref7 ref8 ref9">2, 4, 5, 6, 7, 8, 9</xref>
        ]. These approaches have facilitated
the capturing of spatio-temporal dependencies within STGs.
Despite their remarkable performance in predictive
learning tasks, they often face challenges in handling long-range
temporal dependencies among diferent time steps [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ].
      </p>
      <p>
        STGs often exhibit repetitive patterns over both short and
long periods, which is critical for precise predictions [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ].
Therefore, several methods have adopted self-attention
mechanisms of transformer layers [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] rather than recurrent
units to enhance their capability in exploiting global
temporal information [
        <xref ref-type="bibr" rid="ref10 ref11 ref15">10, 11, 15</xref>
        ]. However, the significant
computational overhead and complexity of attention mechanisms
are being highlighted as major concerns [
        <xref ref-type="bibr" rid="ref10 ref12 ref16 ref17">10, 12, 16, 17</xref>
        ].
      </p>
      <p>
        Meanwhile, structured state space sequence (S4) models
have emerged as a promising approach for sequence
modeling with linear scaling in sequence length [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Those
models take the advantages of recurrent neural networks
and convolutional neural networks, enabling them to
handle long-range dependencies without relying on attention.
However, due to their inability to select information
depending on input data, they have shown limited performance.
      </p>
      <p>
        A recent study has introduced a new S4 model
overcoming the issue, named Mamba, which introduces a selection
mechanism to filter information in an input-dependent
manner [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Mamba has demonstrated notable performance
over transformers across various types of sequence data,
including language, audio, and genomics. In addition, there
have been several studies towards replacing the transformer
with Mamba in graph transformer frameworks [
        <xref ref-type="bibr" rid="ref20 ref21 ref22">20, 21, 22</xref>
        ].
      </p>
      <p>
        In this paper, our focus lies on the predictive learning task
on STGs, specifically STG forecasting. For STG forecasting,
it is vital to capture the evolving behavior of individual
nodes over time and how these changes propagate
throughout the entire graph. Furthermore, leveraging these
dynamics over long spatial and temporal ranges plays a crucial
role in dealing with the intricate spatio-temporal
correlations in STGs [
        <xref ref-type="bibr" rid="ref1 ref23 ref3">1, 3, 23</xref>
        ]. Building upon these insights and
recent advances, we introduce SpoT-Mamba, a new
SpatioTemporal graph forecasting framework with a
Mambabased sequence modeling architecture. With Mamba blocks,
SpoT-Mamba extracts structural information of each node by
scanning multi-way walk sequences and efectively captures
long-range temporal dependencies with temporal scans.
Experiments on the real-world dataset demonstrate that
SpoTMamba achieves promising performance in STG forecasting.
The oficial implementations of SpoT-Mamba are available
at https://github.com/bdi-lab/SpoT-Mamba.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Preliminaries</title>
      <p>
        State Space Model (SSM) The state space model (SSM)
assumes that dynamic systems can be represented by their
states at time step  [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. SSM defines the evolution of
a dynamic system’s state with two equations: h′() =
Ah() + B() and () = Ch() + D(), where
h() ∈ R denotes the latent state, () ∈ R represents
the input signal, ∈(R) ∈× R1, dCenotes the output signal, and
A ∈ R× , B ∈ R× , and D ∈ R are
learnable parameters. SSM learns how to transform the
input signal () into the latent state h(), which is used to
model the system dynamics and predict its output ().
Discretized SSM To adapt SSM for discrete input
sequences instead of continuous signals, discretization is
applied with a step size ∆ . The discretized SSM is defined in a
recurrent form: h = Ah− 1 + B and  = Ch, where
A and B are approximated learnable parameters using a
bilinear method with a step size ∆ [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The term D is
omitted from the equations as it can be considered as a skip
connection. This formulation allows for capturing temporal
dependencies eficiently, resulting in similar computations
to those in recurrent neural networks.
      </p>
      <p>Meanwhile, due to its linear time-invariant (LTI) property,
SSM can be reformulated as discrete convolution: y = K* x
and K ∈ R = (CB, CAB, . . . , CA− 1B), where x ∈
R denotes the input sequence, y ∈ R denotes the output
sequence, * indicates the convolution operation, and  is
the sequence length. This representation facilitates parallel
training for SSM, thereby enhancing training eficiency.</p>
      <p>
        The recurrent and convolutional representations of SSM
for sequence modeling enable parallel training and linear
scaling in sequence length. To further enhance the
computational complexity of SSM, the structured state space
sequence (S4) models have been proposed [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. S4 models
address the fundamental bottleneck of SSM, which involves
repeated matrix multiplications, by employing a low-rank
correction to stably diagonalize the transition matrix A.
Mamba S4 models have demonstrated remarkable
performance in handling long-range dependencies in
continuous signal data, such as audio and time series. However,
S4 models struggle with efectively handling discrete and
information-dense data such as text [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. This limitation
arises from the LTI property inherent in the convolutional
form of SSMs. While the LTI property enables linear time
sequence modeling for S4 models, it requires that the learnable
matrices A, B, and C, as well as the step size ∆ , remain
unchanged across all time steps. Consequently, S4 models
cannot selectively recall previous tokens or combine the
current token, treating each token in the input sequence
uniformly. In contrast, Transformers dynamically adjust
attention scores based on the input sequence, allowing them
to efectively focus on diferent parts of the sequence [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        To address both the lack of selectivity in S4 models and
the eficiency bottleneck in sequence modeling, a recent
study introduced a new S4 model called Mamba, which
removes the LTI constraints [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Mamba incorporates a
selection mechanism that allows its learnable parameters to
dynamically interact with the input sequence. This
mechanism is achieved by modifying the learnable parameters
B and C, as well as the step size ∆ , to functions of the
input sequence. Therefore, Mamba can selectively recall or
ignore information in an input-dependent manner, while
maintaining linear scalability in sequence length.
      </p>
      <p>
        Inspired by the recent advancements in Mamba, we
propose a Mamba-based sequence modeling architecture for
predictive learning tasks on STGs. Our approach employs
a selective mechanism to handle the dynamical changes
in STGs, capturing long-range spatio-temporal
dependencies. In addition, this allows for addressing computational
ineficiencies in transformer-based STGNNs [
        <xref ref-type="bibr" rid="ref10 ref11 ref15">10, 11, 15</xref>
        ].
Spatio-Temporal Graph Forecasting A spatio-temporal
graph (STG) is defined as  = (, ℰ ,  ), where  is
a set of  nodes, ℰ ⊂  ×  is a set of edges,  =
[X1, . . . , X ] is a sequence of observed data for all nodes
at each historical time step, and  is a length of the
sequence. Here, X ∈ R× in denotes the observed data
at time step , where in denotes the dimension of the
input node attributes. STG forecasting aims to predict future
observations for  ′ time steps, given historical
observations for the previous  time steps. This is formulated as
(· )
[X−  +1, . . . , X] →−− [X+1, . . . , X+ ′ ], where  (· )
represents the STG forecasting model.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Spatio-Temporal Graph</title>
    </sec>
    <sec id="sec-4">
      <title>Forecasting with SpoT-Mamba</title>
      <p>We propose SpoT-Mamba (Figure 1), which captures
spatial dependencies from node-specific walk sequences and
learns temporal dependencies across time steps leveraging
Mamba-based sequence modeling. By utilizing the selection
mechanisms of Mamba blocks, SpoT-Mamba can selectively
propagate or forget information in an input-dependent
manner on both temporal and spatial domain.</p>
      <p>
        Multi-way Walk Sequence In STGs, the temporal
sequences for nodes are naturally defined by the time-series
data. On the other hand, since the topological structure
does not have a specific order, a tailored method is required
to define the spatial sequences of nodes in graphs. Hence,
we employ three well-known walk algorithms: depth-first
search (DFS), breadth-first search (BFS), and random walks
(RW), to extract diverse local and global structural
information from each node’s neighborhood. The walk sequences
of length  for node  using these walk algorithms are
defined as   (),   (), and  (), respectively.
These node-specific walk sequences are extracted  times
to exploit more comprehensive structural information.
Walk Sequence Embedding SpoT-Mamba generates
embeddings for node-specific walk sequences by scanning each
sequence. Here, SpoT-Mamba performs bi-directional scans
through Mamba blocks, which makes the model robust to
permutations and captures the long-range spatial
dependency of the sequence more efectively [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Then,
SpoTMamba aggregates the representations of node-specific walk
sequences with pointwise convolution. This allows for
incorporating representations of neighboring nodes to generate
walk sequence embedding for each type of walk sequence.
      </p>
      <p>
        Subsequently, SpoT-Mamba integrates the walk sequence
embeddings into a node embedding w ∈ R using
MultiLayer Perceptron (MLP), where  represents the node index
and  denotes the embedding dimension. Rather than
simply stacking GNN layers, we employ Mamba-based sequence
modeling to generate node embeddings from diverse types
of walk sequences. Therefore, our approach efectively
captures local and long-range dependencies within the graph
by scanning the neighborhood structure of each node.
Temporal Scan with Mamba Blocks We adopt the
learnable day-of-week and timestamps-of-day embeddings to
capture the repetitive patterns over both short and long
periods in STGs [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. SpoT-Mamba concatenates these
embeddings with the node embedding for each time step to
efectively model temporal dynamics from historical
observations. Then, it performs selective recurrent scans with
Mamba blocks across the sequence of embeddings along
the time axis, observing changes in temporal dynamics over
time. This process helps identify critical portions of the
sequence for forecasting and captures periodic patterns in
both short- and long-term intervals. Consequently, the
results efectively encompass spatio-temporal dependencies,
thereby enriching the predictive capabilities.
      </p>
      <p>STG forecasting of SpoT-Mamba Finally, SpoT-Mamba
enhances the representations of nodes scanned along the
temporal axis by incorporating global information from the
entire graph at each time step through transformer layers.
  
  
  

f
o
r
m
e
r</p>
      <p>R
e
g
r
e
s
s
i
o
n
L
a
y
e
r</p>
      <p>∈ ℝ ′× × out</p>
      <sec id="sec-4-1">
        <title>Predicted</title>
      </sec>
      <sec id="sec-4-2">
        <title>Time Series</title>
        <p>′
time</p>
      </sec>
      <sec id="sec-4-3">
        <title>Periodicity and Feature</title>
      </sec>
      <sec id="sec-4-4">
        <title>Embedding</title>
      </sec>
      <sec id="sec-4-5">
        <title>Temporal Scan</title>
        <p>with Mamba Blocks</p>
      </sec>
      <sec id="sec-4-6">
        <title>Spatial Self-Attention and Regression</title>
        <p>
          and Z′ ∈ R× 4 denotes one of the outcomes from the temporal scan, corresponding to the time step ′.
represents the overall procedure of STG forecasting. W ∈ R×  denotes the node embeddings for all nodes in the graph
Then, MLP is applied to forecast the attributes of each node
for future time steps. To accurately predict the temporal
trajectory while ensuring robustness to outliers that deviate
significantly from the expected trajectory, we train
SpoTMamba utilizing the Huber loss, which is less sensitive to
outliers while maintaining the smoothness of the squared
error for small errors [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Experiments</title>
      <p>We compare the performance of SpoT-Mamba with
state-ofthe-art baselines. Additionally, we conduct ablation studies
to further demonstrate the efectiveness of SpoT-Mamba.</p>
      <sec id="sec-5-1">
        <title>4.1. Dataset and Experimental Setup</title>
        <p>Dataset</p>
        <p>
          We evaluate SpoT-Mamba on PEMS04 [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], a
realworld trafic flow forecasting benchmark, following the
experimental setup in [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. PEMS04 contains highway trafic
lfow data collected from the California Department of
Transportation’s Performance Measurement System (PEMS). The
nodes represent the sensors, and edges are created when
two sensors are on the same road. Trafic data in
PEMS04 is
collected every 5 minutes. We set the input and prediction
intervals to 1 hour, corresponding to  =  ′ = 12. The
statistic of PEMS04 is shown in Table 1. In the experiments,
PEMS04 is divided into training, validation, and test sets in
a 6:2:2 ratio in temporal order.
        </p>
        <p>Baselines</p>
        <p>
          We compare SpoT-Mamba with several
baselines using various methods, including GNNs and
Transformers. For STGNNs, we consider GWNet [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], DCRNN [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ],
AGCRN [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], GTS [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], and MTGNN [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. For
attentionbased methods, we include STAEformer [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], GMAN [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ],
and PDformer [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]. Additionally, other methods such as
HI [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], STNorm [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ], and STID [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] are also considered.
Evaluation Metrics
        </p>
        <p>We use three standard metrics for
trafic flow prediction: Mean Absolute Error (MAE), Root
Mean Squared Error (RMSE), and Mean Absolute Percentage
Error (MAPE). MAE is the unweighted mean of the absolute
we conduct a grid search. The grid search covers  ∈
{2, 4}, learning rates of {0.001, 0.0005}), weight decays of
{0.001, 0.0001}, and learning rate decay rates of {0.1, 0.5}.
We fix the feed-forward dimension to 256,  = 32,  =
20, dropout probability to 0.1, batch size to 32, and the
number of layers for Mamba and the transformer to 3. All
experiments are conducted using GeForce RTX 3090 24GB.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Trafic Forecasting Performance</title>
        <p>results are boldfaced, and the second-best results are
under100
200
300
400
500
100
200
300
400
500
ifc200
f
rTa 0
w300
lFo200
iffc100
rTa 0
0
0
100
200
400
500
100
200
400</p>
        <p>
          500
lined. It is observed that SpoT-Mamba consistently achieves
high rankings across all metrics: MAE, RMSE, and MAPE,
recording the highest average rank among all methods. This
suggests the efectiveness of Mamba’s selective recurrent
scan in modeling spatio-temporal dependency. Compared
to other metrics, SpoT-Mamba demonstrates its best
performance in MAPE, achieving the highest ranking. Among
the baselines, STAEformer [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] shows the most comparable
performance to SpoT-Mamba.
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>4.3. Qualitative Analysis &amp; Ablation Studies</title>
        <p>In Figure 2, we visualize the predictions of SpoT-Mamba
and the ground-truth time series on PEMS04. For this
visualization, predictions are made for four randomly selected
nodes, starting from an arbitrary time step within a test
split, covering 576 consecutive time steps (equivalent to two
days). Since we set  =  ′ = 12, multiple predictions
are concatenated to represent the duration. When multiple
predictions exist at a single time step, we average them. It is
observed that the predicted time series closely aligns with
the ground-truth data.</p>
        <p>Additionally, we conduct the ablation studies of
SpoTMamba on PEMS04. We replace the Mamba blocks used
for the walk sequence embedding (indicated as Walk Scan)
and for scanning along the time axis (indicated as Temporal
Scan) with transformer encoders. Note that when replacing
the Mamba blocks for the Temporal Scan with a transformer
encoder, we reduce the batch size from 32 to 8 due to
Out-ofMemory issues. Results are shown in Table 3. We observed
diferences in performance depending on which type of scan
Walk Scan
Transformer
Transformer</p>
        <p>Mamba
Mamba</p>
        <p>Temporal Scan
Transformer</p>
        <p>Mamba
Transformer</p>
        <p>Mamba
module is replaced. Specifically, when the Walk Scan is
conducted by a transformer encoder, the overall performance
of SpoT-Mamba decreases (first and second rows). On the
other hand, replacing only the Mamba blocks for the
Temporal Scan with a transformer encoder shows negligible
performance diferences (third row).</p>
        <p>This disparity can be attributed to the inherent
diferences between Mamba and Transformer, along with the
application of learnable embeddings that impose biases on
the sequence. Mamba blocks scan inputs recurrently,
inherently considering the sequence order. In contrast, the
transformer encoder does not recognize input sequence order
by itself. Furthermore, while SpoT-Mamba utilizes
learnable embeddings for temporal sequences, i.e., day-of-week
and timestamps-of-day, it does not apply such embeddings
for walk sequences. As a result, the Transformer encoder
struggles to perceive the order in walk sequences, despite
performing well with temporal sequences.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusion and Future Work</title>
      <p>
        In this paper, we explore STG forecasting and introduce a
new Mamba-based STG forecasting model, SpoT-Mamba.
SpoT-Mamba utilizes Mamba blocks to scan multi-way
walk sequences and temporal sequences. This approach
allows the model to efectively capture the long-range
spatiotemporal dependencies in STG, enhancing forecasting
accuracy on complex graph structures. SpoT-Mamba shows
promising results on the real-world trafic forecasting
benchmark PEMS04. For future work, we will extend SpoT-Mamba
to handle graphs with complex relations [
        <xref ref-type="bibr" rid="ref37 ref38">37, 38</xref>
        ] and
evolving graphs [
        <xref ref-type="bibr" rid="ref39 ref40">39, 40</xref>
        ].
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This research was supported by an NRF grant funded by
MSIT 2022R1A2C4001594 (Extendable Graph
Representation Learning) and an IITP grant funded by MSIT
2022-0</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Diao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y. Liu,
          <string-name>
            <given-names>K.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Dynamic spatial-temporal graph convolutional neural networks for trafic forecasting</article-title>
          ,
          <source>in: Proceedings of the 33rd AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>890</fpage>
          -
          <lpage>897</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai.v33i01.
          <fpage>3301890</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>Spectral temporal graph neural network for multivariate time-series forecasting</article-title>
          ,
          <source>in: Proceedings of the 34th International Conference on Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>17766</fpage>
          -
          <lpage>17778</lpage>
          . URL: https: //proceedings.neurips.cc/paper_files/paper/2020/file/ cdf6581cb7aca4b7e19ef136c6e601a5-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Spatial-temporal fusion graph neural networks for trafic flow forecasting</article-title>
          ,
          <source>in: Proceedings of the 35th AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>4189</fpage>
          -
          <lpage>4196</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai. v35i5.
          <fpage>16542</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          , I. Segovia,
          <string-name>
            <given-names>Y. R.</given-names>
            <surname>Gel</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z-</surname>
          </string-name>
          <article-title>gcnets: Time zigzags at graph convolutional networks for time series forecasting</article-title>
          ,
          <source>in: Proceedings of the 38th International Conference on Machine Learning</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1684</fpage>
          -
          <lpage>1694</lpage>
          . URL: https://proceedings.mlr.press/v139/chen21o.html.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <article-title>Hierarchical graph convolution network for trafic forecasting</article-title>
          ,
          <source>in: Proceedings of the 35th AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>151</fpage>
          -
          <lpage>159</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai.v35i1.
          <fpage>16088</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <surname>H. Zhang,</surname>
          </string-name>
          <article-title>Spatiotemporal adaptive gated graph convolution network for urban trafic flow forecasting</article-title>
          ,
          <source>in: Proceedings of the 29th ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1025</fpage>
          -
          <lpage>1034</lpage>
          . doi:
          <volume>10</volume>
          .1145/3340531.3411894.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <article-title>Coupled layerwise graph convolution for transportation demand prediction</article-title>
          ,
          <year>2021</year>
          , pp.
          <fpage>4617</fpage>
          -
          <lpage>4625</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai. v35i5.
          <fpage>16591</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y. Liu,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>T-gcn: A temporal graph convolutional network for trafic prediction</article-title>
          ,
          <source>IEEE Transactions on Intelligent Transportation Systems</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>3848</fpage>
          -
          <lpage>3858</lpage>
          . doi:
          <volume>10</volume>
          .1109/TITS.
          <year>2019</year>
          .
          <volume>2935152</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Long</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. Zhang,</surname>
          </string-name>
          <article-title>Graph wavenet for deep spatial-temporal graph modeling</article-title>
          ,
          <source>in: Proceedings of the 28th International Joint Conference on Artificial Intelligence</source>
          ,
          <year>2019</year>
          , p.
          <fpage>1907</fpage>
          -
          <lpage>1913</lpage>
          . URL: https://dl.acm.org/doi/abs/10.5555/3367243.3367303.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xiong</surname>
          </string-name>
          , W. Zhang, Informer:
          <article-title>Beyond eficient transformer for long sequence time-series forecasting</article-title>
          ,
          <source>in: Proceedings of the 35th AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>11106</fpage>
          -
          <lpage>11115</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai.v35i12.
          <fpage>17325</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sun</surname>
          </string-name>
          , R. Jin,
          <article-title>FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting</article-title>
          ,
          <source>in: Proceedings of the 39th International Conference on Machine Learning</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>27268</fpage>
          -
          <lpage>27286</lpage>
          . URL: https: //proceedings.mlr.press/v162/zhou22g.html.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dustdar</surname>
          </string-name>
          , Pyraformer:
          <article-title>Low-complexity pyramidal attention for long-range time series modeling and forecasting</article-title>
          ,
          <source>in: Proceedings of the 10th International Conference on Learning Representations</source>
          ,
          <year>2022</year>
          . URL: https://openreview.net/forum?id=0EXmFzUn5I.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Lai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-C.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          , H. Liu,
          <article-title>Modeling long- and short-term temporal patterns with deep neural networks</article-title>
          ,
          <source>in: Proceedings of the 42th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2018</year>
          , p.
          <fpage>95</fpage>
          -
          <lpage>104</lpage>
          . doi:
          <volume>10</volume>
          .1145/3209978.3210006.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Proceedings of the 31st International Conference on Neural Information Processing Systems</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>6000</fpage>
          -
          <lpage>6010</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <article-title>Attention based spatial-temporal graph convolutional networks for trafic flow forecasting</article-title>
          ,
          <source>in: Proceedings of the 33rd AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>922</fpage>
          -
          <lpage>929</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai.v33i01.
          <fpage>3301922</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <article-title>Historical inertia: A neglected but powerful baseline for long sequence time-series forecasting</article-title>
          ,
          <source>in: Proceedings of the 30th ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2021</year>
          , p.
          <fpage>2965</fpage>
          -
          <lpage>2969</lpage>
          . doi:
          <volume>10</volume>
          .1145/3459637. 3482120.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>R.-G. Cirstea</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Kieu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Pan</surname>
          </string-name>
          ,
          <article-title>Towards spatio- temporal aware trafic time series forecasting</article-title>
          ,
          <source>in: Proceedings of the IEEE 38th International Conference on Data Engineering</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>2900</fpage>
          -
          <lpage>2913</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICDE53745.
          <year>2022</year>
          .
          <volume>00262</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Goel</surname>
          </string-name>
          , C. Re,
          <article-title>Eficiently modeling long sequences with structured state spaces</article-title>
          ,
          <source>arXiv preprint arXiv:2111.00396</source>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .48550/ arXiv.2111.00396.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gu</surname>
          </string-name>
          , T. Dao, Mamba:
          <article-title>Linear-time sequence modeling with selective state spaces</article-title>
          ,
          <source>arXiv preprint arXiv:2312.00752</source>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .48550/ arXiv.2312.00752.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Tsepa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Graph-mamba: Towards long-range graph sequence modeling with selective state spaces</article-title>
          ,
          <source>arXiv preprint arXiv:2402.00789</source>
          (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2402.00789.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Behrouz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hashemi</surname>
          </string-name>
          , Graph mamba:
          <article-title>Towards learning on graphs with state space models</article-title>
          ,
          <source>arXiv preprint arXiv:2402.08678</source>
          (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          . 48550/arXiv.2402.08678.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , A. Coster,
          <article-title>Stg-mamba: Spatial-temporal graph learning via selective state space model</article-title>
          ,
          <source>arXiv preprint arXiv:2403.12418</source>
          (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2403.12418.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Long</surname>
          </string-name>
          , G. Song,
          <string-name>
            <given-names>K.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>Spatial-temporal graph ode networks for trafic flow forecasting</article-title>
          ,
          <source>in: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>364</fpage>
          --
          <lpage>373</lpage>
          . doi:
          <volume>10</volume>
          .1145/3447548.3467430.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <article-title>Spatio-temporal adaptive embedding makes vanilla transformer sota for trafic forecasting</article-title>
          ,
          <source>in: Proceedings of the 32nd ACM International Conference on Information and Knowledge Management</source>
          ,
          <year>2023</year>
          , p.
          <fpage>4125</fpage>
          -
          <lpage>4129</lpage>
          . doi:
          <volume>10</volume>
          .1145/3583780.3615160.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>C.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <article-title>Spatial-temporal synchronous graph convolutional networks: A new framework for spatial-temporal network data forecasting</article-title>
          ,
          <source>in: Proceedings of the 34th AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>914</fpage>
          -
          <lpage>921</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai.v34i01.
          <fpage>5438</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Shahabi</surname>
          </string-name>
          , Y. Liu,
          <article-title>Difusion convolutional recurrent neural network: Data-driven trafic forecasting</article-title>
          ,
          <source>in: Proceedings of the 6th International Conference on Learning Representations</source>
          ,
          <year>2018</year>
          . URL: https://openreview.net/forum?id=SJiHXGWAZ.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Adaptive graph convolutional recurrent network for trafic forecasting</article-title>
          ,
          <source>in: Proceedings of the 34th International Conference on Neural Information Processing Systems</source>
          ,
          <year>2020</year>
          , p.
          <fpage>17804</fpage>
          -
          <lpage>17815</lpage>
          . URL: https://dl.acm.org/doi/abs/10.5555/ 3495724.3497218.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>C.</given-names>
            <surname>Shang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bi</surname>
          </string-name>
          ,
          <article-title>Discrete graph structure learning for forecasting multiple time series</article-title>
          ,
          <source>in: Proceedings of the 9th International Conference on Learning Representations</source>
          ,
          <year>2021</year>
          . URL: https://openreview.net/ forum?id=
          <fpage>WEHSlH5mOk</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Long</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Zhang, Connecting the dots: Multivariate time series forecasting with graph neural networks</article-title>
          ,
          <source>in: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2020</year>
          , p.
          <fpage>753</fpage>
          -
          <lpage>763</lpage>
          . doi:
          <volume>10</volume>
          .1145/3394486.3403118.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <article-title>Gman: A graph multi-attention network for trafic prediction</article-title>
          ,
          <source>in: Proceedings of the 34th AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1234</fpage>
          -
          <lpage>1241</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai. v34i01.
          <fpage>5477</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          , C. Han,
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          , Pdformer:
          <article-title>Propagation delay-aware dynamic long-range transformer for trafic flow prediction</article-title>
          ,
          <source>in: Proceedings of the 37th AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>4365</fpage>
          -
          <lpage>4373</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai.v37i4.
          <fpage>25556</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. W.</given-names>
            <surname>Tsang</surname>
          </string-name>
          ,
          <article-title>Stnorm: Spatial and temporal normalization for multivariate time series forecasting</article-title>
          ,
          <source>in: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2021</year>
          , p.
          <fpage>269</fpage>
          -
          <lpage>278</lpage>
          . doi:
          <volume>10</volume>
          . 1145/3447548.3467330.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>Spatialtemporal identity: A simple yet efective baseline for multivariate time series forecasting</article-title>
          ,
          <source>in: Proceedings of the 31st ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2022</year>
          , p.
          <fpage>4454</fpage>
          -
          <lpage>4458</lpage>
          . doi:
          <volume>10</volume>
          .1145/3511808.3557702.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Karypis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Zhang,</surname>
          </string-name>
          <article-title>Deep graph library: A graphcentric, highly-performant package for graph neural networks</article-title>
          , arXiv preprint arXiv:
          <year>1909</year>
          .
          <volume>01315</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.
          <year>1909</year>
          .
          <volume>01315</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Paszke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Massa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lerer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bradbury</surname>
          </string-name>
          , G. Chanan,
          <string-name>
            <given-names>T.</given-names>
            <surname>Killeen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gimelshein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Antiga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Desmaison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kopf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>DeVito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Raison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tejani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chilamkurthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Steiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chintala</surname>
          </string-name>
          ,
          <string-name>
            <surname>Pytorch:</surname>
          </string-name>
          <article-title>An imperative style, high-performance deep learning library</article-title>
          ,
          <source>in: Proceedings of the 33th International Conference on Neural Information Processing Systems</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>8024</fpage>
          -
          <lpage>8035</lpage>
          . URL: https: //proceedings.neurips.cc/paper_files/paper/2019/file/ bdbca288fee7f92f2bfa9f7012727740-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          ,
          <source>in: Proceedings of the 3rd International Conference on Learning Representations</source>
          ,
          <year>2015</year>
          . URL: https://doi.org/10.48550/arXiv.1412.6980.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Whang</surname>
          </string-name>
          ,
          <article-title>Dynamic relationattentive graph neural networks for fraud detection</article-title>
          ,
          <source>in: Proceedings of 2023 IEEE International Conference on Data Mining Workshops (ICDMW)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1092</fpage>
          -
          <lpage>1096</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICDMW60847.
          <year>2023</year>
          .
          <volume>00143</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Whang</surname>
          </string-name>
          ,
          <article-title>Representation learning on hyper-relational and numeric knowledge graphs with transformers</article-title>
          ,
          <source>in: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>310</fpage>
          -
          <lpage>322</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J.
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vazirgiannis</surname>
          </string-name>
          ,
          <article-title>Dgraph: A large-scale financial dataset for graph anomaly detection</article-title>
          ,
          <source>in: Proceedings of the 36th International Conference on Neural Information Processing Systems</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>22765</fpage>
          -
          <lpage>22777</lpage>
          . URL: https: //proceedings.neurips.cc/paper_files/paper/2022/file/ 8f1918f71972789db39ec0d85bb31110-Paper-Datasets_ and_Benchmarks.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Whang</surname>
          </string-name>
          , InGram:
          <article-title>Inductive knowledge graph embedding via relation graphs</article-title>
          ,
          <source>in: Proceedings of the 40th International Conference on Machine Learning</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>18796</fpage>
          -
          <lpage>18809</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>