<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Benchmark Framework for Trajectory Vectorization in Unsupervised Settings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>CristianoLandi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>PetrosMandali</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>sand Nikos Pelekis</string-name>
          <email>pelekis@unipi.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Univeristy of Pisa</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>KDD-Lab at ISTI-CNR</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Statistics and Insurance Science, University of Piraeus</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>The growth of location-tracking technologies has generated large spatiotemporal datasets that require eficient analysis methods. Recent approaches focus on vectorizing the data as a preprocessing step to enable the application of of-the-shelf classical machine learning techniques. To date, such approaches have primarily been evaluated in supervised tasks. This paper presents a benchmark framework for assessing these methods in unsupervised settings, where no target labels are available. We extend and systematically evaluate four families of vectorization techniques: feature-based, shapelet-based, dictionary-based, and matrix-based. The framework assesses both clustering quality and cross-vectorization diversity through quantitative and qualitative analyses. Overall, the ifndings demonstrate that these vectorization approaches ofer a valuable option for unsupervised trajectory analysis. Additionally, our paper provides insight into which vectorization technique is most suitable for the specific characteristics of the task under analysis.</p>
      </abstract>
      <kwd-group>
        <kwd>Spatial-temporal systems</kwd>
        <kwd>Knowledge representation and reasoning</kwd>
        <kwd>Human Mobility Modeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>https://github.com/cri98(lCi. Landi)</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Spatiotemporal datasets are continually expanding and include videos, social media posts, GPS traces,
and other time-series data. Efective analysis of such data can have a significant societal impact, helping
to address challenges in environmental and climate change, public s1e]c,uarnidtyh[ealthcare2][.</p>
      <p>
        Several researchers in Mobility Data Science (M3D]Sf)oc[us on understanding and analyzing
this data and producing reports on various observed phenomena. This led to the development of
application- and data-specific procedure4]s, [which reduce the reusability of algorithms. Silva et
al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] compared various trajectory analysis methods and found that approximately 70% of them are
tailored to specific applications such as transportation mode identification. To overcome this limitation,
researchers have turned to the field of time series for inspiration and have explored various ways
to represent spatiotemporal data to meet diferent analysis objectives. The concept is to transform
the input sequential data into diferent vector-based feature spaces that can be utilized with standard,
general-purpose machine learning models, thereby minimizing the need for manual feature engineering.
For example, in 6[], the authors explore the concept of shapelet transform, intrGoedoulceitngto
transform input trajectories into a (dis)similarity matrix that compares the input data with specific
discriminant subtrajectories extracted from it.
      </p>
      <p>Although these methods achieve state-of-the-art performance in classification tasks, they have
not yet been evaluated in an unsupervised setting. We therefore propose a unified benchmarking
framework for evaluating trajectory vectorization techniques in such a setting. Within this framework,
we formulate the problem as trajectory clustering, which is the task of identifying groups of trajectories
that exhibit similar movement patterns. Despite multiple definitions of trajectory similarity, most prior
work assumes that trajectories are similar if spatiall7y],cmloaskei n[g clustering based on movement</p>
      <p>CEUR
Workshop</p>
      <p>ISSN1613-0073</p>
    </sec>
    <sec id="sec-3">
      <title>2. Background and Problem Setting</title>
      <p>Trajectory clustering consists of partitioning a set of trajectories into grousipms isluarchtrtahjaetctories
are assigned to the same cluster, while dissimilar trajectories are placed in diferent clusters. In mobility
research, similarity is typically interpreted as spatial proximity between tr7a]j.eIcntocroinetsr[ast, our
objective is to examine a more abstract and dataset-dependent notion of similarity. For instance, given
a dataset in which certain trajectories exhibit hazardous driving events, such as harsh braking, abrupt
turns, or sustained high-speed intervals, we aim to determine whether, and to what extent, diferent
vectorizations reveal these behavioral distinctions in the absence of labeled data.</p>
      <p>We begin by defining the simplest type of mobility data, namely the position of an object in
2dimensional space over time. Formally:
increasing timestam p.</p>
      <p>Definition 1</p>
      <p>(Trajectory.) A trajectory  is a sequence of spatio-temporal poin ts=
{(lat0, long0,  0), … , (lat , long ,   )} ∈ ℝ×3 where the spatial vect o⃗rs= (lat , long ) are sorted by
This type of data is typically not sampled at a constant rate, as observations are often recorded irregularly
due to sensor limitations, energy-saving policies, communication constraints, or event-driven acquisition
processes. Additionally, trajectories into a dataset often have a diferent number of observations, so we
denote by ∈ ℝ ××3</p>
      <p>a set o f trajectories of lengatthmost  . Many mobility analysis tasks can be
formulated as a two-phase process: first, the data are transformed into a tabular vector representation,
then analyzed using classical machine learning methods. For example, consider the following definition
of the trajectory clustering task:
Definition 2 (Trajectory Clusteri.ngG)iven a trajectory dat ase,tTrajectory Clustering is the task of
defining a function from the space of possible input trajectoriteos an outpu t∈ 
denoting its
cluster identity.</p>
      <p>We can define the clustering funct ioans the composition of a vectorization fun c,twiohnich maps
each trajectory to a fixed-size feature vec,taonrd a functioℎn, which maps ( )
to one cluster. By
changing the vectorization funct,iotnhe definition can accommodate diferent clustering objectives.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Related Work</title>
      <p>
        We organize the related work around our definition of a trajectory clustering
afusntchteioconmposition of a vectorization func tiaonnd a clustering functℎiotnhat assign(s ) to a cluster. Sinℎcecan
be any of-the-shelf algorithm, such a-smeans oroptics, we first review the Mobility Data Science
(MDS) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] literature from the perspective of trajectory vectorization. We then summarize existing
clustering approaches in MDS to emphasize the diversity of similarity objectives. We conclude by
situating our framework within the current literature.
      </p>
      <p>
        Vectorization techniques for MDS To better structure this section, we adopt a time series taxonomy
that categorizes approaches by the core idea underlying data transfo8r].mFaetaitounre-[based
vectorizations compute explicit motion descriptors like speed, acceleration, and heading change and aggregate
them into trajectory-level vectors; they often require the practitioner to select features appropriate to
the specific task and data characteristics, and can be sensitive to preprocessing steps like resampling or
segmentation and aggregation choic9e].sI[n shapelet-based vectorization, trajectories are represented
by their similarity to discriminative sub-trajectories, following the shapelet idea from1t0i]m. e series [
Such approaches are fully data-driven, avoiding the heavy feature engineering often required in
featurebased approaches. In MDS, the most prominent approachesMaorveelets [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] andGeolet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In
dictionary-based approaches, the core idea is to discretize the data into tokens, enabling the use of
text-inspired representations such as bag-of-words or tf–idf. In MDS literature, most works focus on
data compression, for example, through tessellation techniques or pathlet learning algorithms. In our
setting, however, discretization must preserve episodic events rather than merely serving compression
or coarse partitionin1g2][. Matrix-based vectorizations encode trajectories into structured objects (e.g.,
images/accumulators) and then embed them into vectors, namely transform-based encodings that can
provide invariances and reduce overlap artifacts, as in rotation/scale-invariant repr1e3s].entations [
Deep-learning vectorization techniquuesse trained encoders to produce task-specific embeddings. Their
results are highly sensitive to architectural and optimization choices, making it challenging to
disentangle the impact of the vectorization strategy itself. A comprehensive assessment of such methods is
deferred to future work.
      </p>
      <p>
        Trajectory clustering and similarity objectives.Trajectory clustering groups trajectories
according to a chosen notion of similarity, largely determined by the representation and distance functions.
Existing surveys1[
        <xref ref-type="bibr" rid="ref4 ref7">4, 7</xref>
        ] note that classical methods primarily focus on spatial or temporal proximity,
whereas many applications requbireehavioral similarity, capturing comparable motion dynamics across
diferent regions. This has motivated sub-trajectory-based approaches sturcahcalsus [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], as well
as semantic and mobility-oriented methods that move beyond pure geom16e,t1r7y].[In our setting,
clustering is used as a diagnostic tool: a vectorization is efective if it yields a cluster structure that
aligns with the intended similarity notion under a controlled evaluation.
      </p>
      <p>Positioning of our framework. To the best of our knowledge, there is no existing framework that
benchmarkscross-family trajectory vectorizations for an unsupervised clustering objective under a
single, fixed evaluation pipeline; most related eforts either standardize mobility primitives (libraries) or
benchmark diferent tasks and stages rather than representation families.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Evaluation framework</title>
      <p>In this section, we propose a unified benchmarking framework, summarized in Fig1u,troe evaluate
of-the-shelf trajectory vectorization techniques. Rather than asking whether two trajectories are
geographically close, our framework allows us to explore how diferent vectorization techniques
naturally group trajectories that share behavioral or structural similarities. To make this evaluatio
transparent and reproducible, we define a common pipeline and measure both (i) clustering quality
within a representation space and (ii) diversity/consistency across representation spaces. The framework
is composed of 3 modules: (i) a vectorization fun ct;i(oiin) an optional dimensionality reduction step;
(iii) a clustering functiℎo. nFinally, the resulting trajectory clustering is evaluated both in terms of
clustering quality within each vectorization and in terms of diversity/consistency across vectorization
spaces.</p>
      <p>In the following paragraphs, we motivate the inclusion of the three modules and outline the algorithms
considered for each in the experimental section. The framework is general and can accommodate a
wide range of algorithms per module; the choices made here are preliminary and will be extended in
future work.</p>
      <p>
        Vectorization functions. The first step in our framework is selecting a set of vectorization functions
to compare. The scope of this module is to transform complex trajectory data into a tabular-like format
such that a classical of-the-shelf machine learning technique can be applied. Motivated by recent
work in mobility, we select 4 diferent functions. Additionally, inspired by the time series literature,
we divided them into feature-based, shapelet-based, dictionary-based, and matrix-based functions. For
feature-based vectorization, i.e., where explicit descriptors are extracted from the data, we selected the
Trajectory Interval ForTesItF)([
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. TIF randomly generates a set of intervals and, for each interval
and trajectory in the dataset, it extracts a wide range of descriptors (speed, acceleration, sinuosity
intensity use etc.). Foshrapelet-based vectorization, which extracts a set of discriminatory subsequences
and uses them as references for data transformation, wGeesoelectt. Geolet relies on statistical
methods to identify discriminatory subsequences, and represents each trajectory by its similarity to
the selected patterns. SinGceeolet was designed for supervised settings, its subsequences selection
requires prior knowledge. In Secti4o.1n, we extendGeolet to unsupervised tasks.
Findingdiactionarybased vectorization was more challenging. In time series analysis, this family of methods discretizes
the input into sequences of symbols, which are then matched against a learned dictionary of patterns.
This allows temporal dynamics to be represented through the frequency or arrangement of symbolic
motifs. In the mobility literature, discretization is typically performed using tessellation techniques
or pathlet-learning approach1e2s],[but these methods mainly capture global spatial characteristics
of movement. As a result, two harsh right turns occurring in diferent geographic regions would be
encoded as diferent symbols. In Sectio4n.2, we propose a simple yet efective approach to address this
limitation. Finally, fmoartrix-based vectorization, we selected the algorithm propos1e9d],inna[mely
RoSITa. This vectorization technique represents each trajectory in a scale- and rotation-invariant
representation by performing the Hough transform.
      </p>
      <p>Optional feature reduction. As some vectorization functions may produce very high-dimensional
or sparse feature spaces, this step aims to mitigate the curse of dimensionality. Due to their widespread
use among practitioners, we selected Principal Component AnaPlCyAs)isa(nd t-distributed Stochastic
Neighbor Embedding t(-SNE). We also introduced a pass-through option that allows the clustering
algorithm to be executed directly on the full vectors.</p>
      <p>Clustering algorithms. The third module applies an of-the-shelf clustering algorithm to the resulting
(reduced) vectorized dataset. We consider two widely used methods from diferent fa-meilaines:
for centroid-based clusteringoapntdics for density-based clustering. This selection will be expanded
in future work.</p>
      <sec id="sec-5-1">
        <title>4.1. Geolet for Unsupervised Tasks</title>
        <p>This section extendGseolet to unsupervised taskGs.eolet is composed of four steps: (i) trajectory
segmentation into a set of candidate discriminative sub-trajectories, (ii) sub-trajectory normalization,
(iii) sub-trajectory filtering to select the most discriminative candidates, and (iv) transformation of the
dataset into a dissimilarity matrix between each trajectory and the selected discriminative candidates
SinceGeolet was originally proposed for supervised tasks, the filtering step relies on dataset labels to
identify, through statistical tests, the most relevant discriminative sub-trajectories. We therefore focu
on replacing this step with a more general approach suitable for unsupervised settings. Inspired by
time series literature, we build upon the conceupntsuopfervised shapelets (u-shapelets)2[0], which
extends the original notion of time series shapelets to scenarios where class labels are unavailable. The
underlying intuition of the proposed work is that, even in unlabeled data, certain short subsequences
tend to recur within groups of instances while remaining dissimilar to the rest of the dataset. Accordingly,
a u-shapelet can be identified by assessing how efectively a candidate subsequence separates the dataset
into two partitions: trajectories that contain a close match to the subsequence and those that do not. W
incorporate the following iterative proceduGreeoinletto. Candidate subsequences are first extracted
from a reference trajectory using spatial, temporal, or spatio-temporal sliding windows. For each
candidate, a distance measure is used to compute a dissimilarity vector for all trajectories in the dataset
The distance measure is defined as the minimal distance over all possible alignments of the subsequence
within each trajectory. Sorting this vector produocrdeesraline, over which all possible split points
are evaluated. For each split, a gap score measures the separation between the two induced groups by
balancing the diference in their mean distances against their respective variances. Given a subsequence
 that partitions the dataset into two trajector y garnodup,sthe gap score is formally defined as:
gap =   −   − (  +   )
(1)
Where  and  are the means of the best fitting oifn all trajectorie s inand in respectively,
while  and  represent the standard deviations. The subsequence that maximizes this separation is
selected as a discriminative subsequence. Trajectories that are well explained by the selected u-shapelet
are then removed from the dataset. The process is repeated on the remaining instances to extract
additional u-shapelets, until the desired number of u-shapelets has been extracted or no further splits
of the remaining instances are possible.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Dictionary-based vectorization</title>
        <p>
          This section lays the groundwork for shifting from continuous to discrete representations of trajectories.
We introduce a set of algorithms designed to transform trajectory data into sequences of symbols.
This discretization provides several important benefits: it enables the use of highly eficient data
structures; it simplifies storage and indexing by reducing continuous data into compact symbolic
forms; and it facilitates faster similarity searches and pattern discovery. We formally define the
problem of trajectory discretization as the composition of a normalization function and a discretization
function, i.ed.,iscretize ∘ normalize. The discretize ∶ ℝ ′×3 → Σ ′ function maps each observation
in the normalized trajectory to a symbol from an alpΣh, apbreotducing a symbolic sequence. The
normalize ∶ ℝ×3 → ℝ ′×3 function transforms the original trajectory into a normalized version.
This step serves two purposes: first, it acts as an approximation function simpialarintsoax [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ],
reducing noise and redundancy; second, it regularizes the trajectory by interpolating or resampling
to ensure a consistent sampling rate or spatial distance. This normalization ensures that trajectories
of comparable “length” yield the same number of symbols, enabling more meaningful comparisons:
trajectories exhibiting similar motion patterns are more likely to share a similar symbolic representation.
Finally, the discretized trajectory data can be vectorized by comtpfu-tidinfgmaatrix over words
extracted via a sliding window.
        </p>
        <p>Normalization We present three possible normalization strategies suited for trajectory data. The first
involves resampling trajectories at constant time intervals, ensuring that observations are uniforml
spaced in time. The second normalizes by constant travel distance, where a new observation is sampled
each time the moving object covers a fixed spatial distance. However, due to positional noise, this
method may produce redundant samples when the object is stationary. To address this, we introduce a
third variant that generates a new observation only after the object has moved a minimum distance
from the last sampled point, rather than after a fixed cumulative distance. In all three strategies, missing
observations are interpolated, and excess ones are subsampled to preserve temporal or spatial regularity.
Discretization We present several strategies for assigning a sequence of symbols to a trajectory. A
ifrst, straightforward approach relies on tessellation techniques already established in the literature
Figure2 (left) illustrates a conceptual example using Geoh2a2s]h, w[here each geohash region crossed
by the trajectory generates a corresponding s1y.mTbhoelsame idea can be extended to other spatial
partitioning methods. However, since our goal is to analyze multiple aspects of trajectories beyond
mere origin–destination movement, we aim to design a discretization strategy that yields comparable
symbolic vectorizations across diferent geographic contexts. One possible approach is to leverage the
hierarchical structure of some tessellation encodings, such as Geohash. Leveraging this property, we
can design discretization techniques based on sufix relationships among GeoHash sequences, enabling
vectorizations that remain structurally consistent even when trajectories span diferent geographic
areas. Figure2 (center-left) illustrates such an example. The main limitation of this approach is its
reliance on a fixed spatial grid: two similar movements are represented by similar subsequences only if
they occur within the same subregion. For instance, consider two right-turn events–oneCfBr”om cell “
to “AD” and another fromBC“” to “BD”. Despite being similar, the discretized subsequences difeCrB”(“
vs. “CD”) solely due to their geographic position. This spatial dependence limits the generalization of
symbolic vectorizations across regions. To overcome this, we introduce two alternative discretization
strategies that are independent of geographic location and instead rely on the intrinsic dynamics of
motion. Specifically, thetraveling direction and theturning angle. The first computes the direction
between consecutive observations and assigns symbols based on predefined angular intervals, as shown
in Figure2 (center-right). The second focuses on the turning angle, i.e., the change in direction between
successive segments. Revisiting the right-turn example, this vectorization consistently captures both
events as the same symbolic pattern, generating the symG”btowli“ce. The resulting longest common
subsequence, “AG”, efectively captures a straight movement followed by a right turn, ofering a
locationand orientation-independent yet behaviorally meaningful vectorization. Additionally, to account fo
stationary behavior, we introduce a dedicated symbol that indicates when the object is not moving.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Experimental Setting</title>
      <p>This section describes the hyperparameters for vectorization, feature reduction, and clustering evaluated
in our experiments. Due to limited prior knowledge of the efectiveness of these vectorization techniques
for unsupervised tasks, we designed the experiments to evaluate all possible combinations of the three
components. Consequently, each vectorization hyperparameter setting was evaluated with every
feature reduction hyperparameter configuration, and each of those combinations was in turn evaluated
with every clustering hyperparameter configuration. As the total number of experiments increases
exponentially with the number of hyperparameters across modules, we prioritize extensive exploration
of the vectorization hyperparameters while restricting the feature reduction and clustering stages to
smaller set of representative methods and settings.
1For clarity, the figure shows the full trajectory, though after normalization each cell (symbol) correspond to a single point.</p>
      <sec id="sec-6-1">
        <title>5.1. Hyperparameters</title>
        <p>Vectorization method hyperparameters. Fortif, we explored intervals defined by the number
of observations, time, or distance-based intervals. The number of intervals ranged from 2 to 20, with
diferent constraints on minimum and maximum lengths. We disabled the use of latitude and longitude
to avoid making position-based tasks trivialG. eFoolret, we tested diferent numbers of subsequences,
from 5 up to 100, using a Geohash-based partitioning with precision levels between 3 and 7. To select
the most discriminative shapelets, we experimented with the algorithm described in4.S1e.cFtoiron
the distance computation between trajectories and shapelets, we tested standard Euclidean,
timeinterpolated Euclidean, and space-interpolated Euclidean distaRnocSesIT.aF,owre used a time-based
sliding window ranging from 10 seconds to 1 hour. For the Hough Transform, the angle resolution
varied between 1° and 10°, and the image size ranged fro×m505t0o 1080×1080 pixels. Consequently,
the Hough accumulator had dimensions ranging f7r0o×m18 to1527 × 180. We tested clustering using
the full accumulator, its vertically summed-squared version, and its FFT transf2o3r]m.F,oasr itnh[e
dictionary-based approach described4.2in,we evaluated all discretization and resampling strategies.
Time-based resampling was tested every 1, 50, or 300 seconds, and distance-based resampling was tested
every 3, 50, or 300 meters. For tessellation-based discretization, we used Geohash with precision levels
3 to 7 and sufix variants, ignoring the first 1 to 3 charactLeorws.er precision values help to capture
movement patterns of objects traversing wide aDrieraesc.tion-based discretization employed alphabets
of 4 to 16 symbols, plus one for stationarity. We then tbuf-iildtfamatrix over words extracted with
sliding windows of 3 to 20 symbols.</p>
        <p>Feature reduction hyperparameters. Regarding the feature reduction step, we evaluated several
approaches commonly adopted in machine learning. In particular, we experimented with Principal
Component AnalysisP(CA) and t-distributed Stochastic Neighbor Embedtd-iSnNgE(). For both methods,
we tested output dimensionalities of 5, 10, 50, and 100 componentts-.SFNoEr, we additionally varied
the perplexity parameter over 5-100. All remaining hyperparameters foPrCAboatnhdt-SNE were left
unchanged and set to their default values as defined in the scikit-learn implem2e.nWteatdieonnote the
output dimension using the symbo.l
Clustering methods hyperparameters. Finally, regarding the clustering step, we limited our
experiments t o-means andOPTICS. For -means, we evaluated values oefqual to 2, 3, 4, 6, and 10.
ForOPTICS, we tested two values for the minimum number of samples required in the neighborhood
of a point for it to be considered a core point: 5 and 10% of the trajectories in the dataset. All
remaining hyperparameters were left at their default v2a.lPulesase note that clustering quality metrics
were computed on the non-reduced vectorized representations. This choice was made to enable a fair
comparison of clustering scores across diferent vectorization methods sharing the same hyperparameter
configurations. Otherwise, such metrics would also be conditioned by the output of the feature reduction
step, thereby confounding the evaluation of the vectorization methods themselves.</p>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Datasets</title>
        <p>This section describes the datasets used in our experiments and the preprocessing steps applied. We
consider both real-world and synthetic datasets. This distinction allows us to capture diferent levels of
similarity to real data. Synthetic datasets are fully controlled and used to test specific, isolated aspect
of the vectorization functions.</p>
        <p>Real-world datasets. We selected five real datase3tfsor GPS trajectory classification with
different sizes, semantics, and classification objectives,ani.iem.,als, vehicles, seabirds, geolife.
The vehicles andgeolife datasets mainly represent road-constrained vehicle movements, whereas
2Specific default hyperparameters are available in sk-leta.lryn/r(lINN)and in the project GitHub repositotr.lyy(/oSmLh).
3animals andvehicles: t.ly/klP5 M,seabirds: t.ly/LpMx3,geolife: t.ly/FbqVa.
animals andseabirds capture unconstrained animal movements on the ground and in tThaeblaei1r.
(top five rows) reports descriptive statistics for the datasets, including the number of trajectories, the
total number of observations, the average and standard deviation of observations per trajectory, th
average sampling rate, and the number and proportion of target classesa.nFiomraltshedataset, the
classification task is to recognize three species.vFeohricles, the objective is to distinguish between
buses and trucks. Fosreabirds, the objective is to recognize the flying trajectories of three species of
seabirds. For thgeeolife dataset, we set the classification problem to recognize trajectories of public
versus private means of transport a6s]i.n [
Synthetic Datasets For synthetic datasets, we use an extended version of those introd2u4c],ed in [
consisting of a larger number of trajectories per dataset. The synthetic datasets are deliberately simple
to isolate the hypothesis under analysis, minimizing confounding variables such as trafic, trafic lights,
and road priorities.</p>
        <p>To measure the extent to which diferent vectorization techniques capture the movement’s geometry,
we generated two synthetic datasets shown in F3i.guTrhee first datasetS,traight-Right (SR),
consists of objects that either move in a straight line or make a 90-degree right turn. Each object in this
dataset travels at a constant speed for precisely 100 seconds. The secondYadoa,tisaisnestp, ired by the
dataset introduced by Yao et al2.5i]n. I[t includes three distinct movement patterns: straight, circling,
and bending. Unlike thSeR dataset, this dataset introduces additional complexity to the analysis task,
as each object travels in a diferent direction with a variable sampling rate for a varying number of
seconds between 70 and 130. We generated 100 trajectories per class, resulting in 20S0Rfdoarttahset
and 300 for thYeao dataset.</p>
        <p>Kinematic-related synthetic datasets.We generated a simple synthetic dataFseatst– Slow (FS)–
of moving objects turning 90 degrees right and traveling at two diferent speeds. By isolating a specific
structural movement pattern while varying the travel speed, we aim to assess the efectiveness of
vectorization techniques in capturing movement dynamics without confounding efects from other
variables. The target labels associated with the trajectories could be either moving “slow” or “fast”.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Experiments Results</title>
      <p>In this section, we analyze how distinct vectorization captures structural patterns in the data unde
an unsupervised setting. Standard clustering practice often relies on the knee method or selects
hyperparameters that maximize the silhouette (SIL) score. However, both heuristics depend on pairwise
distances; hence, they are afected by the shape of the clusters. Additionally, clustering is applied in the
vectorized space; therefore, these metrics can be used only when comparing clustering outcomes derived
from the same vectorization type and configuration. In our framework, we compute the SIL score in the
non-reduced vectorization spaces, enabling the use of this metric to compare clustering results across
various feature reduction techniques. Additionally, since our datasets include class labels, we use them
#trj #obs #obs per trj AvgΔtime (s)</p>
      <p>Classes
Animals 102
15k
as a proxy for ground truth and select, for each component of the pipeline, the hyperparameters that
maximize the Fowlkes–Mallows (FM) score. Alternative heuristics to select the best clustering results
without any prior knowledge are left as future work. SIL and FM scores alone do not characterize
diferences between feature spaces; they only quantify cluster separation and agreement with the
labeled structure. To evaluate how the relative positions of instances vary across vectorizations o
diferent dimensionality, we exploit the Triplet Loss introduc2e6d],ian w[idely used one-shot learning
loss function that encourages embeddings in which similar instances are close together, and dissimilar
instances are far apart. In our framework, we adopt a variant called Random TripletrtAac)c,uracy (
proposed in [27], to measure the impact of feature reduction techniques. This measure operates by
randomly sampling triplets of instances from two feature spaces and evaluating the proportion of
triplets in which the relative distance ordering between instances is preserved across both spaces. A
score near 1 indicates that both spaces encode instance relationships consistently, while a score near 0
suggests that they reverse these relationships, placing samples that are close in one space farther apart
in the other. Figur4eillustrates three scenarios. In the first, 90 instances are randomly projected into
two unrelated feature spaces. Srintcaesamples triplets at random, we repeat the evaluation 100 times
and obtain an average score≈o0f.50, indicating that an ordering sucehda(,s) ≤ ed(, ) in one
space fails to hold in the other in roughly half of the cases. In the second scenario, both vectorizations
yield the same three clusters; however, the points are randomly arranged within each cluster, and the
orange cluster is closer to diferent neighbors across the two spaces. Under the same experimental
conditions, the resulting score is approximately 0.70. Please note that such a scenario would have the
same FM score and approximately the same SIL score. Finally, in the third scenario, the vectorizations
difer only by scale, leading to a score close to 1.</p>
      <p>Tables2 and3 present the experimental results on vectorization diversity, reported thrrtoaugh the
score, and clustering quality, reported through the SIL and FM scores. The optimal pipeline for each
vectorization and dataset is selected using the FM score2. aTdadbilteionally reports the SIL score
associated with the pipeline that maximizes the FM score, as well as the FM and SIL scores of the
pipeline that performs best with respect to the SIL score, thereby simulating a selection process that
relies solely on internal clustering criteria.</p>
      <p>We begin by examining the upper part of Tab2l,ewhich reports the FM and SIL scores of the
pipeline achieving the highest FM. To facilitate interpretation, FM and SIL scores are averaged across
datasets with the same characteristics: shape-based, kinematic-based, and real-world trajectories. By
selecting the pipeline that maximizes FtMif, achieves the overall best performance across all datasets.
NotablyR,oSITa performs particularly well on geometry-based datasets, which aligns with its design.
MoreoverG,eolet consistently attains the overall best SIL scores, indicating an average higher degree
of cluster separation w.r.t. other vectorization techniques, which can be attributed to its sub-trajector
ifltering strategy based on tghaep score. For the other vectorizations, we do not observe any interesting
patter4n. Instead, selecting the best pipeline that maximizes SIL yields slightly lower FM scores across
all vectorization methods and dataset categories. On the other side, SIL scores are extremely high for
most datasets and vectorization techniques. Our interpretation is as follows: we found pipelines with
high FM or SIL scores, which means that, for some vectorization hyperparameters, each vectorization is
able to discover clusters similar to the ground truth or produce vectorization spaces with well-separated
clusters, but with lower similarity to the ground truth. While this highlights the flexibility of such
vectorization in capturing higher-level mobility patterns beyond mere spatial proximity, it also makes it
dificult for practitioners to determine the optimal hyperparameters without additional information
about the data.</p>
      <p>Despite that, several consistent trends emerge from the results2i.nTThaebtleif vectorization
4We would like to recall the dependencies between SIL scores and vectorization spaces. As a result, we can only consider the
single SIL score, comparison between vectorization methods or diferent datasets may be misleading
y tif
etr Geolet
eom RoSITa
G tf-idf
ic tif
ta Geolet
m
in RoSITa
e
K tf-idf
lrd tif
ow Geolet
ea RoSITa
l
R tf-idf</p>
      <p>tif
demonstrates the highest overall FM values across almost all dataset categories. This indicates that
the handcrafted feature extraction pipeline efectively captures both spatial and temporal properties of
movement, producing clusters that closely correspond to the actual labeled classes. GCeoonlveetrsely,
achieves superior performance on real-world datasets, where data are more complex, irregular, and
noisy. This suggests that shapelet-based vectorizations are especially well-suited to capture localized
movement variations or subpatterns (e.g., turns, stops, or loops) that often define behavior in natural
or human mobility. The relatively lower SIL values, combined with high FM scores, indicate that,
whileGeolet clusters may be less compact geometrically, they still align closely with the semantic
ground truth. ThReoSITa vectorization performs best on geometry-related datasets, as expected due to
its design for rotation- and scale-invariant shape encoding. This confirms that geometric alignment
is particularly important when trajectories exhibit clear structural regularitiReos.SIHTao’wsever,
performance drops slightly on datasets dominated by dynamic diferences, suggesting its invariance
properties may suppress meaningful kinematic variation. Fitnfa–lildyf, (dictionary-based) achieves
competitive results on real-world datasets, suggesting significant potential for this technique.</p>
      <p>Table3 reports threta scores between pairs of vectorizations across all dataset categories.
Vectorization hyperparameters follow the configuration that yielded the highest FM score. Overall, the scores
obtained by comparing diferent transformation spaces average a0r.o54u3n,dwith values reaching up
to0.714 for specific dataset–vectorization combinations. As illustrated in4, Feingtuirreely unrelated
vectorizations typically yield a score of about 0.500, while points that fall into the same clusters but
difer in their cluster-wise positioning tend to produce scores near 0.700. Interpreting these results
benefits from revisiting Tabl2e, particularly the FM values. Since the FM score, defined as the geometric
mean of precision and recall, ranges f0rotmo1 and reflects cluster similarity, its generally high values
here suggest that morstta scores correspond to cases where clusters are similar but difer in their
internal positioning.</p>
      <p>The highestrta values ≈( 0.714) occur betweenGeolet andRoSITa for geometry-based datasets,
suggesting that although the methods rely on diferent principles, they capture similar global structure.
In contrast, the lowerstta values ≈( 0.500) appear betweenGeolet andtf–idf in real-world datasets,
reflecting a fundamental diference between local shapelet patterns and symbolic discretization. These
relationships highlight the complementarity among vectorization families. From a practical perspective,
the reporterdta scores suggest promising research avenues for ensemble or hybrid vectorizations.</p>
      <p>To qualitatively observe such complementarity, we now examine representative embedding spaces
for the selected datasets. Fig5uprreesents 2D PCA projections ftoirf, Geolet, andRoSITa (left to
right). For thSeR dataset (top row), all three embedding spaces clearly separate instances by their
true clusters. FoFSr (bottom row), where classes difer only in movement speteidf, yields a clean
separation, whereas the 2D projection ofRtohSeITa space collapses the instances into a single point.
Nonetheless, its FM score is nearly optimal, indicating that the collapse is a visual artefact of PCA.
The Geolet projection also appears scattered randomly. However, its FM score of 0.705 suggests that,
despite the imperfect visualization, meaningful separation persists in this balanced binary setting.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Conclusions</title>
      <p>We introduced a benchmark framework for evaluating alternative trajectory vectorizations for unsu
pervised clustering. By comparing representative vectorization families on real-world and synthetic
datasets, we analyzed how sub-trajectory dynamics and geometry are reflected in diferent vector spaces.
The results show that vectorization choice substantially afects clustering, that methods performing well
under controlled conditions may not generalize to noisy real-world data, and that diferent vectorizations
capture complementary (not redundant) information, as indicated by mrotdaevraaltuees. We also
observe that relying on a single metric (e.g., SIL or FM) can be misleading, whereas combining internal
and external criteria supports more robust model selection. At a finer level, feature-based methods are
efective for motion dynamics, shapelet-based methods capture local geometric variatRiooSnI,Taand
performs strongly on geometry-dominated datasets. These findings motivate future work on
multivectorization ensembles, fusion strategies, and broader evaluation across more datasets, vectorizations,
and clustering algorithms.</p>
    </sec>
    <sec id="sec-9">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author(s) used ChatGPT-4 and Grammarly to perform grammar
and spelling checks. After using these tools, the authors reviewed and edited the content as needed and
take full responsibility for the publication’s content.</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>This work has been partly supported by the University of Piraeus Research Center.</p>
      <p>GeoInformatica 29 (2025) 999–1032.
[25] D. Yao, C. Zhang, Z. Zhu, J. Huang, J. Bi, Trajectory clustering via deep representation learning,
in: IEEE IJCNN, 2017, pp. 3880–3887.
[26] F. Schrof, D. Kalenichenko, J. Philbin, Facenet: A unified embedding for face recognition and
clustering, in: Proceedings of the IEEE conference on computer vision and pattern recognition,
2015, pp. 815–823.
[27] Y. Wang, H. Huang, C. Rudin, Y. Shaposhnik, Understanding how dimension reduction tools work:
An empirical approach to deciphering t-sne, umap, trimap, and pacmap for data visualization, J.
Mach. Learn. Res. 22 (2021) 201:1–201:73.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Dogra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <article-title>Trajectory-based surveillance analysis: A survey</article-title>
          ,
          <source>IEEE Trans. Circuits Syst. Video Technol</source>
          .
          <volume>29</volume>
          (
          <year>2019</year>
          )
          <fpage>1985</fpage>
          -
          <lpage>1997</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hamdi</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. B. Shaban</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Erradi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mohamed</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          <string-name>
            <surname>Rumi</surname>
            ,
            <given-names>F. D.</given-names>
          </string-name>
          <string-name>
            <surname>Salim</surname>
          </string-name>
          ,
          <article-title>Spatiotemporal data mining: a survey on challenges and open problems</article-title>
          , Artif.
          <source>Intell. Rev</source>
          .
          <volume>55</volume>
          (
          <year>2022</year>
          )
          <fpage>1441</fpage>
          -
          <lpage>1488</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Musleh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Mokbel</surname>
          </string-name>
          ,
          <article-title>Let's speak trajectories: A vision to use NLP models for trajectory analysis tasks</article-title>
          ,
          <source>ACM Trans. Spatial Algorithms Syst</source>
          . (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>M. F. M.</surname>
          </string-name>
          et. al.,
          <article-title>Towards mobility data science (vision paper)</article-title>
          ,
          <source>CoRR abs/2307</source>
          .05717 (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C. L.</given-names>
            da
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Petry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bogorny</surname>
          </string-name>
          ,
          <article-title>A survey and comparison of trajectory classification methods</article-title>
          ,
          <source>in: BRACIS</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>788</fpage>
          -
          <lpage>793</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Landi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Spinnato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monreale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nanni</surname>
          </string-name>
          ,
          <string-name>
            <surname>Geolet:</surname>
          </string-name>
          <article-title>An interpretable model for trajectory classification</article-title>
          ,
          <source>in: IDA, volume 13876Leocfture Notes in Computer Science</source>
          , Springer,
          <year>2023</year>
          , pp.
          <fpage>236</fpage>
          -
          <lpage>248</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <article-title>Trajectory data mining: An overview</article-title>
          ,
          <source>ACM Trans. Intell. Syst. Technol</source>
          .
          <volume>6</volume>
          (
          <year>2015</year>
          )
          <volume>29</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>29</lpage>
          :
          <fpage>41</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Middlehurst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ismail-Fawaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guillaume</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Holder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guijo-Rubio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Bulatova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tsaprounis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mentel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Walter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Bagnall</surname>
          </string-name>
          ,
          <article-title>aeon: a python toolkit for learning from time series</article-title>
          ,
          <source>CoRR abs/2406</source>
          .14231 (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Etemad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Júnior</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Matwin</surname>
          </string-name>
          ,
          <article-title>Predicting transportation modes of GPS trajectories using feature engineering and noise removal</article-title>
          ,
          <source>in: Canadian Conference on AI</source>
          , volume
          <volume>10832</volume>
          ,
          <year>2018</year>
          , pp.
          <fpage>259</fpage>
          -
          <lpage>264</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Keogh</surname>
          </string-name>
          ,
          <article-title>Time series shapelets: a novel technique that allows accurate, interpretable and fast classification, Data Min</article-title>
          .
          <source>Knowl. Discov</source>
          .
          <volume>22</volume>
          (
          <year>2011</year>
          )
          <fpage>149</fpage>
          -
          <lpage>182</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Ferrero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. O.</given-names>
            <surname>Alvares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zalewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bogorny</surname>
          </string-name>
          ,
          <article-title>MOVELETS: exploring relevant subtrajectories for robust trajectory classification</article-title>
          ,
          <source>in: SAC</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>849</fpage>
          -
          <lpage>856</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Guibas</surname>
          </string-name>
          ,
          <article-title>Pathlet learning for compressing and planning trajectories</article-title>
          , in: SIGSPATIAL/GIS, ACM,
          <year>2013</year>
          , pp.
          <fpage>382</fpage>
          -
          <lpage>385</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Landi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <article-title>Shape-based methods in mobility data analysis: efectiveness and limitations</article-title>
          ,
          <source>GeoInformatica</source>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <article-title>A survey on trajectory clustering analysis</article-title>
          ,
          <source>arXiv</source>
          (
          <year>2018</year>
          ). arXiv:
          <year>1802</year>
          .06971.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <article-title>GonzalezT,raClass: trajectory classification using hierarchical region-based and trajectory-based clustering</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>1</volume>
          (
          <year>2008</year>
          )
          <fpage>1081</fpage>
          -
          <lpage>1094</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Parent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Spaccapietra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Renso</surname>
          </string-name>
          , G. Andrienko,
          <string-name>
            <given-names>N.</given-names>
            <surname>Andrienko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bogorny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Damiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gkoulalas-Divanis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>F. de Macêdo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pelekis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Theodoridis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Semantic trajectories modeling and analysis</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>45</volume>
          (
          <year>2013</year>
          )
          <volume>42</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>42</lpage>
          :
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>I.</given-names>
            <surname>Ben-Gal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Weinstock</surname>
          </string-name>
          , G. Singer,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bambos</surname>
          </string-name>
          ,
          <article-title>Clustering users by their mobility behavioral patterns</article-title>
          ,
          <source>ACM Transactions on Knowledge Discovery from Data</source>
          <volume>13</volume>
          (
          <year>2019</year>
          )
          <volume>45</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>45</lpage>
          :
          <fpage>28</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Landi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nanni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monreale</surname>
          </string-name>
          ,
          <article-title>The trajectory interval forest classifier for trajectory classification</article-title>
          , in: SIGSPATIAL/GIS, ACM,
          <year>2023</year>
          , pp.
          <volume>67</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>67</lpage>
          :
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Landi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Andrienko</surname>
          </string-name>
          , G. Andrienko,
          <article-title>Rotation‑ and scale-invariant shape extraction from vessel trajectories for human‑in‑the‑loop monitoring</article-title>
          , in: SIGSPATIAL/GIS, ACM,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zakaria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mueen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Keogh</surname>
          </string-name>
          ,
          <article-title>Clustering time series using unsupervised-shapelets</article-title>
          , in: IEEE Computer
          <string-name>
            <surname>Society</surname>
            <given-names>ICDM</given-names>
          </string-name>
          ,
          <year>2012</year>
          , pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Keogh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lonardi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Experiencing</surname>
            <given-names>SAX</given-names>
          </string-name>
          :
          <article-title>a novel symbolic representation of time series, Data Min</article-title>
          .
          <source>Knowl. Discov</source>
          .
          <volume>15</volume>
          (
          <year>2007</year>
          )
          <fpage>107</fpage>
          -
          <lpage>144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>I. S.</given-names>
            <surname>Suwardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Satya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Lestari</surname>
          </string-name>
          ,
          <article-title>Geohash index based spatial data model for corporate</article-title>
          ,
          <source>in: ICEEI</source>
          <year>2015</year>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>478</fpage>
          -
          <lpage>483</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Vagena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Athitsos</surname>
          </string-name>
          ,
          <article-title>Rotation invariant indexing of shapes and line drawings</article-title>
          , in: CIKM, ACM,
          <year>2005</year>
          , pp.
          <fpage>131</fpage>
          -
          <lpage>138</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>C.</given-names>
            <surname>Landi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <article-title>Shape-based methods in mobility data analysis: efectiveness and limitations,</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>