<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>How to Implement your Trajectory Data Warehouse? Design and Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rim Moussa</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandro Bimonte</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Wrembel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>INRAE</institution>
          ,
          <addr-line>Clermont-Ferrand</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Poznan University of Technology</institution>
          ,
          <addr-line>Poznań</addr-line>
          ,
          <country country="PL">Poland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Carthage</institution>
          ,
          <addr-line>Carthage</addr-line>
          ,
          <country country="TN">Tunisia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>The management and analysis of massive trajectory data present significant challenges in storage scalability and computational eficiency. In this paper, we introduce TDWbench - a benchmark designed to assess and compare two alternative trajectory data warehouse schema designs, i.e., point-based and cell-based. TDWbench is based on a large-scale open maritime trajectory dataset provided by the Danish Maritime Authority The benchmark allows assessing trade-ofs between storage eficiency, query performance, and query accuracy. We provide experimental results that demonstrate how TDWbench supports the comparison of trajectory data warehouse models with realistic analytical workloads, for diferent spatial resolutions and scale factors.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;trajectory data warehouse</kwd>
        <kwd>logical data warehouse schema</kwd>
        <kwd>big data</kwd>
        <kwd>OLAP</kwd>
        <kwd>benchmark</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        In this section, we provide an overview of the related work spanning two complementary areas of
research: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) TDW designs, which extend traditional data warehousing concepts to manage and analyze
spatio-temporal data representing the movement of objects and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) TDW benchmarks.
      </p>
      <p>TDW designs. In [3], the authors present the design and implementation of a TDW called
TWAREHOUSE. T-WAREHOUSE integrates all the necessary components for TDW management,
encompassing trajectory reconstruction, ETL processing, and visual OLAP. In [4] the conceptual schema of a
segment-based TDW is proposed, where the fact table contains trajectory segments and their attributes
(like the geometry of a segment route, a distance traveled, speed, and duration). In [5], the authors use
Cartesian grid systems with fixed x-y dimensions to map a regular lattice onto a spherical globe. [ 6]
focuses on and end-to-end method for building a TDW, from data pre-processing, TDW loading, to
building analytical solutions.</p>
      <p>
        TDW benchmarks. BerlinMOD [7] is a benchmark based on human mobility data. It is
parameterized by: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) the number of people and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) the length of observation period; both impact the scale
factor. BerlinMOD defines 17 types of queries. [ 8] surveys spatial analytics systems and conducted
experiments to compare five Spark-based systems on 5 diferent spatial queries (range, kNN, spatial
joins between various geometric datatypes, distance join, and kNN join) and 4 diferent datatypes
(points, linestrings, rectangles, and polygons). Their evaluation uses an Open Street Maps dataset of 500
million rows ( 56.4GB). MobilityDB [9] is an open-source MOD that provides abstract data types and
operations for managing mobility data in PostgreSQL and PostGIS. MobilityDB features an extensive set
of spatiotemporal operations (e.g., distance, temporal predicates, range predicates). The authors assessed
the performance of MobilityDB using BerlinMOD. Distributed query capabilities in MobilityDB were
demonstrated on a data set of 2 billion AIS ship trajectory points (500 GB), obtained from the Danish
Maritime Authority. They utilized Citus—a Distributed PostgreSQL for Data-Intensive Applications
[10]. The logical schema contained two tables: Ship and Port. The authors reported the performance
of four analytical queries, all of which include a spatial join with specific ports and highly selective
temporal filtering. In [ 11], the authors proposed an open-source, anonymized, metropolitan-scale
dataset (YJMob100K) of 100,000 human mobility trajectories within 90 days, from Yahoo Japan
Corporation. Location pings are mapped into 500 × 500 meter grid cells, and timestamps are aggregated into
30-minute intervals.
      </p>
      <p>To conclude, benchmarks for spatial data storage and querying systems are essential for identifying
the most appropriate logical and physical designs. However, to the best of our knowledge, no existing
work proposes a fully-featured, ad hoc benchmark specifically tailored to TDWs. Such a benchmark
would need to encompass multiple design schemas, varying scale factors, query workloads, and relevant
performance metrics.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Benchmark Description</title>
      <p>This section introduces TDWbench1 designed to compare spatio-temporal data models and systems that
manage trajectory data. The benchmark is constructed from a real-world dataset of maritime voyages
(Section 3.1) and enables the generation of TDWs with varying data volumes, in a manner analogous to
well-known TPC benchmarks. The key characteristics of TDWbench are discussed below.
• Scalability: the scale factor of the DW is represented by length of the observation period (the number
of days).
• Schema comparison: the benchmark supports a direct comparison between point-based and cell-based
schemas. For the latter, the resolution parameter defines cell granularity.
• Workload: TDWbench includes a representative analytical workload comprising common analytical
operations such as grouping, joins, and filtering, reflecting realistic trajectory analytical queries.
1https://github.com/rimmmmm/TDWbench
• Performance and accuracy reporting: the benchmark provides a comprehensive evaluation of both
performance and the quality or query results measured by accuracy. The performance metric includes
query execution time, while the accuracy metric quantifies the accuracy loss induced by spatial
aggregation and discretization, enabling an explicit trade-of analysis between eficiency and analytical
precision.
• Implementation: TDWbench is implemented using relational DBMS PostgreSQL v.11 and its spatial
extension PostGIS. Such an implementation does not compromise its genericity, as deploying the
benchmark on another relational DBMS is generally a straightforward task.</p>
      <sec id="sec-3-1">
        <title>3.1. Data Sources Description</title>
        <p>
          We consider three data sources for building the TDW, namely: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) Vessels navigation AIS logs [12]; (
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
Marine geometry data; [13] (3) World port index (WPI) [14]. The AIS logs dataset does not explicitly
contain trips. To obtain the navigation attributes such as departure and arrival ports, it is necessary to
join the AIS data with the spatial coordinates from the WPI database.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Point-based schema</title>
        <p>In the point-based schema, a trajectory  is defined as:  = ⟨1, 2, . . . , ⟩, where each position  is
a spatial-temporal point:  = ⟨, , ⟩ such that: ,  ∈ R represent the geographical coordinates
(longitude and latitude);  ∈ R represents the timestamp of the location (1 &lt; 2 &lt; · · · &lt;  ).</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Cell-based schema</title>
        <p>In the cell-based schema, each point  is mapped to a spatial cell  represented as a polygon. A
trajectory is represented as:  = ⟨1, 2, . . . , ⟩ where  = ⟨, _, _⟩, where
 is a polygon describing the boundaries of the cell containing sequential positions; _
denotes the timestamp corresponding to the vessel’s entry into the cell ; and _ denotes the
timestamp corresponding to the vessel’s departure from the cell .</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Point-based vs. cell-based: resolution</title>
        <p>On the one hand, points representing a trajectory result in the most accurate trajectory, whereas cells
ofer an approximate trajectory. On the other hand point-based storage occupies a larger disk space as
compared to cell-based, and as a consequence queries on a point-based DW schema are more expensive.
Moreover, the size of a cell plays a crucial role in the trade-of between accuracy of a query result, data
storage, and query execution time. Figures 1 and 2 illustrate a trip trajectory represented as a sequence
of points (red curve) alongside its corresponding spatial cell polygons (blue cells). RES2 resolution
illustrates a finer spatial granularity compared to RES1 resolution.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Workload</title>
        <p>
          The analytical workload consists of 12 batches (see Table 1) of 7 queries each. Queries either include
ifltering on the fact tables or incorporate temporal or spatial selectivity predicates. Three spatial
selectivity values (6%, 18%, 36%) and three temporal selectivity values (8%, 16%, 33%) are considered,
resulting in a total of 84 queries. The proposed workload is designed to: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) highlight the diferences in
query processing performance across various DW schemas, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) incorporate complex query patterns, and
(3) address analytical queries that are relevant to maritime transportation. A query can be composed of
the following operations: join, grouping, aggregation, and filtering.
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6. Metrics</title>
        <p>The main metrics we use in TDWbench are: storage size, time performance, and query accuracy, for a
given cell resolution. The time performance is measured as the total runtime for a given TDW volume.
The query accuracy is a new metric that quantifies the deviation of approximate query results on
cell-based schemata from the exact results obtained on the point-based schema.</p>
        <p>To evaluate the query accuracy we use the Root Mean Squared Error (RMSE). For a given measure m,
RMSE is computed using the value of m for the cell-based model  () and the value of m for the
point-based model  (), for all the values of m output by a query. The formula is the following:
⎯
=1</p>
        <p>RMSE = ⎷⎸⎸ 1 ∑︁ (︀  () −  ( ))︀ 2</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. TDWbench deployment</title>
      <p>In this section, we report some experiments to show how TDWbench supports comparing TDW
implementations according to diferent parameters: the type of a model, the scaling factor, the resolution,
and queries.</p>
      <p>
        The trajectory data warehouse was implemented in PostgreSQL v11. The TDW stored AIS logs from
Jan 2023, in two cell resolutions: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) RES1-Cell with approximate dimensions of ∼ 4.9 km ×4.9 km (at
the equator) and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) RES2-Cell with approximate dimensions of ∼ 1.2 km ×0.6 km (at the equator). The
TDW was run on a node on the French national grid computing infrastructure - GRID50002. The node
was equipped with Intel Xeon E5-2620 v4 (2.1 GHz), 16 cores (2×8), and 64 GB RAM node.
Storage: As expected, dimension (e.g., vessels) sizes increase with the scaling factor, but they do not
change according to the schema used, since they do not contain trajectory data. The schemas afect the
storage of fact tables: TripPoint, RES1 Cell polygon, and RES2 Cell polygon.
      </p>
      <p>The point-based schema (TripPoint) needs the highest storage volume since it stores individual vessel
positions with the highest spatial and temporal resolution (i.e., 35 GB for SF=30). The cell-based schema
reduces the storage size according to the resolution used. RES2 has a higher resolution than RES1,
therefore the fact table for RES1 occupies less space than that of RES2 (i.e., 3 GB vs 5.5 GB for SF=30).
Runtime: Figure 3 presents the aggregated runtime results per schema for the two evaluated scale
factors. The cell-based schema provides substantial performance benefits. Figure 4 presents the runtime
distribution across all batches. Batches that include a spatial join consistently show higher runtime.
l]e103
a
c
s
g
o
l
[
)
n
i
(m102
e
m
i
t
n
u
R
l
tToa101
38
point
5,086</p>
      <p>212
2</p>
      <sec id="sec-4-1">
        <title>RES1-Cell</title>
      </sec>
      <sec id="sec-4-2">
        <title>Schema &amp; Scale Factor</title>
        <p>ppRooEiiSnn1tt-((CSSeFFl==l(13S)0F)=1) RRREEESSS221---CCCeeellllll (((SSSFFF===133)00))
491
5</p>
      </sec>
      <sec id="sec-4-3">
        <title>RES2-Cell</title>
        <p>Accuracy: TDWbench allows also to compare the accuracy of query results, represented by the RMSE.
An example is shown in Figure 5 that details the accuracy for each query of Batch 7 and the two
resolutions. As expected, the more fine-grained cell size, the lower RMSE is. The diferences in the
RMSE values between RES1 and RES2 for the same query can be substantial, see SpatialFilter of 6% and
18%.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <p>
        In this paper we presented TDWbench - a benchmark for evaluating TDWs. Built upon over a 2TB
trajectory dataset from the Danish Maritime Authority, TDWbench implements a scaled dataset and a
representative analytics workload to facilitate a detailed investigation of the trade-ofs between: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
storage size (not reported here due to space limit), (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) query performance, and (3) query accuracy. Its
experimental results validate the benchmark utility. We are currently extending this research through
103
l]
e
a
c
s
g
l[)o102
n
i
m
(
e
m
it
lunR101
a
t
o
T
100
      </p>
      <p>1527.27
100
l]
d
scoaeg
l[
E
S
RM10 1
experimental work on alternative DW schemas for storing spatial data and on in-depth analysis of their
impact on query accuracy. Further works with focus on evaluating all 84 queries of TDWbench, and
evaluating alternative storage systems, like MobilityDB.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgement</title>
      <p>The work of S. Bimonte is supported by the CHIST-ERA grant ANR-24-CHR4-0004-0 ’GIS4IoRT’.
The work of R. Wrembel is supported from the National Science Centre (NCN), Poland, grant no.
2024/06/Y/ST6/00136, originally funded from the EU project Chist-Era call 2023.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>The author(s) have not employed any Generative AI tools.
[3] L. Leonardi, G. Marketos, E. Frentzos, N. Giatrakos, S. Orlando, N. Pelekis, A. Rafaetà, A. Roncato,
C. Silvestri, Y. Theodoridis, T-warehouse: Visual OLAP analysis on trajectory data, in: Int. Conf.
on Data Engineering (ICDE), IEEE, 2010.
[4] C. Renso, S. Spaccapietra, E. Zimányi, Mobility Data: Modeling, Management, and Understanding,</p>
      <p>Cambridge University Press, 2013.
[5] L. Wu, Y. Xu, Q. Wang, F. Wang, Z. Xu, Mapping global shipping density from ais data, Journal of</p>
      <p>Navigation 70 (2017).
[6] G. Marketos, E. Frentzos, I. Ntoutsi, N. Pelekis, A. Rafaetà, Y. Theodoridis, Building real-world
trajectory warehouses, in: ACM Int. Workshop on Data Engineering for Wireless and Mobile
Access (Mobide), 2008, pp. 8–15.
[7] C. Düntgen, T. Behr, R. H. Güting, Berlinmod: a benchmark for moving object databases, VLDB</p>
      <p>Journal 18 (2009).
[8] V. Pandey, A. Kipf, T. Neumann, A. Kemper, How good are modern spatial analytics systems?,</p>
      <p>VLDB Endowment 11 (2018).
[9] M. S. Bakli, M. A. Sakr, E. Zimányi, Distributed mobility data management in MobilityDB, in: Int.</p>
      <p>Conf. on Mobile Data Management, IEEE, 2020.
[10] U. Cubukcu, O. Erdogan, S. Pathak, S. Sannakkayala, M. Slot, Citus: Distributed postgresql for
data-intensive applications, in: Int. Conf. on Management of Data (SIGMOD), 2021.
[11] T. Yabe, K. Tsubouchi, T. Shimizu, Y. Sekimoto, K. Sezaki, E. Moro, A. Pentland, Yjmob100k:</p>
      <p>City-scale and longitudinal dataset of anonymized human mobility trajectories (version 2), 2023.
[12] Danish Maritime Authority makes historical AIS data available, 2018. URL: https://safety4sea.com/
danish-maritime-authority-publishes-historical-ais-data/, accessed Jan, 2026.
[13] Natural earth, 2025. URL: https://www.naturalearthdata.com/downloads/, accessed Jan, 2026.
[14] National Geospatial Intelligence Agency. World Port Index, 2019. URL: https://msi.nga.mil/
Publications/WPI, accessed Jan, 2026.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Vaisman</surname>
          </string-name>
          , E. Zimányi,
          <article-title>Trajectory data warehouses</article-title>
          , in: C.
          <string-name>
            <surname>Renso</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Spaccapietra</surname>
          </string-name>
          , E. Zimányi (Eds.),
          <source>Mobility Data: Modeling</source>
          , Management, and
          <string-name>
            <surname>Understanding</surname>
          </string-name>
          , Cambridge University Press,
          <year>2013</year>
          , pp.
          <fpage>62</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>J. von Kistowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Arnold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Huppler</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-D. Lange</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          <string-name>
            <surname>Henning</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Cao</surname>
          </string-name>
          ,
          <article-title>How to build a benchmark</article-title>
          ,
          <source>in: ACM/SPEC Int. Conf. on Performance Engineering</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>