<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Teaching Query-Driven Design in Aggregate-Oriented NoSQL Systems: Resources, Methodology, and Tool Support</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Barbara Catania</string-name>
          <email>barbara.catania@unige.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanna Guerrini</string-name>
          <email>giovanna.guerrini@unige.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amer Al Khoury</string-name>
          <email>amer.alkhoury89@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DIBRIS - University of Genoa</institution>
          ,
          <addr-line>Genoa -</addr-line>
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>Teaching data design in aggregate-oriented NoSQL systems (ANoSQL) is a core topic in Advanced Data Management courses. A key challenge for students is understanding the main diferences between SQL and ANoSQL data design approaches. SQL systems are typically used for integration-oriented databases with fully structured, normalized schemas, and a clear separation between logical and physical levels. In contrast, ANoSQL systems target application-oriented databases and favor flexible, denormalized schemas, optimized for workload-specific access patterns. In these systems, logical and physical levels are tightly integrated to minimize the number of nodes accessed during query execution. Unfortunately, existing ANoSQL design approaches are largely system-specific, often overlook this interplay, and ofer limited reusability in educational contexts. Based on our experience teaching an Advanced Data Management course at the University of Genoa since 2016, we have therefore developed specific educational resources to address this gap and support students in understanding ANoSQL design principles and eficiency considerations. The resources include lesson plans, slides with methodological content and examples, and an optional formative tool. They emphasize query-driven design, clarify the logical-physical interplay, introduce a simple system-independent methodology for generating ANoSQL logical schemas enriched with physical metadata, and provide guidance for translating them into specific ANoSQL systems such as MongoDB and Cassandra.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;aggregate-oriented NoSQL data store</kwd>
        <kwd>query-based design</kwd>
        <kwd>logical-physical design methodology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Teaching an Advanced Data Management (ADM) course in Master’s programs in Computer Science,
Engineering, and related disciplines has become increasingly challenging due to evolving market
expectations and the rapid evolution of DBMS technologies. ADM learning outcomes span competencies
ranging from large-scale data management principles and NoSQL system architectures to data design
and management for AI-driven applications. Given their wide applicability and eficiency in large-scale
distributed settings, aggregate-oriented NoSQL (ANoSQL) systems (key-value, document-based, and
column-family data stores) have become a core topic in ADM courses. All ANoSQL systems rely on the
unifying concept of aggregate [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a self-contained unit of data grouping information accessed together
and managed as a single logical and physical unit.
      </p>
      <p>
        Due to their distributed nature, one of the main dificulties when learning ANoSQL systems is data
design, since architectural choices and data management principles difer substantially from relational
ones. From a teaching perspective, these diferences require educational material that explicitly addresses
the impact of distribution on data modeling, considering the following key factors [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
      </p>
      <p>
        K1. Integration and application databases. Traditional relational systems rely on a single, integrated
database with a shared, normalized schema serving multiple applications, with constraints centrally
enforced by the DBMS. In contrast, large-scale systems often adopt application-specific databases, where
each application owns its data and selects the schema and technology best suited to its workload, moving
part of constraint management to the application [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This architectural change must be explicitly
reflected in the resources employed in ANoSQL data modeling education.
      </p>
      <p>
        K2. Moving towards cluster computing. Large data volumes typical of data-intensive applications
are stored across clusters, where data are partitioned and replicated over multiple nodes. Partitioning
enables intra-query parallelism and requests are executed concurrently on diferent data partitions and
nodes. As a consequence, data design decisions directly afect query execution and performance and
represent a key learning challenge in ANoSQL education [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ].
      </p>
      <p>
        To efectively learn ANoSQL data design, students thus must rethink traditional database design
principles in light of these characteristics. However, the lack of a reference standard for ANoSQL data
modeling and querying makes this task quite challenging [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ]. Existing methodologies are often
specific systems and thus not easily generalizable [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ]; even approaches that abstract from
systemdependent features typically defer decisions on data distribution and storage to later, system-specific
steps [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This limits their efectiveness as reusable teaching artifacts.
      </p>
      <p>Based on these considerations, we propose educational resources supporting ANoSQL data design
teaching through a simple system-independent methodology that integrates logical and physical
requirements and emphasizes abstraction. The resources consist of modular teaching materials (lesson plans,
slides, and worked examples) organized into four methodological slide blocks and one exercise block
with explicit learning objectives, complemented by a lightweight formative tool supporting self-directed
learning. The aim is to provide reusable materials that clarify core ANoSQL data design principles and
support students in reasoning about design choices and their implications, rather than to empirically
evaluate the corresponding learning outcomes.</p>
      <p>The remainder of this paper is organized as follows. Section 2 presents the educational context and
the proposed resources, with guidelines for their reuse. Sections 3 and 4 detail the system-independent
ANoSQL data design methodology at the logical and physical levels, highlighting the related educational
purposes. The translation of the system-independent schemas into system-specific schemas is then
discussed in Section 5. Finally, Section 6 discusses related work while Section 7 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Educational Context and Resources</title>
      <p>Educational Context The proposed artifacts were developed based on our experience in teaching
a graduate course on Advanced Data Management at the University of Genoa since 2016, within
the Master’s programs in Computer Science and Computer Engineering. The course enrolls about 50
students per year (43 in 2025/2026: 37% Computer Science, 44% Computer Engineering, 19% international
programs). Students are expected to have prior knowledge of relational database design, SQL, and basic
database system concepts.</p>
      <p>
        The course covers advanced data management topics, focusing on large-scale processing for
dataintensive applications and ANoSQL data stores. A central learning objective is ANoSQL data design,
requiring students to rethink traditional database design principles in light of large-scale system
characteristics. The main ANoSQL design principles that students are expected to master are the
following [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
      </p>
      <p>P1. Query-based design and nested aggregates. In distributed settings, join operations incur high
communication costs. As a result, ANoSQL logical design relaxes normalization and groups attributes
that are frequently accessed together into a single aggregate, enabling query execution without
crosscollection joins and leading to nested data structures.</p>
      <p>P2. Many ANoSQL data models, one single unifying concept: the aggregate. Despite diferences among
ANoSQL data models, all systems are built around the notion of aggregate, abstractly represented as
a (, ) pair. Models difer in how much of the  can be accessed and processed by the
system. For example, in key-value stores, the  is opaque and can only be processed by applications;
in column-family stores, only first-level components are system-accessible while in document-based
systems the  is fully accessible.</p>
      <p>P3. Interplay of the logical and physical levels. Unlike relational systems, ANoSQL systems tightly
couple logical and physical design. The workload drives both aggregate design and their physical
placement across cluster nodes through the choice of a partition key, which should co-locate aggregates
frequently accessed together to minimize cross-node communication and ensure predictable
performance. Moreover, index selection remains challenging, as indexing strategies must align with both
aggregate structures and data partitioning to eficiently support the workload.</p>
      <p>
        Resource Organization Efective ANoSQL design education should build on these three key
principles. Our review of existing ANoSQL design approaches (see Section 6) shows that all methodologies
adopt a query-driven process based on nested aggregates [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] (P1). However, most are tied to specific
ANoSQL systems or models and are not easily generalizable [
        <xref ref-type="bibr" rid="ref5 ref6 ref7 ref8 ref8 ref9">5, 6, 7, 8, 8, 9</xref>
        ], while a few abstract from
system-dependent features and define aggregate schemas at a more abstract level centered on the notion
of aggregate [
        <xref ref-type="bibr" rid="ref10 ref11 ref4">4, 10, 11</xref>
        ] (P2), from which system-specific schemas are derived through a translation step.
In most cases, physical design aspects, such as partition-key and index selection, are addressed in a
system-specific manner [
        <xref ref-type="bibr" rid="ref4 ref7 ref9">4, 7, 9</xref>
        ]. We argue that postponing physical design aspects in an educational
context may hinder students’ early understanding of performance implications and of the tight coupling
between logical and physical design choices (P3). At the same time, the notion of aggregate provides
an efective abstraction for a system-independent methodology, capturing workload-driven principles
while avoiding unnecessary system-specific complexity. Starting from these considerations, we
developed ANoSQL design resources that integrate logical and physical design within a system-independent
methodology to support coherent and efective learning. The proposed material is organized into four
main blocks of slides:
Block 1 serves as an introduction, motivating query-driven design and nested aggregates (P1) and
justifying the need for a system-independent ANoSQL design methodology (P2).
      </p>
      <p>Block 2 introduces a simple yet efective model for specifying aggregate schemas in a system-independent
way, grounded in existing standards like JSON Schema, and presents a lightweight methodology
for their generation, according to (P1) and (P2).</p>
      <p>Block 3 concerns physical design and its interplay with logical issues (P3). It is used to introduce a
reference notion of query eficiency for ANoSQL systems, upon which a system-independent
methodology for partition-key selection is provided, together with educational material for
understanding how unique constraints can be enforced in ANoSQL systems and why and when
index creation matters. Physical design choices are then represented inside aggregate schemas,
extending the model presented in Block 2.</p>
      <p>Block 4 shows how an aggregate schema, including logical and physical information, can be translated
into concrete schemas for key-value, document-based, and column-family systems, assuming
systems and data models have already been presented in the course.</p>
      <p>All concepts are introduced through a single running example consistently used across the four
slide blocks. A final Block 5 provides alternative examples covering all introduced design concepts.
Blocks 1–4 are intended for frontal lectures (approximately 2 hours for Block 1 and 4 hours for Blocks 2–
4), optionally integrated with instant polling, or delivered in a flipped-classroom setting, particularly
for Block 4. Block 5 is designed for exercise sessions (2 hours for each of Blocks 2–4), which may be
conducted collaboratively. The blocks are designed to take approximately 16—20 hours in total.1</p>
      <p>The slide blocks assume prior knowledge of basic concepts of distributed data management.
Depending on course organization, Block 1 can be omitted; Block 3 must follow Block 2, while Block 4 is
intended for courses covering specific ANoSQL systems. To complement Blocks 2 and 4, an optional
lightweight self-study tool is also available.2 Developed in JavaScript, it supports ER diagram and
1All slide blocks are available as PDF files in the github project available at https://github.com/gioguerrini/
ANoSQL-Methodology.
2The tool is available at https://amazing-benz-6643f3.netlify.app/.
workload specification, generates logical aggregate schemas as JSON schemas, and translates them into
Cassandra commands. As it is based on an earlier version of the methodology, it covers only logical
design. For this reason, it is not further discussed in the next sections that focus instead on Blocks 2–4,
representing the main contribution of this work.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Block 2- Logical Design</title>
      <p>
        Block 2 introduces the core artefact for logical ANoSQL design. It guides students in deriving
aggregateoriented logical schemas from a conceptual diagram and a query workload, emphasizing query-driven
design while remaining system-independent. The proposed methodology simplifies existing approaches
(e.g., [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), builds on students’ prior database design knowledge, and introduces only a minimal set of
new concepts. It is heuristic-based, leaving room for alternative design choices to be refined when
moving from system-independent schemas to system-specific ones (Block 4), encouraging students to
reason about design trade-ofs.
      </p>
      <p>Given an Entity-Relationship (ER) diagram and a workload in natural language, the methodology
produces a set of aggregate-oriented logical schemas, represented in a JSON-inspired schema
metanotation. To reduce cognitive load, the focus is exclusively on queries with equality-based selection
conditions. The process consists of three steps and three main heuristics, described below.
Step 1: Query modeling Each query is represented in a formal and unambiguous way, helping
students clarify which entities and attributes are accessed. More precisely, each workload query Q is
represented as a triple Q(E, LS, LP), where: (i) E is the aggregation entity used to group all selection
and projection attributes accessed by Q; (ii) LS (LP) is the list of entities containing selection (projection)
attributes. For each entity E in LS or LP, the representation specifies the E attributes used in the query
and the path linking E to E.3 Paths are expressed as sequences of association names (or initials, when
unambiguous). When E coincides with E, the empty path is denoted by !.</p>
      <p>Heuristics. To ensure a general, system-independent design (P2), we introduce heuristic H1 for selecting
the aggregation entity E. H1 favors aggregate schemas in which selection attributes appear as simple
(non-nested) properties, reducing selections on nested attributes that may be ineficient or unsupported
in some ANoSQL systems. Accordingly, E is chosen among the entities in LS, or, if none exist, among
those in LP; when multiple candidates are available, preference is given to the entity on the one-side of
the highest number of associations linking selection entities.</p>
      <p>Step 2: ER schema annotation The second step annotates the ER diagram with workload
information, starting from the result of Step 1. For each query Q(E, LS, LP): (i) entity E is marked with label
Q and represented using a double rectangle; Q is referred to as the query associated with entity E; (ii)
for each entity E in LS or LP, the path linking E to E is annotated with arrows directed toward E and
labeled with Q; (iii) attributes involved in selections and projections are annotated with labels Q-S and
Q-P, respectively.</p>
      <p>Step 3: Aggregate schema generation During the last step, students learn how to transform the
annotated ER diagram into a set of system-independent aggregate schemas. For each aggregation entity
E, according to (P1), one logical schema is derived by considering all selection and projection attributes
used in the queries associated with E. The resulting schemas are represented in a JSON-inspired
Schema Meta-notation (JSON-SM), extended to capture ANoSQL-specific design information. JSON-SM
represents the high-level aggregate structure by retaining property names and nesting relationships,
and supports the explicit representation of identifiers. Complex attributes are expressed using curly
brackets {} for nested objects and square brackets [] for collections. The notation intentionally departs
from JSON and JSON Schema standard: JSON-SM is not meant to be parsed or validated, but serves as a
3Listed only if they do not correspond to the whole entity attribute set.
(a) ER diagram for the library book rental domain.
Entities are shown as rectangles, associations as
diamonds, attributes as labeled circles. Cardinality
constraints indicate minimum and maximum
participation as pairs of values (min,max). Black
circles denote identifiers, which may combine entity
attributes with identifiers of related entities (e.g.,</p>
      <p>Rental).</p>
      <p>Q1 Determine the aver- Q1(Reader,[ ],</p>
      <p>age age of readers [Reader(birthdate)_!])
Q2 For each reader, Q2(Reader,[ ],
determine the [Reader(name,surname)_!,
name, surname, and Book(title,authors)_R])
related rated books
Q3 Determine the name Q3(Book,
and the surname of [Book(title,authors)_!],
readers who rated [Reader(name,surname)_R])
the book “1984” by
“George Orwell”
Q4 Given a book, deter- Q4(Book,
mine all information [Book(title,authors)_!],
of copies that con- [Copy_H])
tain it
Q5 Determine all the Q5(Copy,[Copy(format)_!,
copies with format rental(rentalDate)_Rt],
“hardcover”, rented [Copy(location)_!])
from a certain date
Q6 Determine the Q6(Copy,
copies with format [Copy(format)_!,
“hardcover” contain- Book(title,authors)_H],
ing the book “1984” [Copy(location)_!,
by “George Orwell”, Reader(readerId)_MRt])
together with the
readers that rented
them
(b) Workload and query representation. When a query
specifies “all the information of entity E,” it is
assumed to refer to all attributes of E. When a query
refers simply to entity E, it is assumed to refer to
the identifier of E.
lightweight educational notation that remains close to JSON syntax, reflecting the fact that aggregates
are often implemented as JSON documents in ANoSQL systems. At the same time, it extends JSON-based
representations with design-specific information not supported by JSON Schema, such as identifiers
and physical design elements (e.g., partition keys; see Section 4). This enables reasoning about both
logical and physical design within a single, system-independent representation.</p>
      <p>Heuristics. Depending on the cardinality constraints of the associations linking E to other entities,
the aggregate may include complex attributes. Semantically equivalent but syntactically diferent
representations can be generated at this stage. To ensure a general system-independent schema,
students are encouraged to limit the depth of property nesting (heuristic H2). An additional heuristic
concerns identifiers (H3). Although constraint checking in ANoSQL systems is typically delegated
to the application, document-based and column-family systems still support uniqueness constraints,
though with limitations. It is therefore important that students learn that one identifier for E should be
selected at this stage. To avoid introducing irrelevant attributes, the identifier that shares the greatest
number of attributes with the selection and projection attributes should be selected.
Example 3.1. Figure 1b shows a reference workload and the corresponding query representation for the
ER diagram in Figure 1a. Query Q6 includes selection conditions on attribute format of entity Copy and
on attributes title and authors of Book, projections over Copy and Reader entities. According to
H1, the aggregation entity should be Copy since it appears in LS and on the one side of the association
connecting the two selection entities. All entities except Reader are connected to the aggregation entity
Copy through a single association; Reader instead is linked through a path composed of two associations,
denoted as MRt in the query representation (Makes + RelatedTo). An alternative would be to select
Book as the aggregation entity; however, since one book may correspond to multiple copies, this would
introduce nested selection attributes (e.g., location and format) in the schema, in contrast with H1.
Book: { title, authors,
rated_by: [{name, surname}],
has_copies: [{location, format}] }
Reader:{ readerId, name, surname, birthdate,</p>
      <p>rates: [{title, authors}] }
Copy: { location, format,
rentals: [{rentalDate, readerId}],
title, authors }
(a) Annotated ER diagram.</p>
      <p>(b) JSON-SM aggregate schemas.</p>
      <p>Figure 2a shows the ER diagram annotated with the query workload, which serves as input to the final
step producing the aggregate schemas in Figure 2b.4 As an example, in the Book schema, Book attributes
annotated with queries Q3 and Q4 are included as simple properties in the JSON-SM representation. Since
the Book identifier is used in query selections, no further attributes should be added, according to H3.
The identifier attributes are underlined. To complete the schema, attributes of entities connected to Book
through annotated associations (Reader and Copy) are included as collection properties since they appear
on the many side of the relationships. ♢</p>
    </sec>
    <sec id="sec-4">
      <title>4. Block 3 - Physical Design</title>
      <p>To introduce physical design in a system-independent way, Block 3 focuses on (P3) and explains
how logical design choices afect query execution eficiency in aggregate-oriented ANoSQL systems.
Students learn to distinguish between batch queries, which scan the entire dataset, and random-access
queries, which target specific data subsets and for which partition-key selection is crucial. We further
show that query eficiency depends on (i) limiting execution to one or few nodes and (ii) avoiding full
partition scans, through data ordering or indexing. Based on these principles, Block 3 guides students
in understanding three complementary physical design aspects and one heuristic.
Partition key selection We show that a query is eficiently executed only when its selection
predicates include equality conditions on all partition-key attributes, ensuring query routing to a single node.
Additional selection conditions may restrict results to a subset of a partition but still require scanning
the entire partition, increasing disk accesses. When some equality conditions on the partition key are
missing, relevant aggregates may be distributed across multiple nodes, preventing eficient execution.
Heuristics. Based on these observations, we introduce a heuristic for partition-key selection (H4)
to support eficient execution for as many associated random-access queries as possible. Given an
aggregate schema  and the set of associated random-access queries , the partition key is derived
from the intersection int of simple selection attributes across the queries in , with diferent choices
depending on whether this intersection is empty or coincide/does not coincide with each set of query
selection attributes.</p>
      <p>Enforcing uniqueness constraints Unlike relational systems, enforcing a UNIQUE constraint in a
distributed setting is costly, since conflicting aggregates may reside on diferent nodes. To avoid global
4Notice that nested properties could also be represented with an additional nesting level (e.g.,
has_copies : [{copy : {location, format}}]); however, this increases schema depth and, according to H2,
should be avoided.
checks, ANoSQL guarantee uniqueness only locally within each node. As a consequence, identiefirs
must include the partition key, and minimality, typical of conceptual and relational identifiers, is no
longer preserved. To clarify this distinction and simplify the translation into system-related schemas,
we introduce the notion of super-identifier, defined as an identifier from the ER schema extended with
the partition-key attributes.</p>
      <p>Index creation We first note that in ANoSQL systems user-defined indexes are local, indexing the
aggregates stored on a single node, typically through ordered tree-based structures (e.g., B-trees). Local
indexes serve two main purposes: (i) an index on a super-identifier (with partition-key attributes first)
enables eficient uniqueness checking and may be created by the system (e.g., in Cassandra) or by the
user (e.g., in MongoDB); (ii) secondary local indexes can be defined on selection attributes of queries
that are not eficiently supported by design, improving aggregate retrieval on each node.</p>
      <p>To help student understand the strict relationship between logical and physical schemas, we extend
JSON-SM to represent partition keys and super-identifiers.</p>
      <p>Example 4.1. Consider the Book schema. Queries Q3 and Q4 share the same selection attributes (title
and authors); selecting them as the partition key ensures eficient execution, since both queries are routed
to a single node. Now consider the Copy schema. Queries Q5 and Q6 do not share the same selection attributes,
but their intersection is non-empty and corresponds to format. Selecting format as the partition key
routes both queries to a single node; however, additional selection attributes may require scanning the
partition. Secondary indexes on these attributes (e.g., (title, authors) for Q6) improve query eficiency.
For Book, the super-identifier coincides with the partition key, while for Copy it is (format, location).
In both cases, a unique local index can enforce the unique constraint. We use superscripts and overlined
text to represent partition keys and super-identifiers in JSON-SM (e.g., format, location for the Copy
entity). ♢</p>
    </sec>
    <sec id="sec-5">
      <title>5. Block 4 - System-oriented Design</title>
      <p>The fourth slide block shows how a JSON-SM schema representing both logical and physical information
can be systematically translated into schemas for key–value, document-based, and column-family
systems, exposing students to the power of abstraction across data models while clarifying the expressive
power and limitations of each ANoSQL system.</p>
      <p>To perform the translation, the block first guides students in understanding how each system
interprets a (, ) pair and the degree of system-level processing allowed on the value. For each
category of ANoSQL system, the block relies on general rules and worked examples to explain: (a)
whether the JSON-SM schema can be fully retained in the system (as in document-oriented systems
such as MongoDB) or needs to be adapted (as in column-family systems such as Cassandra), while
highlighting which queries remain eficiently supported at the system level and which do not; (b) how
super-identifiers map to system-specific mechanisms (e.g., unique indexes in MongoDB and primary
key declarations in Cassandra); (c) which system-specific commands implement the logical/physical
schema. This block is complemented by hands-on activities on MongoDB and Cassandra deployed on a
small departmental cluster.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Related Work</title>
      <p>
        The workload-driven nature of ANoSQL data modeling has been recognized since the early stages of
NoSQL systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and alternative schemas are known to afect performance [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Over the last decade,
numerous approaches have supported DBAs and data engineers in the complex task of workload-based
data design (see [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for a survey). Most of them separate conceptual and logical design, transforming
ERor UML-based models into ANoSQL schemas. Existing works difer in (i) focusing on a single system
(i.e., system dependent) [
        <xref ref-type="bibr" rid="ref12 ref7 ref9">7, 9, 12</xref>
        ] or on multiple systems (i.e., system independent) [
        <xref ref-type="bibr" rid="ref10 ref11 ref13 ref14 ref15 ref4">13, 10, 11, 4, 14, 15</xref>
        ],
(ii) considering only conceptual models [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] or also workload queries, with [
        <xref ref-type="bibr" rid="ref14 ref15 ref7 ref9">7, 14, 15, 9</xref>
        ] or without
query frequencies [
        <xref ref-type="bibr" rid="ref10 ref12 ref6">10, 6, 12</xref>
        ], and (iii) generating one [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref15">10, 11, 15, 12</xref>
        ] or multiple target schemas [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
or evaluating alternatives based on a single [
        <xref ref-type="bibr" rid="ref12 ref4 ref6">4, 6, 12</xref>
        ] or multiple criteria [
        <xref ref-type="bibr" rid="ref14 ref7 ref9">7, 14, 9</xref>
        ]. Other approaches
address schema design from diferent perspectives, such as object-NoSQL mappers [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and automated
sharding approaches [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] from application-level classes, focusing on complementary issues.
      </p>
      <p>
        Despite the wide range of complexity of the mentioned approaches, they are often system-specific
or highly detailed, rather than focusing on general principles, and primarily targeted at software
practitioners rather than students. An exception is [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which proposes a conceptually grounded
ERto-JSON mapping methodology, although limited to fully query-capable document-based systems.
In designing a more general methodology with educational goals, intermediate abstract models that
ensure technology independence become essential (see, e.g., [
        <xref ref-type="bibr" rid="ref13 ref4">4, 13</xref>
        ]), in line with the principles outlined
in Section 2. Our methodology can be viewed as a pedagogically oriented simplification of such
system-independent approaches (e.g., [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), integrating both logical and physical design.
      </p>
      <p>
        A unified methodology that captures common principles across diferent ANoSQL models is very
valuable for data system education. Overall, data systems education remains a challenging component
of many computer science curricula. Most research in this area focuses on query languages (primarily
relational languages and SQL [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ] and, to a lesser extent, NoSQL query languages [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]). For
what concerns modeling, some approaches target conceptual modeling [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and relational logical
modeling focusing on normalization and schema design [
        <xref ref-type="bibr" rid="ref22 ref23">22, 23</xref>
        ]. In contrast, NoSQL data modeling
is less explored [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Exceptions include: (i) reports on experiences of introducing NoSQL systems
in curricula [25, 26], emphasising query-driven modeling, data aggregation and denormalization; (ii)
proposals ofering insights into comparative modeling strategies and the flexibility of schema-less and
graph-based systems [27]; (iii) works collecting learning analytics on the use of SQL and NoSQL systems
through a React-based web application [28, 29].
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Concluding Remarks</title>
      <p>In this paper, we presented a set of educational resources to support the teaching of ANoSQL data
design in graduate-level Advanced Data Management courses. The approach grounds this design in
key properties of large-scale distributed environments, building on foundational data management
knowledge acquired in undergraduate curricula. Compared to existing approaches, the resources are
system-independent, emphasize abstraction, and highlight the interplay between logical and physical
design.</p>
      <p>The resources were deployed during the first term of the 2025/26 academic year and used in lectures,
exercise sessions, and a two-task autonomous assignment. Formative feedback focused on aggregate
schema correctness, partition-key selection, and consistency between logical and physical design choices.
The optional tool supported self-checking and exploration of alternative designs. While further activities
are needed to demonstrate measurable learning impact, preliminary observations are encouraging:
more than 50% of the submissions applied the methodology correctly and, in the first two exam sessions,
76% of students achieved a passing grade in the modeling exercise, improving on the 67% recorded in the
2024/25 academic year. This suggests that the resources support an efective and structured approach
to ANoSQL design learning.</p>
      <p>Future work includes a comparison with traditional domain-focused design and an experimental
exploration of the performance impact of alternative schemas, to be addressed through an additional
block and hands-on experiments on MongoDB and Cassandra.</p>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT-5.2 for: paraphrasing and rewording;
improving writing style; grammar and spell checking. After using this tool, the authors reviewed and
edited the content and take full responsibility for the publication’s content.
of Systems and Software 226 (2025) 112391.
[25] S. Mohan, Teaching NoSQL databases to undergraduate students: A novel approach, in: Proc.</p>
      <p>SIGCSE, 2018, pp. 314–319.
[26] S. Kim, Seamless integration of NoSQL class into the database curriculum, in: Proc. ITiCSE, 2020,
pp. 314–320.
[27] A. Alawini, P. Rao, L. Zhou, L. Kang, P.-C. Ho, Teaching data models with TRIQL, in: Proc. DataEd,
2022, pp. 16–21.
[28] V. Meyer, L. Wiese, A. I. A. Al-Ghezi, Analyzing student feedback to assess NoSQL education, in:</p>
      <p>Proc. IDEAS, Springer, 2025, pp. 211–223.
[29] V. Meyer, L. Wiese, A. I. A. Al-Ghezi, A unified teaching platform for (No)SQL databases, in: Proc.</p>
      <p>ICEIS, 2024, pp. 374–381.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Sadalage</surname>
          </string-name>
          , M. Fowler,
          <article-title>NoSQL distilled: A brief guide to the emerging world of polyglot persistence</article-title>
          ,
          <source>Pearson Education</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Roy-Hubara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sturm</surname>
          </string-name>
          ,
          <article-title>Design methods for the new database era: A systematic literature review</article-title>
          ,
          <source>Software and Systems Modeling</source>
          <volume>19</volume>
          (
          <year>2020</year>
          )
          <fpage>297</fpage>
          -
          <lpage>312</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bansal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. K.</given-names>
            <surname>Awasthi</surname>
          </string-name>
          ,
          <article-title>Are NoSQL databases afected by schema?</article-title>
          ,
          <source>IETE Journal of Research</source>
          <volume>70</volume>
          (
          <year>2024</year>
          )
          <fpage>4770</fpage>
          -
          <lpage>4791</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Davoudian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>A workload-driven method for designing aggregate-oriented NoSQL databases</article-title>
          ,
          <source>Data &amp; Knowledge Engineering</source>
          <volume>142</volume>
          (
          <year>2022</year>
          )
          <fpage>102089</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bansal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sachdeva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. K.</given-names>
            <surname>Awasthi</surname>
          </string-name>
          ,
          <article-title>Schema generation for document stores using a workloaddriven approach</article-title>
          ,
          <source>The Journal of Supercomputing</source>
          <volume>80</volume>
          (
          <year>2024</year>
          )
          <fpage>4000</fpage>
          -
          <lpage>4048</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W. Y.</given-names>
            <surname>Mok</surname>
          </string-name>
          ,
          <article-title>A conceptual-model-based design methodology for MongoDB databases</article-title>
          ,
          <source>in: Proc. ICICT</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>151</fpage>
          -
          <lpage>159</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Reniers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Landuyt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rafique</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Joosen</surname>
          </string-name>
          ,
          <article-title>A workload-driven document database schema recommender (DBSR)</article-title>
          ,
          <source>in: Proc. ER</source>
          <year>2020</year>
          , volume
          <volume>12400</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2020</year>
          , pp.
          <fpage>471</fpage>
          -
          <lpage>484</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Carey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. Y.</given-names>
            <surname>Alkowaileet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Digeronimo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Smotra</surname>
          </string-name>
          , T. Westmann,
          <article-title>Towards principled, practical document database design</article-title>
          ,
          <source>Proc. VLDB Endowment</source>
          <volume>18</volume>
          (
          <year>2025</year>
          )
          <fpage>4804</fpage>
          -
          <lpage>4816</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hewasinghage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nadal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abelló</surname>
          </string-name>
          , E. Zimányi,
          <article-title>Automated database design for document stores with multicriteria optimization</article-title>
          ,
          <source>Knowledge and Information Systems</source>
          <volume>65</volume>
          (
          <year>2023</year>
          )
          <fpage>3045</fpage>
          -
          <lpage>3078</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>A. de la Vega</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>García-Saiz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Blanco</surname>
            ,
            <given-names>M. E.</given-names>
          </string-name>
          <string-name>
            <surname>Zorrilla</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Sánchez</surname>
          </string-name>
          ,
          <article-title>Mortadelo: Automatic generation of NoSQL stores from platform-independent data models</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>105</volume>
          (
          <year>2020</year>
          )
          <fpage>455</fpage>
          -
          <lpage>474</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Atigui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Azough</surname>
          </string-name>
          , N. Travers,
          <article-title>ModelDrivenGuide: An approach for implementing NoSQL schemas</article-title>
          ,
          <source>in: Proc. DEXA</source>
          <year>2020</year>
          , volume
          <volume>12391</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2020</year>
          , pp.
          <fpage>141</fpage>
          -
          <lpage>151</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mozafari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Nazemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Eftekhari-Moghadam</surname>
          </string-name>
          ,
          <article-title>CONST: Continuous online NoSQL schema tuning</article-title>
          ,
          <source>Software: Practice and Experience</source>
          <volume>51</volume>
          (
          <year>2021</year>
          )
          <fpage>1147</fpage>
          -
          <lpage>1169</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Atzeni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bugiotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cabibbo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Torlone</surname>
          </string-name>
          ,
          <article-title>Data modeling in the NoSQL world</article-title>
          ,
          <source>Computer Standards &amp; Interfaces</source>
          <volume>67</volume>
          (
          <year>2020</year>
          )
          <fpage>103149</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Kuszera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Peres</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. D. D. Fabro</surname>
          </string-name>
          ,
          <article-title>Exploring data structure alternatives in the RDB-toNoSQL document store conversion process</article-title>
          ,
          <source>Information Systems</source>
          <volume>105</volume>
          (
          <year>2022</year>
          )
          <fpage>101941</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Roy-Hubara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sturm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shoval</surname>
          </string-name>
          ,
          <article-title>Designing NoSQL databases based on multiple requirement views</article-title>
          ,
          <source>Data &amp; Knowledge Engineering</source>
          <volume>145</volume>
          (
          <year>2023</year>
          )
          <fpage>102149</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>U.</given-names>
            <surname>Störl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hauf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klettke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Scherzinger</surname>
          </string-name>
          ,
          <article-title>Schemaless NoSQL data stores - object-NoSQL mappers to the rescue?</article-title>
          ,
          <source>in: Proc. BTW</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>579</fpage>
          -
          <lpage>599</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Scherzinger</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Thor, AutoShard
          <article-title>- declaratively managing hot spot data objects in NoSQL document stores</article-title>
          ,
          <source>CoRR abs/2111</source>
          .01086 (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Taipalus</surname>
          </string-name>
          ,
          <article-title>SQL: A trojan horse hiding a decathlon of complexities</article-title>
          ,
          <source>in: Proc. DataEd</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gilad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Miao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Stephens-Martinez</surname>
          </string-name>
          ,
          <article-title>What teaching databases taught us about researching databases: Extended talk abstract</article-title>
          ,
          <source>in: Proc. DataEd</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Alkhabaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alawini</surname>
          </string-name>
          ,
          <article-title>Student's learning challenges with relational, document, and graph query languages</article-title>
          ,
          <source>in: Proc. DataEd</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>30</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lopez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ferrada</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Hogan,
          <article-title>ERDoc: A web interface for entity-relation modelling</article-title>
          ,
          <source>in: Proc. DataEd</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>T. M. Connolly</surname>
            ,
            <given-names>C. E.</given-names>
          </string-name>
          <string-name>
            <surname>Begg</surname>
          </string-name>
          ,
          <article-title>A constructivist-based approach to teaching database analysis and design</article-title>
          ,
          <source>Journal of Information Systems Education</source>
          (
          <year>2005</year>
          )
          <fpage>43</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>C.</given-names>
            <surname>Köhnen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Heuer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zumbrägel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Scherzinger</surname>
          </string-name>
          ,
          <article-title>Making a case for visual feedback in teaching database schema normalization</article-title>
          ,
          <source>in: Proc. DataEd</source>
          ,
          <year>2025</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tripathi</surname>
          </string-name>
          ,
          <article-title>NoSQL database education: A review of models, tools, and teaching methods</article-title>
          ,
          <source>Journal</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>