<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparison of Graph- and Collection-Based Representations of Early Modern Biographical Archives</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Meadhbh Healy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas O'Connor John Keating</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Maynooth University Maynooth</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <fpage>60</fpage>
      <lpage>85</lpage>
      <abstract>
        <p>The ingestion and digital storage of historical records has had a profound impact on scholarly practices. Yet in order for digitized archives to be used to full advantage, it is imperative that they are searchable and organized in a coherent and consistent way. These requirements are particularly evident in the case of historical records pertaining to realworld individuals: personal data is likely to be intricate and may originate from disparate sources, whose rules of data collection and data storage vary greatly. While this makes graph-oriented databases a very attractive option for storing historical records due to their emphasis on attributes and relationships, document-oriented databases may oefr similar advantages in terms of flexibility and precision of record storage. In the present paper, both kinds of database are analyzed and compared in terms of the ease of ingestion and accuracy of record presentation within the database. The historical data used was gathered from a number of diverse collections of historical records referring to persons of Irish descent who served in European armies or studied at European universities between the sixteenth and the nineteenth century.</p>
      </abstract>
      <kwd-group>
        <kwd>In</kwd>
        <kwd>Tara Andrews</kwd>
        <kwd>Franziska Diehr</kwd>
        <kwd>Thomas Efer</kwd>
        <kwd>Andreas Kuczera and Joris van Zundert (eds</kwd>
        <kwd>)</kwd>
        <kwd>Graph Technologies in the Humanities - Proceedings 2020</kwd>
        <kwd>published at http</kwd>
        <kwd>//ceur-ws</kwd>
        <kwd>org</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <sec id="sec-1-1">
        <title>Overview</title>
        <p>The data hosted on the Virtual Research Environment (VRE) of the Irish in
Europe Project was collected between 2001 and 2008 by researchers at the
Universities of Leuven, Oxford, Toulouse, Dublin (Trinity College),
Madrid (Complutense), and Maynooth. Their work was part of a coordinated
eofrt to capture and host biographical material, held mainly in European
archival repositories, on Irish soldiers, merchants, and clerics in Europe and
the European empires in the early modern period (1550–1800). In total,
basic biographical records on about 15,000 individuals were identified and
harvested for hosting in a virtual research environment that would facilitate
advanced querying, graphical representation, and mapping. The assembled
material provides us with access to information about a specific migrant
population and its evolution over a significant time scale. More importantly,
within their digital environment, these sources help to deepen our
understanding of early modern migrant populations in general, while serving as a
template for the reconstitution of parallel migrant populations. They also
open up, for the first time, the possibility of quantitative, as well as
qualitative, comparative studies.</p>
        <p>
          One of the most crucial issues facing the digital humanities, and the Irish
in Europe Project VRE specifically, is the ingestion of ‘unstructured’ data
into a digital database and the manipulation thereof. By its very nature,
historical data can consist of uncertain and/or divergent primary source
material, which is why systems centered on modeling and representing
personal historical data are often extensive, complex, and heterogeneous
          <xref ref-type="bibr" rid="ref25 ref26 ref36 ref40 ref8">(Mosquera and Piedra, 2017)</xref>
          . We have found that while these complications can
be resolved to a certain extent by using NoSQL collection-based systems,
such as MongoDB, a great deal of redundancy remains. Similarly,
ingestion and search involve complex and time-consuming algorithms. We believe
that graph databases provide a more elegant solution to these problems, as
they enable eficient storage of data with intricate relationships and dynamic
schema. As graph databases are not rigid in their structure and organization,
it is important to consider integrity constraint (IC) support to ensure that
data insertion and other processes are performed in such a manner as to
ensure that data integrity is not aefcted. For the purposes of this study, we
will focus on providing a preliminary overview of the graph-oriented
database Neo4j and a comparison with the most prolific NoSQL database,
MongoDB, using the prosopographical data described above. We will compare
ingestion and search performances for the collection and graph systems
together with an evaluation of schema and IC creation approaches for both
systems.
1.2
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Topics to Be Addressed</title>
        <p>
          NoSQL – commonly referred to as ‘Not Only SQL’ (with SQL standing
for Structured Query Language) – databases have been developed as a
response to the limitations of existing relational database management
systems (RDBMS). While traditional RDBMS are capable of large-scale data
management for structured data, NoSQL databases adeptly manage large
amounts of structured, unstructured, semi-structured, and hybrid data at
reduced complexity and cost
          <xref ref-type="bibr" rid="ref25 ref26 ref36 ref40 ref8">(Mohmmed and Osman, 2017)</xref>
          . Yet graph
databases, a particular form of NoSQL databases, are becoming increasingly
significant in real-world applications: they provide an adequate framework for
representing complex relationships in diverse datasets, which can lead to the
discovery of causal relationships by combining disparate sources of
information
          <xref ref-type="bibr" rid="ref20">(Le May et al., 2020)</xref>
          . With the development of graph databases it has
become possible to combine the performance of NoSQL databases and the
representativity of graphs
          <xref ref-type="bibr" rid="ref26 ref35 ref40 ref8">(Castelltort and Martin, 2018)</xref>
          .
        </p>
        <p>
          NoSQL databases vary greatly in how they store and manage data. One of
the primary distinctions is that in relational models the schema is extremely
rigid and possible relationships are fixed in advance. In graph databases,
the information is stored in a schemaless format, (key-value pairs), which
allows several related values to be stored at the same node
          <xref ref-type="bibr" rid="ref37">(Čerešňák and
Kvet, 2019)</xref>
          . When it comes to archiving historical data in particular, many
NoSQL databases provide a distinct advantage over SQL databases. This
is again due to the rigidity of the SQL schema, which is inferior where the
modeling and absorbing of unstructured historical data is concerned. Both
types of database compared in this paper are NoSQL; MongoDB is the most
prominent collection-based database available at present, and Neo4j is a top
performing graph database.
1.3
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>Motivation</title>
        <p>
          One of the immediate incentives for historians and digital archivists who
wish to digitize historical data is that the archives become more accessible
instead of remaining in storage with no public access. The immense changes
taking place in the past number of years have begun to reshape digital
curation and digital historiography, as detailed extensively by Sabharwal
          <xref ref-type="bibr" rid="ref31">(Sabharwal, 2015)</xref>
          . It is imperative for history scholars and academics, however, that
certain precautions be taken when reproducing a historical collection of
miscellaneous nature in a digitized format, in order to precisely replicate the
original
          <xref ref-type="bibr" rid="ref7">(Borissova, 2018)</xref>
          .
        </p>
        <p>In particular, this means detailing relationships between the information
chunks of a record, as well as delineating common relationships between
records in a precise manner. This is a significant challenge when compiling a
database of historical data, as it is easy to convolute multiple records from
diefring datasets if they are not correctly represented. What results is an
interpolation problem, as it must be determined how best to reproduce
intricate, heterogeneous data without loss of complexity, making an analysis
of integrity constraint support vital for digital archivists.</p>
        <p>
          The data being examined in this study originally formed part of a now
obsolete and no longer publicly accessible VRE website where each dataset was
modeled, and each record stored, as an intricate XML file. XML databases
could not be considered due to the volume of data, as the syntactical
redundancy of XML would have resulted in a prohibitively large transport and
storage eofrt. Moreover, the records collected have been archived in multiple
diefrent languages, which means that it is entirely possible for a military
record in the French collection, for example, to have details listed in English,
Irish (Gaelic), Latin, and French. Given that any or all of the names Sean De
Paor, John Power, or Jean LePoer could be cataloged in a record for the same
individual, querying a database for precisely one of the above would not
return results for the others. This is why a database that permits regex (short
for regular expression) queries, which allow the fields being searched to be
matched to a query pattern, is an essential requirement. XML databases are
limited in this regard, as nested tags are very dificult to manage with
regular expressions
          <xref ref-type="bibr" rid="ref17 ref34">(Taktek and Thakker, 2020)</xref>
          . Due to the way information is
stored in MongoDB and Neo4j, the two platforms are much better suited to
the use of regular expressions.
1.4
        </p>
      </sec>
      <sec id="sec-1-4">
        <title>Problem Statement</title>
        <p>
          Integrity constraints are rules that restrict the information that may be
present in the database. They play a major role in maintaining the precise
structure of a record or set of records. The constraints in SQL databases can
be divided into two main components: entity integrity, which tests the
validity of the data by providing primary keys; and referential integrity, which
adds meaningful structure to the data by using foreign keys to tie relations
together
          <xref ref-type="bibr" rid="ref6">(Bono, 2007)</xref>
          . As referenced previously, many NoSQL databases
suffer from a lack of strictly designed schema structure, which can impede any
prospect of securely defining and confirming the rules of referential
integrity, and can result in a lower degree of control over data values
          <xref ref-type="bibr" rid="ref5">(Bjeladinovic
et al., 2020)</xref>
          . However, this is not the case in graph-oriented databases such
as Neo4j, where referential integrity is maintained by establishing a
relationship between two nodes. Much of the support for data integrity is solely
available in SQL databases as it is implemented in the SQL language, and this
can leave non-relational databases at a severe disadvantage when attempting
to maintain data integrity. The justification for this is easy to fathom: by
deliberately maintaining low-level systems in a NoSQL environment, it is
possible to create, store, and analyze vast quantities of data at high speeds. In
order to obtain higher data consistency and an increase in reliability of NoSQL
applications, transactional services, which vary according to the NoSQL
system, have been developed
          <xref ref-type="bibr" rid="ref13">(González-Aparicio et al., 2018)</xref>
          . Various
measures can be implemented on the user application side, including designing
test cases which check various possibilities during the execution of a
transaction in order to detect potential faults or inconsistencies in the data
          <xref ref-type="bibr" rid="ref1">(Agnelo
et al., 2020)</xref>
          .
1.5
        </p>
      </sec>
      <sec id="sec-1-5">
        <title>Approach</title>
        <p>
          When processing large amounts of real-world data, entities may be
represented in a variety of formats, such as JSON (JavaScript Object Notation) or
relational records, and representations may contain redundant or
inconsistent information
          <xref ref-type="bibr" rid="ref33">(Simonini et al., 2019)</xref>
          . Data entity types are another
important issue, as they prevent certain schematic anomalies – for example, it
should not be possible to enter a string into an integer-specific column and
vice versa (Šestak et al., 2016). Several distinct factors can be identified when
discussing data integrity in graph database models
          <xref ref-type="bibr" rid="ref2">(Angles, 2012)</xref>
          :
• schema-instance consistency, which prevents incomplete or
inconsistent data from being inserted into the database and ensures that each
entity can only have the attributes and relations previously established
in the desired schema
• node or edge identity, which demands that each value in the database
can be identified by a value, such as name or id, or the values of its
attributes
• cardinality integrity, which stipulates that each node in the database has
a unique identifier which is the equivalent of a primary key constraint
in relational databases
• referential integrity, which requires that only existing nodes in the
database can be referenced, similar to foreign key constraint
• functional dependencies, which tests if an entity determines the value
of another database entity
• graph pattern constraints, which identify structural restrictions such as
path constraints in the data
The elimination or control of data redundancy, which would decrease the
volume of superfluous information stored in the database, is another
important consideration. For the purpose of archiving historical data,
functional dependencies are largely not a priority – as the data is static, the
entities do not depend on each other and the values do not change. However,
as will become clear over the course of this paper, the other factors listed
above are all crucial when it comes to ensuring adequate referential integrity
constraint support. In order to determine the optimal method of
archiving historical data-sets, a comparison test was conducted between two types
of NoSQL databases, Neo4j and MongoDB. In MongoDB, a
documentoriented database which processes semi-structured data, each record and its
associated attributes are considered a document. MongoDB stores data as
a hierarchy of key-value pairs and provides a rich query language allowing
for easier transition from relational databases. Neo4j is a native graph
database which is geared towards the storage and processing of graphs, and
allows the management of interconnected data. Graph databases help find
relationships between data, and have index-free adjacency. This means that
while relational or other non-native graph databases have central indexes and
processing overheads with every index lookup, graph databases load every
relationship associated with an entity when a node is accessed
          <xref ref-type="bibr" rid="ref14">(Henderson,
2020)</xref>
          .
1.6
        </p>
      </sec>
      <sec id="sec-1-6">
        <title>Metrics</title>
        <p>In order to accurately compare the performance of MongoDB against
Neo4J, the research has been divided into three distinct categories:
• to observe and determine the dificulty, or lack thereof, of ingesting a
record into both types of database, for an individual record and a set of
records
• to examine the structure of the data in both sets of software, placing
particular emphasis on the precision of conversion from the XML
record to its reproduction within the database and the lack of dificulty
of merging records between datasets
• to analyze the complexity at which an individual record can be
extracted from both types of database
The databases were evaluated according to both shared and unique features.
1.7</p>
      </sec>
      <sec id="sec-1-7">
        <title>Data</title>
        <p>The data being examined and archived comes from a variety of sources. It is
diverse, real world data that is both complex and heterogeneous. The library
containing the data consists of five primary components:
• Brockliss &amp; Ferté; this data was gathered by historians Laurence
Brockliss and Patrick Ferté in collaboration with institutions in Oxford and
Toulouse, and details Irish students that studied in Irish colleges in
Europe in the 1700s and 1800s
• Spain; records of Irish soldiers serving in Spain during the seventeenth
and eighteenth century, collected by Óscar Recio Morales of the
Complutense Institute for International Studies.
• France; records of Irish regiments in France in the 1700s and 1800s,
collected by Colm Ó Conaill of Trinity College Dublin
• King’s Inns; records of Irish students who attended King’s Inns in
the pre-Cromwellian era, archived by Brid McGrath of Trinity College
Dublin
Each dataset contains intricate and varied records, which can make it
dificult to establish a comprehensive pattern. On the other hand, as mentioned
above, the records are currently stored in XML files, and the fact that each
class of data has already been resolved into elements and attributes is of great
help when it comes to examining the structure of an individual record.
2
2.1</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Technical Background</title>
      <sec id="sec-2-1">
        <title>Topic Material</title>
        <p>Neo4j and MongoDB are compared in terms of both speed of processing and
the formation of the syntax necessary to create a query. As stated above, the
records that are being processed are complex and involve nested data, which
makes it advisable to create an encoding of each data library before they are
absorbed by the software. The model is based on the structure of the XML
ifle of each collection, and therefore no assumptions were made in
transforming the records to the model produced. A parent-child model of the Brockliss
&amp; Fertè dataset is shown in Figure 1.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Technical Material</title>
        <p>
          MongoDB (as well as Neo4j) does not use Structured Query Language
(SQL) to interact with the database, but it is compatible with a number of
languages, including Go, C++, and Python. MongoDB’s primary querying
language is JavaScript.1 The documents are hierarchical tree data structures
which can consist of maps, collections, and primitive values. MongoDB uses
the following hierarchy: database, collection, and document (Mahajan et al.,
1When commands are composed in this language, it creates a JSON document, BSON
(Binary JSON) object, or sub-documents, which are the primary components of collections
in the database
          <xref ref-type="bibr" rid="ref17 ref34">(Jose and Abraham, 2020)</xref>
          .
2019). There are a number of bespoke querying languages that can be used
in conjunction with Neo4j, the most popular of these being Cypher. The
syntax of Cypher is similar to SQL and uses an ASCII-Art syntax to
denominate patterns. Ideally, Cypher queries are constant strings, so they can be
cached by the database as compiled queries. The returned entities can be
nodes with all attributes, selected attributes, or aggregated data, depending
on the query
          <xref ref-type="bibr" rid="ref15">(Holzschuher and Peinl, 2016)</xref>
          . Both types of database are fully
ACID compliant.2 ACID properties guarantee that data integrity is
maintained for every transaction in the database. This ensures strict consistency
in the sense that all read operations must return the precisely same entities
as the latest completed write operation
          <xref ref-type="bibr" rid="ref21">(Lotfy et al., 2016)</xref>
          .
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Software and Syntax Analysis</title>
      <p>
        In MongoDB, records are stored in collections. These are analogous to
tables in relational databases, although the implementation and operational
concept of MongoDB and RDBMS is diefrent. Unlike in SQL databases,
the schema of a table does not have to be determined before inserting data.
For MongoDB, it is not imperative for all records within a collection to
have the same schema, which makes it possible to change the framework
2ACID stands for atomicity, which means that a database transaction must be entirely
ifnished, or it will not commit; consistency, which means that a database must remain
consistent before, during, and after the transaction occurs; isolation, which means that when
multiple transactions are executed simultaneously, transactions are processed exclusively
and consecutively – data from one transaction cannot be transferred to another when the
transaction has not been completed; and durability, which indicates that once database
transactions are registered, events are recorded to a permanent medium which will not be
modified outside a transaction.
of a schema within a collection by adding, removing, or updating fields
        <xref ref-type="bibr" rid="ref37">(Čerešňák and Kvet, 2019)</xref>
        . Creating an entity-relationship model is
vital in order to display the record or object nesting correctly
        <xref ref-type="bibr" rid="ref10">(Edward and
Sabharwal, 2014)</xref>
        . It can assist in illustrating the embedding and
denormalization necessary for scaling the data in MongoDB. Denormalization may be
thought of as tables being refined and transformed into secondary simplified
data structures, where redundancy is regulated in order to optimize
performance. Denormalization allows the data to be wholly retrieved without using
a join
        <xref ref-type="bibr" rid="ref19">(Kingdon et al., 2016)</xref>
        . To insert a record, the following syntax is used:
db.Brocklissferte.insert({ id : "1",
personalDetails : {"name" :
{"forename" : "Dionysius", "surname" : "O'Beirne"}
gender : "male",
origin : {"residence" : {"address" :
"diocese" : "Ardagh",
"province" : "Armagh",
"country" : "Ireland"}}}},
education : {"qualification" : {"course" :
"received the four minor orders",
"achieved" : "yes",
"date" : {"end" : "1771-05"},
"institution" : {"name" : "University of Paris",
"address" :
{"town" : "Paris",
"country" : "France"}}}},
notes : {"note" : "Boyle, 'St. Nic.', p. 490;
      </p>
      <p>Boyle, I.C.P., p. 200."}})
The fields of the schema can be then adjusted for diefrent records within a
collection (Figure 2).</p>
      <p>
        Graph databases, such as Neo4j, provide the most sophisticated and
evolved method of data modeling, making it easy to update the schema
according to the user’s needs (
        <xref ref-type="bibr" rid="ref27">Perçuku et al., 2017</xref>
        ). Normalization and
denormalization are largely redundant here, as graph databases provide as much
or as little structure as the data requires. Traditional SQL databases have
rigid schema and a convoluted schema migration process, making the creation
and continuous ingestion of records a challenging exercise, particularly for
highly relational data
        <xref ref-type="bibr" rid="ref32">(Schulz et al., 2016)</xref>
        . Alternatively, graph databases
can be employed as an ideal method of managing highly connected data, as
they prioritize the modeling and retrieval of relationship-rich data (
        <xref ref-type="bibr" rid="ref28">Pokorný
et al., 2017</xref>
        ). In fact, once a comprehensive graph data model which
incorporates the complex nested fields present in each particular record has been
composed, it can be absorbed directly into the database given that the
structure of the graph model corresponds exactly to the structure of the schema
within Neo4j
        <xref ref-type="bibr" rid="ref7">(Vágner, 2018)</xref>
        . This then automatically generates a Cypher
command (the domain language of Neo4j), which can be run in the database
to create the schema. The following Cypher command was used to create the
schema in the Brocklissferte database:
CREATE('0':Person),
('1':personalDetails),
('2':Name {forename:'$forename,',
surname: '$surname'}) ,
('3' :gender {gender:'$male'}) ,
('4' :origin),
('5':residence {diocese:'$county,',province:
'$province,',country:'$country'}) ,
('6':Education),
('7':Qualification
{course:'$course,',achieved:'$achieved,',
date: '$date'}),
('8':Institution
{institutionName:'$institutionName,',
address:'$address'}),
('9':Notes {notes:'$notes'}) ,
('1')-[:'INDIVIDUAL' ]-&gt;('2'),
('0')-[:'DESCRIPTION' ]-&gt;('1'),
('1')-[:'RELATED\_TO' ]-&gt;('3'),
('1')-[:'HOMEPLACE' ]-&gt;('4'),
('4')-[:'ADDRESS' ]&gt;('5'),
('0')-[:'STUDIES']-&gt;('6'),
      </p>
      <p>
        Although graph databases are constantly under development in order to
increase their stability and expand their range of features, they have not yet
reached the maturity level of other data management solutions, such as
relational databases. Nonetheless, a number of mechanisms exist which can
help to increase the applicability of graph database technology in a real-world
context. In some instances, traditional solutions developed for relational
databases can be reworked and adapted to the context of graphs, integrity
constraints being a case in point: (Šestak et al., 2021) for example, much of
the integrity constraint support available for Neo4j has been written into the
Cypher query language in a similar style to SQL (e.g. UNIQUE or NOT NULL)
        <xref ref-type="bibr" rid="ref22">(Ma
et al., 2020)</xref>
        .
      </p>
      <p>In order to directly ingest documents into the MongoDB database, the
mongoimport command can be used. This utility allows data to be imported
from JSON, CSV, or TSV files. It is not necessary to specify a collection
when importing data into the database; however, a collection will be created
upon the addition of the records. Like other NoSQL databases, MongoDB
has a dynamic schema design, allowing the documents in a collection to have
varied fields and structures.</p>
      <p>
        While Neo4j oefrs an option to eefctively import large datasets in
several diefrent formats, it recommends the LOADCSV command as the optimal
method. This facilitates the conversion from relational or other type
databases to a graph database format
        <xref ref-type="bibr" rid="ref18">(Karan, 2016)</xref>
        . This command will
directly map input data into a complex graph/domain structure. When
performing the operation on a significant amount of data, the command can be
appended with the fragment USING PERIODIC COMMIT, which reduces memory
overhead when the transaction is being conducted.3
      </p>
      <p>
        Within the MongoDB database, each document is stored in the JSON
format. As JSON documents support embedded fields and nested data,
related data can be included within the document instead of having to be
stored externally within the collection. Embedded fields act as placeholders
that can be added to text fields to dynamically display entity-specific content.
Each JSON field consists of unordered key-value pairs, a form of NoSQL
database that has become increasingly prevalent in recent years, with each
pair stored in a key-based lookup structure
        <xref ref-type="bibr" rid="ref1">(Agnelo et al., 2020)</xref>
        . The value
3However, other file types such as JSON and XML are also supported and can be
processed using APOC, an add-on library in Neo4j that is accessible with a very simple
command.
is represented as a document encoded in standard semi-structured format.
      </p>
      <p>
        As mentioned previously, Neo4j has native graph storage, which means
that each part of the graph data model is stored separately. There are diefrent
notions as to what makes up the key components of a graph database, one
of them being the property graph model, which is schemaless and allows
the user to represent the data close to a real-world conception. The records
in a Neo4j graph database are structured with each entity as a node, which
are linked through directed connections named relationships. Nodes and
relationships refer to their attributes as properties
        <xref ref-type="bibr" rid="ref12">(Giabelli et al., 2021)</xref>
        .
      </p>
      <p>
        A fundamental challenge in Neo4j data modeling is classifying a
categorical variable as a property, label, or node. A categorical variable may be
described as having two or more distinct categories with no intrinsic
ordering
        <xref ref-type="bibr" rid="ref3">(Baak et al., 2020)</xref>
        . The records being ingested here are rich in
categorical variables; for example, in the Brocklissferte architecture, residence,
diocese, province, and country all have a finite, discrete set of values. In graph
data modeling, categorical variables can often result in more irregularity,
due to the options there are for representing them. By way of example, let
us consider the category of gender, which is characterized as a label as
follows: (:Person:male). As a property value, it can be represented as (:Person
{gender:'male'}), and as a distinct node as (:Person)-[:GENDER]-&gt;(:Gender {name:
'male'}). These selections become even more complex when considering the
cardinality for each categorical variable – for example, almost fifty dioceses
are represented in the Brocklissferte dataset, along with five provinces. While
by the standards of Big Data neither of these attributes can be described as
having high cardinality, consistency must nonetheless be maintained in the
approach to modeling them
        <xref ref-type="bibr" rid="ref15 ref24">(Moeyersoms and Martens, 2015)</xref>
        . The
merits of each property type become apparent when endeavoring to retrieve
variables using regular queries. A regular query may be described as a
nonrecursive query that traverses the graph and returns pairs of nodes connected
by a common expression
        <xref ref-type="bibr" rid="ref4">(Bagan et al., 2020)</xref>
        . Labels attach simple types to
nodes and relationships. They provide fast look-ups for Neo4j and are used
to describe the nodes’ role in a graph as well as for grouping nodes in fixed
sets. However, they are a poor quality option for medium or high
cardinality values, as a variable with a large amount of categories can make a data
model extremely unwieldy. In addition, while relationships between
connected data can be referenced directly by Neo4j, relationships between
corresponding labels are hidden inside nodes and are not expressed explicitly
        <xref ref-type="bibr" rid="ref42">(Zhu
et al., 2019)</xref>
        . Properties are expressed as name-value pairs and can store any
data type. Thus, they can accommodate high cardinality data with ease, and
can be subjected to database constraints which preserve integrity (
        <xref ref-type="bibr" rid="ref16">Jiménez
et al., 2016</xref>
        ). (They cannot be set to NULL, as this is equivalent to
deleting the property). A small-scale property lookup is quicker that traversing a
node. The disadvantage of properties is that when multiple small categoric
variables are applied to a node, performance can be adversely aefcted. The
same can be said for parsing multiple properties to a large string or large array.
When searching for all nodes that share a specific property as part of a
regular query pattern, properties are not the optimal choice
        <xref ref-type="bibr" rid="ref30">(Ruetter et al., 2015)</xref>
        .
The command below illustrates the syntax that was used in an attempt to
return all common dioceses of a particular record (under the assumption that
all fields had been modeled as properties of the ‘person’ node):
      </p>
      <sec id="sec-3-1">
        <title>MATCH (b:Person \{ surname: 'Moore' \})</title>
      </sec>
      <sec id="sec-3-2">
        <title>WITH b</title>
      </sec>
      <sec id="sec-3-3">
        <title>MATCH (allCommonDiocese:Person \{ diocese: b.diocese \})</title>
      </sec>
      <sec id="sec-3-4">
        <title>RETURN allCommonDiocese;</title>
        <p>
          When searching for nodes that share a common property, or if the
cardinality of the categorical variables is inordinately high, modeling each variable
as a separate node can be the most eficient choice. Retrievability of nodes
has been the subject of extensive research for a number of decades
          <xref ref-type="bibr" rid="ref11">(Gacem
et al., 2020)</xref>
          . In this instance, if each record were to be filtered by shared
diocese, the following query might be used:
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>MATCH (b:Person \{ surname: 'Moore' \})</title>
      </sec>
      <sec id="sec-3-6">
        <title>WITH b</title>
      </sec>
      <sec id="sec-3-7">
        <title>MATCH (b)-[:HAS]-&gt;(j:PersonalDetails)</title>
        <p>[:HAS]-&gt;(o:Origin)[:HAS]-&gt;(r:Residence)[:HAS]-&gt;(d:diocese)&lt;-[:HAS]-(r:Residence)
&lt;-[:HAS]-(o:Origin)&lt;-[:HAS]&lt;-(j:PersonalDetails)
&lt;-[:HAS]-(other:Person)</p>
      </sec>
      <sec id="sec-3-8">
        <title>RETURN count(other);</title>
        <p>We can see how cumbersome this syntax can be, as the levels of the graph
descend from the root. Moreover, if data is too densely connected, it can
result in the creation of supernodes, i.e. vertices with a disproportionately high
number of relationships, which will have a negative impact on any queries
that attempt to access them.</p>
        <p>
          Therefore, our approach to modeling the data in Neo4j had to be tailored
to the questions the user will need to ask of the database, and an adaptable
attitude is required. Unlike traditional RDBMS, there is no significant
decrease in performance for highly connected data ingested, stored, and
navigated in Neo4j – as can be noted from Figure 3, it has no dificulty in
representing structured, semi-structured, or unstructured data (
          <xref ref-type="bibr" rid="ref27">Perçuku et al.,
2017</xref>
          ).
4
4.1
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>The Solution</title>
      <sec id="sec-4-1">
        <title>Analytical Work</title>
        <p>Both types of database support a Python driver.4 Although both types of
database support several diefrent methods of bulk import, some of which
have been outlined in the previous chapter, this approach was deemed the
fastest and most objective way of conducting a fair comparison test for
multiple data ingestion and retrieval.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>High Level</title>
        <p>Retrieving data from the database in MongoDB is also done via JavaScript
queries, using a simple command.</p>
        <p>This can then be appended depending on the specificity of the results that
need to be recovered. Each record in a collection will have a unique ID, which
can also be accessed when the ‘find’ command is used with a single criterion.</p>
        <p>The node always has its own variable name, an alias that is assigned by the
user the first time the node is referenced within a query. This makes the
retrieval of individual nodes a less complex process. However, returning an
entire record is a more formidable task in Neo4j. Graph databases, on the other
4PyMongo is the recommended method of interacting with MongoDB from Python,
while the oficial Python driver for Neo4j is ‘neo4j-python-driver.’
hand, are well suited for ontology-oriented data, representing a record as a
set of concepts and relationships. These entities are all dispersed and stored
separately, meaning that their retrieval requires a more explicit approach.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Schema and Data Migration</title>
        <p>Due to the static nature of the data, updating the records would be an
infrequent operation, but it is still prudent to examine the intricacy of such an
undertaking. In MongoDB, the command is straightforward:
db.brocklissferte.update_one(
{'id' : origin.get('id') },
{'$set': {"occupationalHistory":</p>
        <p>({"occupation": "Cleric", "role": "Priest"})}})</p>
        <p>Indexing is paramount in MongoDB, and the ID is required in order to
retrieve and update the record. Updating an individual record in Neo4j is a
similarly easy task:</p>
        <sec id="sec-4-3-1">
          <title>MATCH (n) WHERE n.surname = "Clarke" SET n.address = "Meath"</title>
          <p>The situation becomes more complicated when making changes to the
structure of an individual record. Although MongoDB is commonly
referred to as ‘schemaless’ data, it is important to contemplate how data is
stored in order to optimize database performance. If we consider the
Spanish Military parent-child model, it can be observed that the framework of the
record is as follows:
&lt;employment&gt;
&lt;employer&gt;
&lt;address&gt;</p>
          <p>&lt;country&gt;"Spain"&lt;/country&gt;
&lt;/address&gt;
&lt;/employer&gt;
&lt;/employment&gt;
&lt;employment&gt;
&lt;employer&gt;</p>
          <p>&lt;address&gt;
In some of the records, an extra set of parameters is present, recording the
details of military inspection:
&lt;country&gt;"Spain"&lt;/country&gt;
&lt;/address&gt;
&lt;date&gt;1748&lt;/date&gt;
&lt;record field="age"&gt;21&lt;/record&gt;
&lt;record field="height"&gt;5.2&lt;/record&gt;
&lt;record field="eyes"&gt;Black&lt;/record&gt;
&lt;record field="hair"&gt;Black&lt;/record&gt;
&lt;/employer&gt;
&lt;/employment&gt;
In order to accurately represent this structure across multiple records in
MongoDB, it is necessary to invoke the $addFields(aggregation) function,
which appends extra fields into each document.</p>
          <p>The format of a Neo4j update query is quite diefrent. A node can be
created and then accessed and added to directly using the SET clause.</p>
          <p>This is an important issue when attempting to accurately represent this
data in the database: the complexity of the data is such that it is necessary to
append and update the data encoding habitually in order to ensure accurate
reproduction. All told, both of the command syntaxes in question provide
a distinct advantage over schema migration in SQL databases, as they are less
complex and more intuitive.</p>
          <p>
            When considering multiple collections, as is being done in this instance,
one feature that needs to be discussed is merging multiple records on a
particular field or attribute. In database technology, these are known as
aggregate functions. The dominant aggregation framework of MongoDB uses a
pipeline concept, as it provides eficient data aggregation using native
operations. A pipeline is an array consisting of distinct operators which modifies
a collection (or sub-collection) in stages. As the collection passes through
each stage, certain operators modify the collection documents according to
various techniques
            <xref ref-type="bibr" rid="ref23">(Mahajan et al., 2019)</xref>
            . A less eficient alternative is the
MapReduce framework, a data processing technique that uses two stages or
tasks, namely Map and Reduce. The map function takes a MongoDB
document and maps each individual element to a key-value tuple. During the
subsequent reduce stage, the elements are condensed and aggregated data is
collected. An aggregation framework is generally faster than MapReduce,
but MapReduce is useful for aggregating extremely large collections. If an
archivist wishes to merge records from two of the previously described
collections on a particular value – for example, the Brockliss &amp; Fertè and King’s
Inns collections, both of which include a sponsor or employer address field
and contain records from a similar time period – the following syntax could
be used:
db.collection.aggregate([
{ "$lookup": {
"from": Brocklissferte,
"let": { "address": "$address" },
"pipeline": [
{ "$match": { "$expr": { "$in":
          </p>
          <p>[ "$employer", "$$address" ] } } },
{ "$lookup": {
"from": Innstudents,
"let": { "address": "$address" },
"pipeline": [
{ "$match": { "$expr": { "$in":
["$place", "$$address" ] } } } ],
"as": "address" }} ],
"as": "address"}}])</p>
          <p>An updated parent-child model of the Brockliss &amp; Fertè records may be
observed in Figure 4. This syntax, while looking quite intricate upon initial
observation, is of great help in the consolidation of nested data, as it covers
the range of each record to find and merge the relevant field. In this instance,
it provides an assortment of records to the user, giving access to
information regarding relevant benefactors and their level of influence in a
particular area. Meanwhile, $lookup is a pipeline operator that allows the user to
perform a left join to combine two collections in the same database. A left
join will contain all documents from the collection referenced first, as well
as the matching documents from the second collection. All unmatched
documents from the second collection will be omitted. The aggregation
functions in Neo4j are comparable to those in relational databases with a
syntax similar to SQL. The amount of matching rows can be queried using the
count(*). The collect() command returns a list of heterogeneous elements in
a single list.</p>
          <p>
            One major advantage NoSQL databases have over relational databases is
the increase in scalability. A highly scalable database is one which can
increase its workload and throughput when additional resources are added to
it. Scalability may be supported by the following approaches: horizontal
scaling, the process of adding more hardware to the system; and vertical
scaling, which increases the memory of the existing server. Although our data
is static, and is infrequently increased or altered, it is worth considering this
aspect to fully understand the capabilities of our database. MongoDB
supports horizontal scaling through sharding; a method for distributing data
across several servers
            <xref ref-type="bibr" rid="ref29">(Ravat et al., 2020)</xref>
            . Neo4j is both vertically and
horizontally scalable. The graph database platform provides high quality scaling,
and Neo4j uses cypher query language, an external Domain Specific
Language (DSL), which is tailored to a specific application domain. DSLs are
expressive and concise and therefore aid scalability by reducing complexity
            <xref ref-type="bibr" rid="ref26 ref35 ref40 ref8">(Yoon and Lee, 2018)</xref>
            .
5
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>The comparison between both types of non-relational database yielded a
number of illuminating observations and results. The primary objective was
to ingest the data with as little loss of data integrity as possible. As the
complexity of the data was illustrated using a parent-child relationship model, it
made the composing of queries and interactions with both types of database
a less arduous task. The syntax of creating this structure with both
MongoDB and Neo4j was exacting. However, the fact that a Cypher command
could be automatically generated in Neo4j upon the creation of a graph data
model, as opposed to being painstakingly produced for each individual
dataset by the user, gives graph databases a distinct advantage.
5.2</p>
      <sec id="sec-5-1">
        <title>Examination of Schema Traversal</title>
        <p>It is worth noting the contrast between the syntax of retrieving a nested data
object from both types of database, especially with regards to the complexity
of each query. From traversing the parent-child model in Figure 1, it can be
observed that the entity ‘country’ is a child of the entity ‘residence,’ which is
a child of the entity ‘origin,’ which is a child of the entity ‘personal details.’
The parent entity is described as ‘person.’ To correctly traverse the data in
order to retrieve this entity in MongoDB, each property has to be labeled in
the command.</p>
        <p>This results in the following record being retrieved:
{'name': {'forename': 'Thomas', 'surname': 'Williams'},
'_id': ObjectId('5e6396629578d4b117ea55fa'),
'origin': {'residence': 'province': 'Connaught',
'country': 'Ireland',
'diocese': 'Wexford'}},
'id': 128016,
'gender': 'male',
'education': {'qualification': {'achieved': 'yes',
'date': 1747-10-22 08:47:32',
'course': 'Law'},
'address': {'town': 'Brussels', 'country': 'Belgium'},
'institution': 'University Of Brussels'}}</p>
        <p>This is a more complex syntax than the Neo4j variant. As can be observed
in Figure 3, a graph data model which is representative of the structure of
the record within the database, the entity ‘country’ is stored as a property
of the node address. This node can be directly queried by the user with the
following cipher query:</p>
        <sec id="sec-5-1-1">
          <title>MATCH (n:address) RETURN address.country</title>
          <p>The contrast in complexity between both commands emphasizes that
graph-based data models provide a much more eefctive and adept method
of data traversal when interacting with the database.</p>
          <p>As stated earlier, Neo4j provides a graphical environment which allows
the user to observe and examine patterns in the data. This can aid historians
in gathering knowledge or forming an impression of a particular aspect of the
dataset. To give an example: one of the smaller datasets being ingested in this
study is the King’s Inns dataset, which provides information regarding Irish
students who attended the Inns of Court in London in the pre-Cromwellian
era, 1603–1633. The records are divided into four datasets, one for each of
the four preliminary colleges that students could attend to gain admittance
to King’s Inns in Dublin, namely Gray’s Inn, Lincoln’s Inn, Inner Temple,
and Middle Temple. Each dataset has approximately 250 records. While the
datasets may have slightly diefrent relationships, they all possess the
structure of the property graph model shown in Figure 3.</p>
          <p>All Irish students that attended the institution during this period were
obliged to be sponsored either by their family or by an afluent member of their
community. Therefore, by analyzing each dataset by area as it is absorbed
into Neo4j, a pattern emerges of the most wealthy areas during this time
period. The Lincoln’s Inn dataset linked students to 25 diefrent counties
represented by the blue nodes, with the more prominent areas having more
student nodes linked to them. This can be observed in Figure 5. When the
fourth dataset was added to the database a diefrent visualization can be seen
in Figure 6.</p>
          <p>
            After absorbing all four datasets, 399 nodes and 359 relationships are
present in the database. Clear areas of prominence have emerged, allowing
the user to gain insight into areas of influence during this period:
NoSQL databases have grown in popularity over the last decade, with
MongoDB emerging as the forerunner for large-scale data management and
processing. The advantage of adopting a collection-based approach to data
administration is that each set of key-value pairs can be accessed in a flexible
manner. This allows a collection to be composed of similar but diversified
documents, and makes the storage and retrieval of semi-structured records
more convenient for the user. Another benefit is that it is possible to
interact with the MongoDB engine using the JavaScript programming language,
rather than having to learn and master SQL. However, this can be regarded
as a mixed blessing: SQL is a standardized programming language, designed
for processing data stored in relational database management systems.
Technically, it is less complicated to connect diefrent relational DBMS than it is
to connect relational systems with NoSQL databases
            <xref ref-type="bibr" rid="ref25 ref26 ref36 ref40 ref8">(Vathy-Fogarassy and
Hugyák, 2017)</xref>
            . In MongoDB, merge join queries and hash join queries,
typically a more eficient algorithm than a nested loop join, are not possible,
rendering the commands for data retrieval rather convoluted.
          </p>
          <p>
            Throughout this paper, it can be observed that the graph database Neo4j
provides an elegant alternative to other NoSQL databases for users
confronted with highly complex relationships and entities. As well as providing a
platform which makes the absorption of records into the database
significantly easier, it also maintains uniformity between the structure of an
individual record and the architecture of the data within the database. With a
graph database schema, schema-instance consistency is required (
            <xref ref-type="bibr" rid="ref28">Pokorný
et al., 2017</xref>
            ). This consistency makes graph databases an ideal tool for digital
archivists and historians who wish to observe patterns in semi-structured
and hybrid data. The Cypher querying language allows the user to easily
establish links and extract linked records for observation. Although there are
limitations to using a graphical database to store complex nested data, these
obstacles can be surmounted by monitoring and applying the rules of data
modeling to ensure the data is reproduced accurately within the database.
As we have demonstrated, MongoDB outperformed Neo4j slightly in terms
of speed of ingestion. However, for historians and those who work in the
digital humanities, the priority will most likely be the accuracy of data
transcription rather than speed, which makes Neo4j the more suitable method
of data processing and management.
          </p>
          <p>
            It cannot be concluded from this that traditional relational database
management systems are ideal for the digitization of all types of historical source
material. When confronted with complex and diverse real-world data,
relational systems suefr from the fact that they cannot capture their inherent
graph structure
            <xref ref-type="bibr" rid="ref9">(De Virgilio et al., 2014)</xref>
            . It is vital to prevent a loss of data
integrity when digitally archiving records, so that each record is precisely
preserved for those who wish to access it at a future point in time. It is therefore
necessary to maintain an accurate representation of the data throughout the
entire process of digital archiving, from modeling the structure of the data,
to inserting the data into the database, to observing the structure within the
database, to retrieving the data. All this makes the analysis of integrity
constraints in NoSQL databases a crucial desideratum. While we hope that this
paper will prove useful as a preliminary appraisal of the usefulness of
nonrelational database technology for the processing of unstructured data, there
is still ample potential for further research, be it in the form of a more
comprehensive evaluation of NoSQL databases, a more detailed breakdown of
domain and query languages used to construct integrity constraints, or
extensive testing on other machines to afirm the validity of our tentative
findings.
          </p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Agnelo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laranjeiro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bernardino</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Using Orthogonal Defect Classification to Characterize NoSQL Database Defects</article-title>
          .
          <source>Journal of Systems and Software</source>
          ,
          <volume>159</volume>
          :
          <fpage>110451</fpage>
          , DOI: 10.1016/j.jss.
          <year>2019</year>
          .
          <volume>110451</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Angles</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>A Comparison of Current Graph Database Models</article-title>
          .
          <source>In 2012 IEEE 28th International Conference on Data Engineering Workshops</source>
          , pages
          <fpage>171</fpage>
          -
          <lpage>177</lpage>
          . DOI:
          <volume>10</volume>
          .1109/ICDEW.
          <year>2012</year>
          .
          <volume>31</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Baak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoek</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Klous</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>A New Correlation Coeficient Between Categorical, Ordinal and Interval Variables With Pearson Characteristics</article-title>
          .
          <source>Computational Statistics &amp; Data Analysis</source>
          ,
          <volume>152</volume>
          :
          <fpage>107043</fpage>
          , DOI: 10.1016/j.csda.
          <year>2020</year>
          .
          <volume>107043</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Bagan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bonifati</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Groz</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>A Trichotomy for Regular Simple Path Queries on Graphs</article-title>
          .
          <source>Journal of Computer and System Sciences</source>
          ,
          <volume>108</volume>
          :
          <fpage>29</fpage>
          -
          <lpage>48</lpage>
          , DOI: 10.1016/j.jcss.
          <year>2019</year>
          .
          <volume>08</volume>
          .006.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Bjeladinovic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marjanovic</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Babarogic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>A Proposal of Architecture for Integration and Uniform Use of Hybrid SQL/NoSQL Database Components</article-title>
          .
          <source>Journal of Systems and Software</source>
          ,
          <volume>168</volume>
          :
          <fpage>110633</fpage>
          , DOI: 10.1016/j.jss.
          <year>2020</year>
          .
          <volume>110633</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Bono</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Historical Seismometry Database Project: A Comprehensive Relational Database for Historical Seismic Records</article-title>
          .
          <source>Computers &amp; Geosciences</source>
          ,
          <volume>33</volume>
          (
          <issue>1</issue>
          ):
          <fpage>94</fpage>
          -
          <lpage>103</lpage>
          , DOI: 10.1016/j.cageo.
          <year>2006</year>
          .
          <volume>05</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Borissova</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Cultural Heritage Digitization and Related Intellectual Property Issues</article-title>
          .
          <source>Journal of Cultural Heritage</source>
          ,
          <volume>34</volume>
          :
          <fpage>145</fpage>
          -
          <lpage>150</lpage>
          , DOI: 10.1016/j.culher.
          <year>2018</year>
          .
          <volume>04</volume>
          .023.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Castelltort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Handling scalable approximate queries over NoSQL graph databases: Cypherf and the Fuzzy4S framework</article-title>
          . DOI:
          <volume>10</volume>
          .1016/j.fss.
          <year>2017</year>
          .
          <volume>08</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>De Virgilio</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maccioni</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Torlone</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Model-Driven Design of Graph Databases</article-title>
          . In
          <string-name>
            <surname>Yu</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Dobbie</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jarke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Purao</surname>
          </string-name>
          , S., editors,
          <source>Conceptual Modeling. ER</source>
          <year>2014</year>
          , volume
          <volume>8824</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>172</fpage>
          -
          <lpage>185</lpage>
          . DOI:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -12206-914.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Edward</surname>
            ,
            <given-names>S. G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Sabharwal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2014</year>
          ). Practical MongoDB: Architecting, Developing, and Administering MongoDB. Apress.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Gacem</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papadopoulos</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Boukhalfa</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Scalable Distributed Reachability Query Processing in Multi-Labeled Networks</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          ,
          <volume>130</volume>
          :
          <fpage>101854</fpage>
          , DOI: 10.1016/j.datak.
          <year>2020</year>
          .
          <volume>101854</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Giabelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malandri</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mercorio</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mezzanzanica</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al. (
          <year>2021</year>
          ).
          <article-title>Skills2Job: A Recommender System That Encodes Job Oefr Embeddings on Graph Databases</article-title>
          . Applied Soft Computing,
          <volume>101</volume>
          :
          <fpage>107049</fpage>
          , DOI: 10.1016/j.asoc.
          <year>2020</year>
          .
          <volume>107049</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>González-Aparicio</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Younas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tuya</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Casado</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Testing of Transactional Services in NoSQL Key-Value Databases</article-title>
          .
          <source>Future Generation Computer Systems</source>
          ,
          <volume>80</volume>
          :
          <fpage>384</fpage>
          -
          <lpage>399</lpage>
          , DOI: 10.1016/j.future.
          <year>2017</year>
          .
          <volume>07</volume>
          .004.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Henderson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Using Graph Databases to Detect Financial Fraud</article-title>
          .
          <source>Computer Fraud &amp; Security</source>
          ,
          <year>2020</year>
          (7):
          <fpage>6</fpage>
          -
          <lpage>10</lpage>
          , DOI: 10.1016/S1361-
          <volume>3723</volume>
          (
          <issue>20</issue>
          )
          <fpage>30073</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Holzschuher</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Peinl</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Querying a Graph Database - Language Selection and Performance Considerations</article-title>
          .
          <source>Journal of Computer and System Sciences</source>
          ,
          <volume>82</volume>
          (
          <issue>1</issue>
          ):
          <fpage>45</fpage>
          -
          <lpage>68</lpage>
          , DOI: 10.1016/j.jcss.
          <year>2015</year>
          .
          <volume>06</volume>
          .006.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Jiménez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diez</surname>
            ,
            <given-names>J. V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ordieres-Mere</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>HOSHIN KANRI Visualization with Neo4j</article-title>
          .
          <article-title>Empowering Leaders to Operationalize Lean Structural Networks</article-title>
          .
          <source>Procedia CIRP</source>
          ,
          <volume>55</volume>
          :
          <fpage>284</fpage>
          -
          <lpage>289</lpage>
          , DOI: 10.1016/j.procir.
          <year>2016</year>
          .
          <volume>08</volume>
          .023.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Jose</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Abraham</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Performance Analysis of NoSQL and Relational Databases With MongoDB and MySQL</article-title>
          .
          <source>Materials Today: Proceedings, 24, Part</source>
          <volume>3</volume>
          :
          <fpage>2036</fpage>
          -
          <lpage>2043</lpage>
          , DOI: 10.1016/j.matpr.
          <year>2020</year>
          .
          <volume>03</volume>
          .634.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Karan</surname>
            ,
            <given-names>K. K.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Visualizing and Searching Relationships Between Academic Papers Using Neo4j Graph Database</article-title>
          .
          <source>PhD thesis</source>
          , Thapar Institute.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Kingdon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nayembil</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richardson</surname>
            ,
            <given-names>A. E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>A Geodata Warehouse: Using Denormalisation Techniques as a Tool Fordelivering Spatially Enabled Integrated Geological Information Togeologists</article-title>
          .
          <source>Computers &amp; Geosciences</source>
          ,
          <volume>96</volume>
          :
          <fpage>87</fpage>
          -
          <lpage>97</lpage>
          , DOI: 10.1016/j.cageo.
          <year>2016</year>
          .
          <volume>07</volume>
          .016.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Le</given-names>
            <surname>May</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Carter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Gehly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Flegel</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , et al. (
          <year>2020</year>
          ).
          <article-title>Representing and Querying Space Object Registration Data Using Graph Databases</article-title>
          .
          <source>Acta Astronautica</source>
          ,
          <volume>173</volume>
          :
          <fpage>392</fpage>
          -
          <lpage>403</lpage>
          , DOI: 10.1016/j.actaastro.
          <year>2020</year>
          .
          <volume>04</volume>
          .056.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Lotfy</surname>
            ,
            <given-names>A. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saleh</surname>
            ,
            <given-names>A. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>El-Ghareeb</surname>
            ,
            <given-names>H. A.</given-names>
          </string-name>
          , et al. (
          <year>2016</year>
          ).
          <article-title>A Middle Layer Solution to Support Acid Properties for NoSQL Databases</article-title>
          .
          <source>Journal of King</source>
          Saud University - Computer and Information Sciences,
          <volume>28</volume>
          (
          <issue>1</issue>
          ):
          <fpage>133</fpage>
          -
          <lpage>145</lpage>
          , DOI: 10.1016/j.jksuci.
          <year>2015</year>
          .
          <volume>05</volume>
          .003.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , et al. (
          <year>2020</year>
          ).
          <source>Graph Classification Algorithm Based on Graph Structure Embedding. Expert Systems with Applications</source>
          ,
          <volume>161</volume>
          :
          <fpage>113715</fpage>
          , DOI: 10.1016/j.eswa.
          <year>2020</year>
          .
          <volume>113715</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Mahajan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blakeney</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zong</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Improving the Energy Eficiency of Relational and NoSQL Databases via Query Optimizations</article-title>
          .
          <source>Sustainable Computing: Informatics and Systems</source>
          ,
          <volume>22</volume>
          :
          <fpage>120</fpage>
          -
          <lpage>133</lpage>
          , DOI: 10.1016/j.suscom.
          <year>2019</year>
          .
          <volume>01</volume>
          .017.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Moeyersoms</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Martens</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Including High-Cardinality Attributes in Predictive Models: A Case Study in Churn Prediction in the Energy Sector</article-title>
          .
          <source>Decision Support Systems</source>
          ,
          <volume>72</volume>
          :
          <fpage>72</fpage>
          -
          <lpage>81</lpage>
          , DOI: 10.1016/j.dss.
          <year>2015</year>
          .
          <volume>02</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Mohmmed</surname>
            ,
            <given-names>A. G. M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Osman</surname>
            ,
            <given-names>S. E. F.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Study on SQL vs. NoSQL vs</article-title>
          .
          <source>NewSQL. Journal of Multidisciplinary Engineering Science Studies</source>
          ,
          <volume>3</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1821</fpage>
          -
          <lpage>1823</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Mosquera</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Piedra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Use of Graph Database for the Integration of Heterogeneous Data About Ecuadorian Historical Personages</article-title>
          .
          <source>In 2018 7th International Conference On Software Process Improvement (CIMPS)</source>
          , pages
          <fpage>95</fpage>
          -
          <lpage>100</lpage>
          . DOI:
          <volume>10</volume>
          .1109/CIMPS.
          <year>2018</year>
          .
          <volume>8625618</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Perçuku</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Minkovska</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Stoyanova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Modeling and Processing Big Data of Power Transmission Grid Substation Using Neo4j</article-title>
          .
          <source>Procedia Computer Science</source>
          ,
          <volume>113</volume>
          :
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          , DOI: 10.1016/j.procs.
          <year>2017</year>
          .
          <volume>08</volume>
          .276.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Pokorný</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valenta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kovačič</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Integrity Constraints in Graph Databases</article-title>
          .
          <source>Procedia Computer Science</source>
          ,
          <volume>109</volume>
          :
          <fpage>975</fpage>
          -
          <lpage>981</lpage>
          , DOI: 10.1016/j.procs.
          <year>2017</year>
          .
          <volume>05</volume>
          .456.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Ravat</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teste</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Trojahn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Eficient Querying of Multidimensional Rdf Data With Aggregates: Comparing NoSQL, RDF and Relational Data Stores</article-title>
          .
          <source>International Journal of Information Management</source>
          ,
          <volume>54</volume>
          :
          <fpage>102089</fpage>
          , DOI: 10.1016/j.ijinfomgt.
          <year>2020</year>
          .
          <volume>102089</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Ruetter</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romero</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vardi</surname>
            ,
            <given-names>M. Y.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Regular Queries on Graph Databases</article-title>
          . In Arenas, M. and
          <string-name>
            <surname>Ugarte</surname>
          </string-name>
          , M., editors,
          <source>18th International Conference on Database Theory (ICDT'15), number 31 in Leibniz International Proceedings in Informatics (LIPIcs)</source>
          , pages
          <fpage>177</fpage>
          -
          <lpage>194</lpage>
          . DOI:
          <volume>10</volume>
          .4230/LIPIcs.ICDT.
          <year>2015</year>
          .
          <volume>177</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Sabharwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Digital Curation in the Digital Humanities : Preserving and Promoting Archival and Special Collections</article-title>
          .
          <source>Chandos Publishing</source>
          , Oxford.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>W. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nelson</surname>
            ,
            <given-names>B. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Felker</surname>
            ,
            <given-names>D. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Durant M.D.</surname>
            ,
            <given-names>T. J.</given-names>
          </string-name>
          , et al. (
          <year>2016</year>
          ).
          <article-title>Evaluation of Relational and NoSQL Database Architectures to Manage Genomic Annotations</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          ,
          <volume>64</volume>
          :
          <fpage>288</fpage>
          -
          <lpage>295</lpage>
          , DOI: 10.1016/j.jbi.
          <year>2016</year>
          .
          <volume>10</volume>
          .015.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Simonini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gagliardelli</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergamaschi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Jagadish</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          (
          <year>2019</year>
          ). Scaling Entity Resolution:
          <string-name>
            <given-names>A Loosely</given-names>
            <surname>Schema-Aware Approach</surname>
          </string-name>
          .
          <source>Information Systems</source>
          ,
          <volume>83</volume>
          :
          <fpage>145</fpage>
          -
          <lpage>165</lpage>
          , DOI: 10.1016/j.is.
          <year>2019</year>
          .
          <volume>03</volume>
          .006.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Taktek</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Thakker</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Pentagonal Scheme for Dynamic XML Prefix Labelling</article-title>
          .
          <source>Knowledge-Based Systems</source>
          ,
          <volume>209</volume>
          :
          <fpage>106446</fpage>
          , DOI: 10.1016/j.knosys.
          <year>2020</year>
          .
          <volume>106446</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Vágner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Store and Visualize EER in Neo4j</article-title>
          .
          <source>In Proceedings of the 2nd International Symposium on Computer Science and Intelligent Control, ISCSIC '18</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , New York, NY. Association for Computing Machinery, DOI: 10.1145/3284557.3284694.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Vathy-Fogarassy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Hugyák</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Uniform Data Access Platform for SQL and NoSQL Database Systems</article-title>
          .
          <source>Information Systems</source>
          ,
          <volume>69</volume>
          :
          <fpage>93</fpage>
          -
          <lpage>105</lpage>
          , DOI: 10.1016/j.is.
          <year>2017</year>
          .
          <volume>04</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Čerešňák</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kvet</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Comparison of Query Performance in Relational a Non-relation Databases</article-title>
          .
          <source>Transportation Research Procedia</source>
          ,
          <volume>40</volume>
          :
          <fpage>170</fpage>
          -
          <lpage>177</lpage>
          , DOI: 10.1016/j.trpro.
          <year>2019</year>
          .
          <volume>07</volume>
          .027.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>Šestak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heričko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Družovec</surname>
            ,
            <given-names>T. W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Turkanović</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Applying K-Vertex Cardinality Constraints on a Neo4j Graph Database</article-title>
          .
          <source>Future Generation Computer Systems</source>
          ,
          <volume>115</volume>
          :
          <fpage>459</fpage>
          -
          <lpage>474</lpage>
          , DOI: 10.1016/j.future.
          <year>2020</year>
          .
          <volume>09</volume>
          .036.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <surname>Šestak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rabuzin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Integrity Constraints in Graph Databases - Implementation Challenges</article-title>
          . In Hunjak, Tihomir; Kirinić,
          <string-name>
            <surname>V. K. M.</surname>
          </string-name>
          , editor,
          <source>Proceedings of Central European Conference on Information and Intelligent Systems</source>
          , pages
          <fpage>23</fpage>
          -
          <lpage>30</lpage>
          . https://urn. nsk.hr/urn:nbn:hr:
          <volume>211</volume>
          :
          <fpage>102684</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <surname>Yoon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <source>leted From a MongoDB</source>
          .
          <volume>10</volume>
          .1016/j.diin.
          <year>2017</year>
          .
          <volume>11</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <given-names>A</given-names>
            <surname>Method</surname>
          </string-name>
          and Tool to Recover Data DeDigital Investigation,
          <volume>24</volume>
          :
          <fpage>106</fpage>
          -
          <lpage>120</lpage>
          , DOI:
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>A Novel Approach Based on Neo4j for Multi-Constrained Flexible Job Shop Scheduling Problem</article-title>
          .
          <source>Computers &amp; Industrial Engineering</source>
          ,
          <volume>130</volume>
          :
          <fpage>671</fpage>
          -
          <lpage>686</lpage>
          , DOI: 10.1016/j.cie.
          <year>2019</year>
          .
          <volume>03</volume>
          .022.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>