<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>An Approach to Population Linkage using Graph Databases</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alan Dearle</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Graham Kirby</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Özgür Akgün</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science, University of St Andrews</institution>
          ,
          <addr-line>St Andrews, KY16 9SX, Scotland</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>We report on a database project which is in the process of linking 29 million vital event records encompassing the entire population of Scotland from 1856 until 1973. Since these records contain no common identifiers, the challenge is to form a pedigree by performing probabilistic linkage over the records. We describe the linkage methodology used to create links between records, for example identifying the birth and marriage records of a single person, and discuss the database technologies employed in the project. A graph database (Neo4j) is used to store both the original vital event records and the links made between them. A metric index is used to find potential links eficiently. Finally, we demonstrate how linkage can be improved by augmenting links based on record distance thresholds with local graph analysis.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data linkage</kwd>
        <kwd>graph databases</kwd>
        <kwd>similarity search</kwd>
        <kwd>metric indexing</kwd>
        <kwd>metric search</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The SHiPP Scotland project [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is in the process of linking the contents of the civil registers of
births, marriages and deaths for Scotland covering 1856-1973. This project aims to undertake
a family reconstitution exercise which will encompass the entire population of Scotland over
these 12 decades: some 14 million births, 11 million deaths and 4 million marriages.
      </p>
      <p>The records have now been transcribed, and must be linked in order to reconstruct the
population structure. This involves deciding which particular subsets of records should be
linked, and in what way, and selecting appropriate linkage algorithms. Although this is not a
massive dataset by modern standards it does present a database challenge, since complex data
relationships must be established by linkage algorithms. The output of this project, intended
primarily as a resource for further research in the social sciences, will be linked pedigrees
containing the individual people represented in the civil records, along with the relationships
between them. This paper will describe an approach to establishing these relationships.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Population Data</title>
      <p>
        In the period of study, the Scottish birth records collected by the General Register Ofice for
Scotland include the following fields (reproduced from [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]):
• register entry number, year, registration district number and sufix, child’s forename(s)
and surname, child’s sex, date and place of birth, mother’s forename(s), surname and
maiden surname, father’s forename(s) and surname, parents’ date and place of marriage,
father’s occupation.
      </p>
      <sec id="sec-2-1">
        <title>Death records include the following fields:</title>
        <p>• register entry number, year, registration district number and sufix, deceased’s forename(s)
and surname, deceased’s sex, date, place and cause of death, deceased’s date of birth (or
their age at death), deceased’s occupation, deceased’s marital status, deceased’s spouse’s
name and occupation, deceased’s mother’s forename(s), surname and maiden surname,
deceased’s father’s forename(s) and surname, whether deceased’s parents were deceased.
Marriage records include the following fields:
• register entry number, year, registration district number and sufix, groom’s forename(s)
and surname, bride’s forename(s) and surname, date and place of marriage, religious
denomination, bride and groom’s dates of birth (or their ages at marriage), bride and
groom’s addresses, bride and groom’s occupations, bride and groom’s previous marital
status, bride and groom’s mothers’ and fathers’ forenames, surnames and maiden
surnames, bride and groom’s fathers’ occupations, whether bride and groom’s parents were
deceased.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Linkage Methodology</title>
      <p>
        Probabilistic linkage is the process by which entries in database records may be determined to
be related to the same underlying entity [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ], without any common identifier. A probabilistic
linkage process must account for errors, inconsistencies and omissions in the data, and also
ambiguity where there are alternative possible links.
      </p>
      <p>We distinguish between entity linkage, in which a single individual is identified as appearing
in multiple records, and relationship linkage, in which relationships between diferent individuals
appearing in multiple records are established. Our approach involves both entity linkage and
relationship linkage. Examples include:
• entity linkage: linking a woman’s birth record to her marriage records, or linking a
man’s birth record to the death records of his daughters
• relationship linkage: linking the birth records of full siblings
The amount of relevant information available on which to make linking decisions varies with
the type of link. In the first entity linkage example above, the fields containing the woman’s
mother’s names and her father’s names might be compared with the mother’s and father’s names
on a marriage record. The relationship linkage example might be performed by comparing the
values of the fields containing the parents’ names and the places and dates of their marriages.</p>
      <sec id="sec-3-1">
        <title>3.1. Comparing Records</title>
        <p>
          Comparator functions are used to determine the distance (i.e. lack of similarity) between two
records, by comparing corresponding pairs of fields drawn from the vital event records. Each
comparator function operates over selected fields of the records. A base comparator compares a
single field of one record with a single field of another record. A composite comparator compares
two records by combining a set of base comparators, perhaps each with a separate weighting.
For example, a simple composite comparator for comparing birth and death records—to link the
birth records of individuals to their own death records—may be defined as follows:
-(ℎ, ℎ) =
-(ℎ. , ℎ. ) +
-(ℎ., ℎ.)
We employ true metrics as comparators, in order to access the power of metric indexing [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. A
metric is a comparator that satisfies the postulates of non-negativity, identity, symmetry, and
triangle inequality. This excludes some functions that are often referred to as distance metrics
(such as Jaro-Winkler [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]), but do not satisfy the triangle inequality, and are therefore not true
metrics. We discuss the use of metric indexing techniques in Section 5.
        </p>
        <p>Initial linkage decisions are made by comparing the distance between each candidate pair
of records, calculated by the appropriate composite comparator, with a predetermined
distance threshold. Thresholds for each type of linkage are calibrated using known ground-truth
examples. Initial linkage decisions are refined later in the process, as discussed in Section 6.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Linkage Opportunities</title>
        <p>Each record contains information on a diferent number of individuals, depending on the record
type:
• three on a birth record (child, mother, father)
• six on a marriage record (bride, groom, two sets of parents)
• four on a death record (deceased, spouse, deceased’s parents)
• The combinations shown in red are logically not possible, since at most one birth and
death event can be recorded for an individual. For example: the child on one birth record
cannot be linked to the child on another birth record.
• The combinations shown in pink are also logically excluded, due to the sexes of the
individuals. For example: the groom on a marriage record cannot be linked to the mother
on a birth record.
• The combinations shown in green are logically possible. Those shown in dark green
involve a significant amount of common information between the records, allowing links
to be made with greater confidence than those shown in light green. We denote the
former strong linkage opportunities.</p>
        <sec id="sec-3-2-1">
          <title>Overall, we identify 64 diferent types of logically possible entity linkage.</title>
          <p>There are many logically possible types of relationship linkage. It is not necessary to
perform linkage to establish the most straightforward types, such as parent-child and
spousespouse links, since they are captured within a single record. Conversely, for relatively obscure
linkages such as linking a person’s birth record to their great-aunt’s death record, there is
unlikely to be suficient common information to make the linkage feasible. The most practical
type of relationship linkage is sibling linkage: while siblings are not recorded on any single
record, there is a significant amount of information in common between the birth, death or
marriage records of siblings.</p>
          <p>Figure 2 lists the opportunities for relationship linkage between siblings on two records.
While all combinations are logically possible, the feasibility of establishing links varies with
the amount of common information. Those shown in dark green involve include the names
of both sets of parents for a potential sibling pair, in addition to temporal and geographical
information. The combinations shown in light green are less likely to be feasible to establish,
since less common information is available.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. The Linkage Process</title>
      <p>In our prototype, we define a number of aspects common to the various types of linkage, in a
generic linker function. This takes as input either one (when linkage is over a single type of
event record) or two sets of records (when linkage is over diferent types of event record). For
example, one set of records is supplied for linkage of birth records of siblings or linkage of the
births of children to the births of their mothers. Two sets are supplied for linkage of the births
of children to their death records.</p>
      <p>The output of the linker is a set of record pairs for which the distance between the records is
below some defined threshold. Each pair is annotated with the type of linkage. A specialised
linker for a particular type of linkage is obtained by configuring the generic linker with a recipe
defining the aspects specific to that linkage:
• the types of the records
• the source(s) of the records (for example, particular files, or a database)
• the fields to be compared
• the distance threshold
• any link viability constraints (for example, temporal rules such as death occurring after
birth, or minimum age at marriage)
For example, birth-bride-entity linkage involves linking the birth records of female children to
their marriage records. The recipe includes the following:
• types: birth records, marriage records
• sources: database queries to retrieve the records
• fields :
– birth records: forename, surname, mother_forename, mother_maiden_surname,
father_forename, father_surname
– marriage records: bride_forename, bride_surname, bride_mother_forename,
bride_mother_maiden_surname, bride_father_forename, bride_father_surname
• threshold: 0.78 (as determined by previous experimentation)
• constraints: child sex is female, marriage date after birth date, diference between dates
greater than 14 years and less than 120 years
The first stage of the linkage process involves running a linker for each of the linkages
highlighted in dark or light green in Figures 1 and 2. These initial linkage results are then refined as
discussed in Section 6, and finally combined (not discussed in this paper).</p>
    </sec>
    <sec id="sec-5">
      <title>5. Database Requirements</title>
      <p>This project has some unique database requirements:
1. the storage of the original 29 million source records,
2. the ability to eficiently search the records by distance (for various definitions of distance)
in order to find records for linking,
3. the storage of the links formed between these records, along with meta data including the
linker used, distances between records and/or probabilities of each link being correct, and
4. (as described below) the ability to navigate the database to analyse and improve linkage
quality.</p>
      <p>For requirements 1 and 3, any traditional relational database could be employed. This would
not be suitable for requirement 4 however, due to the inability of relational systems to perform
transitive queries eficiently, while requirement 2 cannot be met by any conventional database.</p>
      <p>
        To address requirements 1, 3 and 4 we use the Neo4j graph database [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]; to address requirement
2 we employ a metric index, BitPart [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This addresses the issue that comparing each record
with every other record would be prohibitively expensive. Blocking [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is sometimes employed
to avoid polynomial complexity, but it has the disadvantage of unavoidable false negatives [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The BitPart metric index creates a set of inclusion zones encoding the inclusion of data
points in a set of database partitions in a binary fashion with respect to a set of reference
points. A relatively small number of reference points (in the order of 20-40) is enough to
characterise the search space, and each metric query only requires distances to be calculated to
the reference points. For large datasets such as the Scottish vital event records, this represents a
large performance increase. Furthermore, the index is highly compressed (as a set of bits) and
obviates the need to directly interact with the stored database records when making queries.</p>
      <p>When linkage is performed between two datasets, the larger set is loaded into the metric
index. The composite comparator and the threshold are supplied. Each record from the smaller
set is used to query the BitPart index, and any records with a distance from the query term
below the threshold are retrieved and recorded in the database, as described below.</p>
      <p>We use Neo4j to store both the original source records and the links formed between them.
Such a database is well suited for this purpose, as it supports attributed relationships between
nodes of arbitrary type. The original vital event records (birth, death and marriage records) are
stored as Neo4j nodes containing the field data from the original records. Thus the database
initially contains a number of unconnected data records containing the original data. The
linkage process involves linking these records by creating relationships between them.</p>
      <p>Once a linker has found links between records using the BitPart index, it encodes the links
between these original records as relationships between the stored nodes in the graph database.
Each of the relationships has attributes which encode the name of the linker used to form the
link (and thus all the provenance in the source code), the distance between the nodes (from the
metric search), and the type of relationship (for example mother, father, sibling, entity (same
individual on two records)). In some cases additional information is stored in the relationship
attributes to disambiguate the links established (for example, in the case of entity links, which
individuals on the records are involved). Figure 3 shows a link formed between a birth record
and a corresponding death record. The relationship shows the metric distance between the
records, and the record identifiers (571766 and 571764).</p>
      <p>In addition to the ease with which linkage information may be encoded, Neo4j facilitates
querying such graph structures, using the rich Cypher query language. For example, the nodes
shown in Figure 3 could be retrieved via a number of queries specifying varying degrees of
detail, as shown in Figure 4.</p>
      <p>MATCH (d:Death) WHERE ID(d) = 571764 RETURN d;
MATCH (b)-[r]-(d) WHERE r.linker = "Birth-Death-Entity" RETURN b,r,d;
MATCH (b:Birth)-[r]-(d) WHERE r.distance = 0.013 RETURN b,d;
MATCH (b:Birth)-[r]-(d:Death) RETURN b,r,d;
MATCH (b:Birth)-[r linker:"Birth-Death-Entity"]-(d:Death) RETURN b,r,d;
The first example matches a node based on its identity; the second matches based on the linker
attribute of a relationship; the third uses the distance attribute of a relationship. Note that the
types of the nodes may be specified, or not, and that the queries may return multiple nodes,
relationships or a combination of both. The final example returns all nodes and relationships
between birth and death records that were created by the Birth-Death-Entity linker.</p>
      <p>Once all linkers have completed, the original graph containing only the source records has
been transformed into a labelled graph containing relationships between the nodes. Figure 5
shows a simplified example of the output, omitting many of the links. The purple and blue nodes
in the diagram represent birth and marriage records respectively. The purple nodes denote the
birth records of siblings, linked with child-child relationships established by a Birth-Birth-Sibling
linker. Each birth record is linked to a marriage record by (redundant) mother-bride, father-groom
and child-couple relationships. Such redundant links may seem wasteful, but they can be formed
by diferent linkers, comparing diferent fields, thus adding confidence (or otherwise) to the
links that have been formed. In the diagram, the nodes are labelled with the identity of the
father of the children and the groom in the marriage record. We can see that, in this case, since
the identities match, the linkers have made consistent assertions about the relationships. In this
example the ground truth is extracted from a known dataset.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Linkage Refinement</title>
      <p>Linkage is followed by a refinement process which aims to improve the quality of the linkage.
The graphs that are created in the linkage phase may contain various classes of error:
1. Errors of omission: where a link should exist in the linkage graph, but the linkers have
not established the link.
2. Errors of inclusion: where a link has been created in error. For example, a cluster of
siblings may form one graph but in fact represents multiple families.
3. Uniqueness constraint errors: where a single relationship should exist in the graph but
more than one has been established. For example, if entity links are established between
the child on one birth record and multiple death records, at least one link must be in error.</p>
      <sec id="sec-6-1">
        <title>6.1. Detecting Errors</title>
        <p>We use Cypher queries to detect such errors. In some cases we can then decide how to rectify
the error based on information in the retrieved nodes and their neighbours in the relationship
graph.</p>
        <p>Some errors can be detected by considering transitive relationships, and searching for cases
where the logical consequences of transitivity are not reflected in the graph. For example, the
relationship is-a-sibling-of is transitive: if person A is a sibling of person B, and B is a sibling of
C, then A must be a sibling of C. Similarly, if a person recorded on record X is the same person
as one recorded on record Y, and the person recorded on Y is the same person as one recorded
on record Z, then the person recorded on X must be the same person as recorded on Z.
The simplest sub-graph to which a transitivity check can be be applied is a triangle of three
nodes:
• If the sub-graph is not connected, i.e. at least one node is not connected to another, the
triangle can be ignored.
• If the sub-graph is fully connected, the transitivity condition is satisfied, and no error is
apparent.
• If the sub-graph contains exactly two relationships (edges), we consider this an open
triangle, which indicates an inconsistency in transitivity, arising from either an error of
omission or an error of inclusion.</p>
        <p>Open triangles can be identified via a simple Cypher query, as illustrated in Figure 6.
MATCH (x:Birth)-[:SIBLING]-(y:Birth)-[:SIBLING]-(z:Birth)
WHERE NOT (x)-[:SIBLING]-(z) RETURN x,y,z
Open triangles may contain nodes of the same type (e.g. three birth records representing
siblings), or nodes of diferent types (e.g. a birth, death, and marriage record all representing the
same person). More complex patterns of non-transitivity may also be identified, for example
chains of nodes that are not fully connected. Searching for open triangles is suficient to identify
such patterns, although consideration of the surrounding graph may be beneficial in deciding
how to rectify errors.</p>
        <p>For some types of linkage we can define constraints such that a record should not be linked
to more than one other record. Examples of potential errors include:
1. Multiple births corresponding to one death
2. Multiple deaths corresponding to one birth
3. Multiple parents’ marriages corresponding to one child birth
4. Multiple bride births/deaths corresponding to a single marriage
Such errors may be identified via Cypher queries such as that illustrated in Figure 8.
MATCH (b:Birth)-[r linker:"Birth-Death-Entity"]-&gt;(d:Death) WITH d, count(r) as countlinks
WHERE countlinks &gt; 1
RETURN d;</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Correcting Errors</title>
        <p>We cannot determine with certainty whether an open triangle indicates an error of omission
or inclusion. In the example of Figure 7, it is possible that there should be a link between A
and B (omission), or that there should not be links between A and C, and B and C (inclusion).
Given that two out of three links have been established, it may be reasonable to conclude that
the error of omission is more likely. The distances AC and BC may also be considered: if they
are relatively high and close to the threshold, an error of inclusion may be more likely.</p>
        <p>We can also look for potential graph isomorphisms when we consider entity and relationship
linkage graphs together. For example, Figure 9 shows an open triangle for sibling linkage using
birth records. Since an entity link has also been established between each birth record and
a corresponding death record, and those death records form a complete triangle for sibling
linkage, we can have greater confidence that the missing birth sibling link should be present.
In the example above, the presence of entity links is used to support the existence of missing
relationship links. Conversely, if both relationship triangles were complete but one of the entity
links was missing, its existence could be inferred.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions</title>
      <p>
        We have reported on a project to link 29 million vital event records. We have sketched how
the linkage code operates and its use of two diferent database technologies. A state of the art
metric index is used to perform distance calculations over the records and to determine which
records should be linked. A graph database is used to store both the original record and the
links created between them, and to facilitate analysis of the graph structures to refine linkage
decision. The project code is published [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Work continues on aspects including:
1. Combining linkages: we are investigating approaches to synthesising a coherent
population structure from the results of the large number of diferent linkages (74 linkage
types were identified in Section 3). We aim to exploit the high degree of redundancy
in the structure of the data, in that genealogical relationships are encoded in multiple
ways (for example, sibling relationships may be deduced independently from birth, death
and marriage records). Although some of the types of linkage are dificult to perform
in isolation due to a limited amount of common information between the records, we
take encouragement from the success of ensemble methods in machine learning; multiple
weak linkages may still improve overall results.
2. Calibrating parameters: our approach requires various parameters to be configured,
in particular the distance thresholds for each type of linkage. In order to establish
appropriate parameter values we need ground truth data against which linkage quality can
be optimised. The scarcity of ground truth is a general problem in automated approaches
to population reconstruction, due to the relatively low number of available population
datasets and the high cost of expert annotation. Furthermore, even when available, such
ground truth may be incomplete, biased, and contain significant numbers of errors [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
We are fortunate to have access to a relatively large ground truth dataset from the CEDAR
group in Umeå, Sweden [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This data is encoded in a similar form to the Scottish data
and thus permits us to experiment with some confidence that results can be extrapolated
to the Scottish data. Another promising avenue to obtaining ground truth is statistically
faithful simulation of large-scale populations [
        <xref ref-type="bibr" rid="ref10 ref12">10, 12</xref>
        ].
3. Evaluating approaches: similar to the calibration issues discussed above, it is also
desirable to be able to experiment with various linkage approaches. Again, access to
realistic ground truth data at scale is essential to evaluation.
4. Delivering results: the simplest approach to delivery of a complete population
reconstruction would be to collate all of the final links into a unified genealogical structure,
apply any necessary anonymisation, and make this data available to approved researchers.
However, the utility of this resulting data would be dependent on having made optimal
decisions on linking thresholds and graph refinements throughout the process. Furthermore,
the trade-of between Type 1 and Type 2 errors in the population structure might, ideally,
be set diferently depending on the research question. We are considering approaches to
allowing some degree of control over this trade-of by the end-user researcher.
This work was supported by ESRC grants ES/K00574X/2 “Digitising Scotland”, ES/L007487/1
“Administrative Data Research Centre – Scotland”, ES/S007407/1 “Administrative Data Research
Centres 2018” and ES/W010321/1 “2022-2026 ADR UK Programme”.
      </p>
      <p>We thank our collaborators on the SHiPP project: Chris Dibben, Peter Christen, Lee
Williamson and Eilidh Garrett. We are also indebted to our colleagues in Umeå University who
created the ground truth used in this project, especially Maria Larsson and Pär Vikström.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] SHiPP,
          <source>Scottish Historic Population Platform (SHiPP)</source>
          ,
          <year>2023</year>
          . URL: https://www.scadr.ac. uk/our-research/shipp.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Ö.</given-names>
            <surname>Akgün</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dearle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kirby</surname>
          </string-name>
          , E. Garrett,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dalton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Christen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dibben</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Williamson</surname>
          </string-name>
          ,
          <article-title>Linking Scottish Vital Event Records using Family Groups, Historical Methods: A Journal of Quantitative and Interdisciplinary History (</article-title>
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          . URL: https://doi.org/10.1080/ 01615440.
          <year>2019</year>
          .
          <volume>1571466</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I. P.</given-names>
            <surname>Fellegi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Sunter</surname>
          </string-name>
          ,
          <article-title>A Theory for Record Linkage</article-title>
          ,
          <source>Journal of the American Statistical Association</source>
          <volume>64</volume>
          (
          <year>1969</year>
          )
          <fpage>1183</fpage>
          -
          <lpage>1210</lpage>
          . URL: https://doi.org/10.1080/01621459.
          <year>1969</year>
          .
          <volume>10501049</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Christen</surname>
          </string-name>
          , Data Matching:
          <article-title>Concepts and Techniques for Record Linkage</article-title>
          , Entity Resolution, and Duplicate Detection, Springer Publishing Company, Incorporated,
          <year>2012</year>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>642</fpage>
          -31164-2.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Ö.</given-names>
            <surname>Akgün</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dearle</surname>
          </string-name>
          , G. Kirby,
          <string-name>
            <given-names>P.</given-names>
            <surname>Christen</surname>
          </string-name>
          ,
          <article-title>Using Metric Space Indexing for Complete and Eficient Record Linkage</article-title>
          , in: D.
          <string-name>
            <surname>Phung</surname>
            ,
            <given-names>V. S.</given-names>
          </string-name>
          <string-name>
            <surname>Tseng</surname>
            ,
            <given-names>G. I.</given-names>
          </string-name>
          <string-name>
            <surname>Webb</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ganji</surname>
          </string-name>
          , L. Rashidi (Eds.),
          <source>Advances in Knowledge Discovery and Data Mining</source>
          , Springer International Publishing,
          <year>2018</year>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>101</lpage>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -93040-
          <issue>4</issue>
          _
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.</given-names>
            <surname>Winkler</surname>
          </string-name>
          ,
          <article-title>String Comparator Metrics and Enhanced Decision Rules in the Fellegi-Sunter Model of Record Linkage</article-title>
          ,
          <source>Proceedings of the Section on Survey Research Methods</source>
          (
          <year>1990</year>
          ). URL: https://eric.ed.gov/?id=
          <fpage>ED325505</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <issue>Neo4j</issue>
          ,
          <string-name>
            <given-names>Neo4j</given-names>
            <surname>Graph Data Platform</surname>
          </string-name>
          ,
          <year>2023</year>
          . URL: https://neo4j.com.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dearle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Connor</surname>
          </string-name>
          ,
          <article-title>Bitpart: Exact Metric Search in High(er) Dimensions, Information Systems 95 (</article-title>
          <year>2021</year>
          )
          <article-title>101493</article-title>
          . URL: https://doi.org/10.1016/j.is.
          <year>2020</year>
          .
          <volume>101493</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dearle</surname>
          </string-name>
          , G. Kirby, T. Dalton, Population Linkage,
          <year>2023</year>
          . URL: https://github.com/stacs-srg/ population-linkage.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Dalton</surname>
          </string-name>
          ,
          <article-title>Evaluating Data Linkage Algorithms with Perfect Synthetic Ground Truth</article-title>
          ,
          <source>Ph.D. thesis</source>
          , University of St Andrews,
          <year>2022</year>
          . URL: https://doi.org/10.17630/sta/247.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>M. J. Wisselgren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Edvinsson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Berggren</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Larsson</surname>
          </string-name>
          ,
          <source>Testing Methods of Record Linkage on Swedish Censuses, Historical Methods: A Journal of Quantitative and Interdisciplinary History</source>
          <volume>47</volume>
          (
          <year>2014</year>
          )
          <fpage>138</fpage>
          -
          <lpage>151</lpage>
          . URL: https://doi.org/10.1080/01615440.
          <year>2014</year>
          .
          <volume>913967</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Dalton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kirby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dearle</surname>
          </string-name>
          , Ö. Akgün,
          <source>ValiPop: a Micro-simulation Model for Generating Synthetic Genealogical Populations</source>
          ,
          <year>2023</year>
          . URL: https://github.com/stacs-srg/ population-model.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>