<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>On the Choice of Vocabularies for Archival Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>VeruskaZamborlini</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leon van Wissen</string-name>
          <email>charles.van.den.heuvel@huygens.knaw.n</email>
          <email>l.vanwissen@uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AlbertMeroño-Peñuel a</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Charlesvan</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Archival Data, Vocabulary Reuse, Provenance</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Humanities, University of Amsterdam (UvA)</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Informatics Department, King's College London</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Knowledge &amp; Art Practices, Huygens Institute</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Ontology &amp; Conceptual Modeling Research Group (NEMO), Federal University of Espírito Santo (UFES)</institution>
          ,
          <addr-line>ES</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Time and time again researchers are faced with the issue of choosing the most appropriate vocabulary for publishing archival data, particularly in the Semantic Web. Options range from most popular ones, such as schema.org, or more comprehensive ones such as CIDOC-CRM. There are pros and cons in each of them, but no guidelines on how to decide about it. This paper aims at providing some guidance based on an analysis of data at hand but also the requirements of data providers and users. For example, archives often refrain to add much interpretation by providing simple access to categorised documents with simple annotations such as person's names or location names. Moreover, the archival data as well as its digitized versions may present subtleties, such as is the document original or has it been modified, simplified, copied or translated, which is often omitted. Therefore, depending on how much detailed information is actually accessible, but also what are the requirements of the data providers/users, the data can be ”placed” at diferent levels of content literacy/granularity and provenance. By having a clear understanding of the possibilities and limitations of each level, the choice of one or more vocabularies are down to the one(s) that should provide the necessary expressiveness. Naturally, choosing more than one vocabulary also requires some integration task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>When publishing data in the Semantic Web, researchers are often faced with the challenge of
choosing the most appropriate vocabulary. This is far from a simple task for a few reasons: (i)
it is not easy, and even not recommendable, to build it from scratch; (ii) there are a variety of
options to choose from and (iii) it profoundly impacts the expressiveness, the interpretation
and the connectivity of your data on the web.</p>
      <p>This problem is not exclusive for archival data, but this paper focuses specifically on this kind
of data that come from historical documents, records, and artifacts preserved in various national
and regional archives. Researchers often struggle with dilemmas: choosing for the most popular
nEvelop-O
vocabulary, such asschema.org, or more comprehensive ones, such as tChIeDOC-CRM, yet
something in between. An extensive list of vocabularies can be found on the report of a study by
the European Commission for management, exchange and publication of archival1d].aTthai[s
study highlights that (i) providing data using the Semantic Web framework has the potential to
make data more accessible and interoperable, and that (ii) although a consensus on a body of
standard vocabularies to use exists, a heterogeneity of practices needs to be acknowledged. In
fact, there are pros and cons in each vocabulary, but no guidelines on how to decide which ones
to choose.</p>
      <p>This paper aims at providing some guidance for this choice by looking at data in two
dimensions of content granularity/literacy and provenance. By having a clear understanding of
the possibilities and limitations of each level in those dimensions, the choice of one or more
vocabularies is down to the one(s) that should provide the necessary expressiveness, also taking
into account the conventions and most commonly used vocabularies in the field. An analysis to
position the data in these dimensions not only considers the features of the data but also the
requirements of both data providers and users.</p>
      <p>As an example for requirements of data providers, archivists often only provide access to
categorised documents with simple annotations such as person’s names or locations (called
indexes), leaving out deeper and possibly arguable interpretations of the archival resources.
After all, originally these indexes on the material were made for easier access to the material
for researchers and genealogists that would further read and investigate the contents of the
respective material anyway. Time-wise this was possible, because only a few records had
to be inspected. As an example for requirements of data users, nowadays, researchers and
data scientists are interested in analysing thousands of archival documents at the same time,
requiring other means of modeling the data and also of obtaining it (e.g. through Handwritten
Text Recognition and Natural Language Processing).</p>
      <p>Common features of historical data in general are uncertainties and incompleteness.
Interpretation is mostly unavoidable and users may require detailed description of it. For example,
information about the origin and the manipulations of the archival resources in multiple
versions, such as copies, simplifications, translations etc. are often omitted, because it is not (fully)
known or because the providers lack the means to express what is known. However, those
processes may add several layers of interpretation that ideally should be made accessible.</p>
      <p>Naturally, it may be necessary to consider using more than one vocabulary, which would
bring its own challenges. Semantic interoperability is far from obvious, and might require
a proper ontological analysis of the choices (often implicitly) made in each vocabulary to
avoid misleading integration. For example, 2in] t[he authors show that concepts such as
period and temporal entity do not have the same meaninCgIDiOnC-CRM, OWL-Time and
PeriodO vocabularie1sthrough an ontological analysis. Moreover, vocabularies may choose
diferent modeling styles that are also not trivial to integrate. For example, evenCtI-DbOasCe-d
CRM models birthc(rm:E67_Birth) as a class while attribute-bascehdema.org models birth
(schema:birthDate) as a property that expects a literal as value. Mistaken alignments of
classes and/or properties lead to logically invalid models or misleading conclusions, which,
even worse, are not detectable by a reasoner.
1Respectively accessible awtww.cidoc-crm.org, www.w3.org/TR/owl-timeandhttps://perio.do/e.n/</p>
      <p>The strateg2yhere proposed was developed in the context of the Golden Agents p3roject
aimed at analyzing interactions between the production and consumption of cultural goods (e.g.
paintings, books, silverware) during the long Dutch Golden Age (ca. 1570s-1750) using linked
data from various cultural heritage institutions. To understand this consumption and explain
the flourishing of Amsterdam’s cultural industries, millions of digitized archival documents and
metadata records from the Amsterdam City ArchSiAveAs)(with data on producers, consumers
and the goods that were interchanged were brought together as linked data.</p>
      <p>In the remainder of this paper, Sect2iopnresents the related work, Sect3iopnresents a
case study motivating the data analysis strategy proposed in 4S,eacntdioSnection5 presents
discussion and future work.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related Work</title>
      <p>
        Almost every knowledge engineering methodology identifies the reuse of existing ontologies and
vocabularies as a crucial step when modeling and publishing linked data on th4e, w5,e6b, [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
However, there is a broad landscape in how exactly this reuse must be en9a]c:tfoedr e[xample,
some methodologies recommend a direct or indirect re1u0s],eo[r reuse based on design patterns
[11, 12]. More recently, thFeAIR data principles1[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] have established a framework for making
data Findable, Accessible, Interoperable and Reusable; with nanopublica1t4i,o1n5s] [being a
successful implementation that distinguishes betweeanssaenrtion (i.e. observations, literal
content)p,rovenance (i.e. backing references, workflows; typically modeled using the WP3RCOV
vocabulary 1[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) and publication info (i.e. metadata, documents/collections). Although this
addresses some of our challenges, nanopublications (a) have mainly been used in the life sciences
domain, where there is often little space for interpretation compared to archival data; (b) do not
address data reconstruction explicitly, as they often relate to the result of contemporary scientific
processes; and (c) do not establish diferent degrees or levels of required provenance. Similarly
to nanopublications, the Dutch Historical CensCusEeDsA( R) [17] deploy a similar approach
for historical data that distinguishes betrawweeonbservationsa,nnotations that are subject to
interpretation, asntdatistical data cubes (using the RDF Data Cube vocabula1r8y])[. However,
this focuses on the specific domain of historical statistics, rather than ofering an abstract
framework for any archival data; and records only provenance of data transformation processes,
rather than historical processes preceding archival documents. Various vocabularies exist that
are specifically designed to model cultural heritage objectsC(eID.gO.C-CRM [19]), bibliographic
information (e.g.FRBR [20]), and glossaries for archival recor2d1s].[ Unfortunately, these
are all specific to their corresponding domains, and do not extend to more general notions of
provenance.
2A first version was discussed at LODLAM Summit 2020h(ttps://lodlam.net/challenge-entr)ieasn/d later at the RSA
2022 Conference [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
3Golden Agents: Creative industries and the making of the Dutch GoldewnwAwg.egoldenagents.orfgunded by
Dutch Research Council (NWOw)ww.nwo.nl.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Case Study</title>
      <p>An archival organization could provide digitized information about archival records as follows:
1. indexed information, i.e. just key information and metadata of the record (not the full
text), e.g. a notice of marriage between John and Mary on a certain date;
2. scanned books/documents under some classification, e.g. all the marriage banns at the
city hall;
3. index information connected to the physical and (if available) scanned page in which it
appears;
4. full text extraction from a record.</p>
      <p>Let’s consider as a case study a record documenting the notice of marriage between Jan
Ceuler and Susanna de Bock, taken from the Amsterdam City Archives (SAA) and developed
during the Golden Agents project. Ta1bliellustrates raw data available as items 1-4 for the
mentioned example. It shows (1) indexed data including name, location, date and others, it
also shows (2) a scanned page containing the referred record, (3) the URL connection between
the index data and the scanned page and (4) the full text extracted from the record providing
the name, age and location of each person with the intention to marry the other and of their
witness.</p>
      <p>Observe that users are not provided with information on its provenance and reliability. How
and by whom were the indexes and scans crea4taendd published and who extracted the texts
in which ways. And more importantly for provenance (and helpful for end-users as a visual
aid), where in the scan can information be found?
4Due to legacy at the archives, this would be impossible for old indexes. However, nowadays it is very much needed
(e.g. due to any bias that can be present in data) to report on the coming to being of indexes. What is more, it is
good to know whether the data came from the archive itself, any external research project, or by crowdsourcing
and citizen science. Both for attribution as well as to understand the modeling choices or classifications that were
made (and from what perspective).</p>
      <p>In order to provide more reliable data, ideally the users would have access to (i) the source of
the data and (ii) information on the production or transformation processes of the data. Both are
often referred to as provenance data, in addition to what is often called provenance in cultural
heritage, referring to the processes of creation and ownership of the cultural object (see ISO
8000-120: 2016). It should allow for users to inspect and validate the original handwritten
data or the original text or whom came with which interpretations. Thereby, this information
increases the credibility of the data ofered by the archive and the likelihood of being used in
research. Some ways in which more provenance data could be provided are:
1. detailed information on the location of an entire record on a scanned page (e.g. by
including coordinates of bounding boxes or regions, cf. the IIIF Presentation API and its
FragmentSelectors and SVG Selectors);
2. even more detailed information on the location of each indexed information in the record
section on the scanned page or the recognized or transcribed text (e.g. by including
character ofsets);
3. detailed information on the processes of obtaining the scans, the indexed information
and/or the full text and their connections, including any changes made to the (metadata)
record.</p>
      <p>Figure1 displays the scan of a double page spread with the referred notice of marriage record
highlighted within a (red) box, at the bottom right. This selection can be produced by indicating
the record’s xywh pixel coordinates on the scanned page, and be classified as a record section.
The chosen vocabulary in this particular example was inspired by Web Annotation Vocabulary
[22] The raw text of such a section can also be extracted as the content of the record section.
Moreover, in a similar fashion, more specific sub-sections can be created to indicate where a
particular name is mentioned in the record, as depicted in F2i,gaulrsoe with its raw content.</p>
      <p>As for the information on the archiving processes, Fi3gduerpeicts the processes of creating
and digitizing documents as indexes and scans at the Amsterdam City Archives. The two events
depicted more to the left, called coverage events, represent a complex event covering the whole
creation of the documents in a collection or book. Naturally, these events, which happened in
the 17th century, outputs the original documents and precedes the ones of digitization. First
point to highlight is that the current index, provided as open data in XML, was not digitized
directly from the documents, but from another (paper or rolodex) index previously created.
This is important for understanding that any interpretation, at least regarding the early modern
Dutch handwriting, happened in the first process of indexing, not in the second one. Another
point is that the process of scanning the archival resources happened separately. Therefore, the
connection between the indexed documents and the scanned pages in which they appeared,
as well as the extraction and annotation of particular mentions, could only have happened a
posteriori in another complimentary digitization process.</p>
      <p>Now, from the records and sections one can identify references to (supposedly real world)
entities, such as the specific mention of the name Jan Ceuler van Hamborgh.
Fi4guilrluestrates a detailed extraction of references from the record section, including references within
references. Here the vocabulary choices were inspired by the Factoid Prosopography Ontology
[23], according to which any piece of information containing references can be called a factoid,
including references themselves.</p>
      <p>Naturally, one could consider skipping these steps and, for example, connect the record
directly to the person-entity named Jan Ceuler. That only means that intermediate steps were
made, but are left implicit. This simplifies the representation, but also may leave less space to
accommodate disagreements. The extracted references are meant to be less arguable, as they
should contain the literal content without any adjustment of language evolution or translations.
Meanwhile, those references referenttoities of certain types, sometimes also called observations.
At this point, adjustments of the language used are often performed and also reliability of the
data is presumed. It is a moment where one could consider thainttetrhperetations would start,
which is almost inevitable for historical data. Moreover, diferent levels of interpretation can
take place, more or lesdsirectly drawn from the raw-reference data.</p>
      <p>Figure5 shows two interpretations that are not directly stated in the references. The first,
assuming a cultural rule, infers that a marriage event might have happened not earlier than
2-3 weeks after the registered notice of marriage. The second assumes that the reference to the
groom, Jan Ceuler van Hamborgh, might be related to his city of origin, namely Hamborgh
(currently called Hamburg, Germany). It could result in associating Jan Ceuler to a place entity
called Hamborgh (back then) as his hometown. This extracted information could potentially
provide information on Ceuler’s town of origin, even though it was not explicitly stated as his
hometown.</p>
      <p>Finally, Figure6 illustrates another level of interpretation, in which entities that are more
or less directly drawn from the references, are inferred to be the same. It adds to our running
example a baptism record of a child born to a Jan Keulder and Susanna de Bock in April 1668.
Using some identity criteria, it is possible to infer that the groom Jan Ceuler is the same as
the father Jan Keulder (similarly the mentioned bride and mother could be inferred to be the
same). Moreover, it can be concluded (or inferred) that the marriage event that happened few
days after the notice of marriage, must be same as the marriage expected, based on domain
knowledge, to have happened before the baptism of their child.</p>
      <p>By combing the information from two archival resources, wreecaornestructing the life of
these two persons: we are making a reconstruction of the entities based on our observations in
individual records. Again, the processes and decision of getting to such conclusions could be
documented in detail providing arguments to agree/disagree w24i]t.h. [</p>
    </sec>
    <sec id="sec-5">
      <title>4. Strategy</title>
      <p>An strategy for analysing diferent ways of describing archival data was developed while
investigating the issues mentioned in Sect3i.oInt considers two orthogonal dimensions for
data: content literacy/granularity and provenance.2.Tparbolevides a summary illustrating
how the dimensions can be combined in some scenarios.</p>
      <p>First, regarding the provenance dimension, three levels are considered: minimum, detailed
and advanced. They correspond with having as little as possible to very detailed provenance
information. Naturally, there could be more than three.</p>
      <p>Minimum provenance means literally as little as possible, or provenance information is not
a priority. An example can be seen in Tab1lein Section3, where indeed nothing is known
about how the data came to be, as well as Fig6uwrheere nothing is known about how
the entities were reconstructed.</p>
      <p>Detailed provenance describes the means to access the source(s) of the main information.</p>
      <p>For example, the location of a registry in a book, or the precise position of a registry or a
particular mention in a scanned document or in the transcribed text, or yet the specific
sequences of (text) references in which an entity was mentioned. Examples from Section
3 are found in Figure1s to5. Regarding integration/reconstructions, this level would
require some sort of reification, either for each identity link (same as) or for a group of
them (linksets) where the criteria for integration could be described.</p>
      <sec id="sec-5-1">
        <title>Documents/ Collections</title>
      </sec>
      <sec id="sec-5-2">
        <title>Literal Content</title>
      </sec>
      <sec id="sec-5-3">
        <title>Direct Interpretation</title>
      </sec>
      <sec id="sec-5-4">
        <title>Indirect Interpretation</title>
      </sec>
      <sec id="sec-5-5">
        <title>Integration/</title>
      </sec>
      <sec id="sec-5-6">
        <title>Reconstruction</title>
      </sec>
      <sec id="sec-5-7">
        <title>Interpretation</title>
      </sec>
      <sec id="sec-5-8">
        <title>Minimum provenance</title>
      </sec>
      <sec id="sec-5-9">
        <title>Detailed provenance</title>
      </sec>
      <sec id="sec-5-10">
        <title>Advanced provenance</title>
        <p>General information about Information of access to Description of the archiving
documents, eventually the original document, its processes. Who
created/diggrouped in collections under scanned version or literal itized the document? Is the
certain criteria/classifica- transcription. process reliable? Is it the
tion. original, copy or
transcription? Are there damages to
the documents?</p>
      </sec>
      <sec id="sec-5-11">
        <title>Mentions/references to in- Information of access to the Description of the processes</title>
        <p>dividuals, objects, events, mentions/references in the for obtaining the
mentionplaces, types and roles ob- document and chain of ref- s/references. Where they
served literally in a docu- erences inside references. modified, adapted or
transment. lated in the process?
Mentions/references are Information of access Description of the processes
taken as (supposed) facts to (parts of) mention- of ”materialization” of the
(individuals, objects, events, s/references underlie the mentioned facts. Are there
places), for which the observed/interpreted facts. ambiguities or
uncertainrespective entities and their ties? To which degree?
properties are represented. Could the degree change in
face of new facts?</p>
      </sec>
      <sec id="sec-5-12">
        <title>New facts are derived from Information of access to Description of the processes</title>
        <p>the existing ones, i.e. in- (parts of) mentions/refer- of deriving new facts.
direct interpretation of the ences underlie the indirectly Which domain rules and
document. interpreted facts. which type of inference? Is
there probability involved?
Integration/disambiguation Information about which Description of the processes
of mentions to one single properties where used as cri- of
integration/disambiguaindividual concept in several teria for integration. tion. How precise is it? Are
documents. the criteria enough to
individuate? If not, which is the
degree of confidence? Was
it validated?
Advanced provenance Describes the processes of obtaining the data. For example, the
processes of digitization of a collection or the processes of integrating and disambiguating
mentioned entities. It may also consider describing information regarding the reliability
of the data, including uncertainties inherent to the data, to the inference rules or yet
the ones introduced in the process of obtaining the data. Who created the document?
Is the creator trustworthy (e.g. by afiliation or status)? Is it a original document or
has it been copied/translated? There has been any damage/modification to the original
document? There can be errors introduced in the process? Is it likely that someone may
have lied, for example, about the age, gender or place of origin? Particularly regarding
uncertain assertions about individuals, it is still a challenge how this should be precisely
represented, either qualitatively or quantitatively. Regarding integration/reconstructions
in particular, this would require for identity links (same as) to be reified and qualified
with a metric often referred to as similarity.</p>
        <p>Second, regarding the content literality/granularity dimension, five levels are considered:
Documents/Collections: includes more general data about documents and their digitization,
for example, a book containing baptism registries or several books registering baptism
events. Here data is not about the content of one document in particular, but regards who,
how and when the documents were (re)produced and preserved or their compilation in
an index. It may indicate in which page(s) of which book a document/registry is, or even
its precise position in a scanned page. It also might concern, for example, the church who
produced a baptism registration book, which would not only indicate the location but
also the denomination of the involved people. The location could indicate if the dates are
to be interpreted as Gregorian or Julian cal5e.ndar
Literal content: includes descriptions of the content of document/registries and entities
mentioned in it as literally as possible, ideally as references in the text. It may also indicate
where exactly the mentions are located in a scanned page. The annotation of such entities’
mentions may already indicate the kind of entity mentioned or the role played by them.</p>
        <p>For example, a document mentions names of people and locations.</p>
        <p>Direct interpretation: includes interpretations that are drawn directly from the
information/references in the document, regarding events and the roles of the entities. For
example, a child baptized on a certain date and the parents.</p>
        <p>Indirect interpretation: includes interpretations that are drawn indirectly from the
information in the document, also regarding events and the roles of the entities. For example, the
name of person could indicate his/her profession and/or his/her place of birth. Or even, a
baptism could imply some assumption/approximations about the birth date and location,
or burial and death analogously, according to religion and customs. Traceability regarding
the data that were used as evidence for the conclusions, as well as for the processes and
uncertainties involved so that more detailed provenance could be provided.
5Not all cities in the Netherlands used the same calendar. While Amsterdam used the Gregorian calendar from the
end of the sixteenth century onwards, the city of Utrecht for instance kept on using the Julian calendar until the
end of the seventeenth century.</p>
        <p>Integration/Reconstruction interpretation: includes the identity links connecting
mentions from diferent documents. Traceability regarding the data that was used as evidence
for the conclusions, as well as the processes and uncertainties involved are more detailed
provenance that could be provided.</p>
        <p>It makes clear that a vocabulary choice will directly influence how much detail can be
expressed or how much detail is necessary, and vice versa.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion and Conclusion</title>
      <p>This paper presents a strategy meant to support the choice of vocabularies and vocabulary
reuse for digitization of archival data. It proposes understanding the data by looking into two
orthogonal dimensions: (i) content literality/granularity and (ii) provenance. There is not a
single answer to the question which vocabulary to use, as it will depend on the level of detail
in both dimensions in which the data is meant to be described. However, the provided matrix
based on diferent levels of literality/granularity of the contents of documents/collections and
their provenance, might be useful to formulate the requirements to make a more structured and
fundamental choice of suitable vocabularies in the future. This implies not only understanding
the features of the data at hand but also the requirements of data providers and users. When
this is clear, then one or more suitable vocabularies can be chosen as the ones that provide the
necessary expressivity.</p>
      <p>
        The idea was developed during the Golden Agents project where data from the Amsterdam
City Archives were analyzed in order to be published as RDF data with enrichments. It was
influenced by the Resource-Observation-Reconstruction structure presenRtOiAnRthveocabulary6
and later further developed as ROAR3+,+2[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Some interesting things became clear in this
process: (1) the archives’ role as data provider was not to provide enrichments that would add
much interpretation to the original data, while (2) the project required to enrich data as much
as possible, for example integrating mentions to create evidence-based storylines of events for
Amsterdamers in the 17th centur24y][and ultimately to get insights in the production and
consumption of cultural goods of the Dutch Golden Age. Therefore, the data produced under
the authority of the archive would go until a certain level, while the project had to bring it
further. Naturally, the vocabulary choices, reflecting diferent requirements, were not the same.
      </p>
      <p>During the Golden Agents project it was not possible to perform a thorough analysis of digital
humanities vocabularies and how they would correspond with in the proposed dimensions.
This would be an interesting area for future work, although often one vocabulary will not
completely cover one level, nor be contained within one. In particular, correspondence of these
dimensions with ontologies of provenance and historical assertion records,PsRuOchV-aOs7
and theSTAR model developed in the Releven proje8catnd specially Records in ContexRtI C()9,
a promising model and ontology for describing archival record resources and their contextual
entities and also the reference model from ISO called Open Archival Information Systems (OAIS
6https://w3id.org/roar
7www.w3.org/TR/prov-o/
8releven.univie.ac.at/
9www.ica.org/resource/records-in-contexts-ontology/
ISO 14721)10. Another important investigation would be an ontological analysis of the general
entities present in each level/dimension to support their correspondence with vocabularies.
The outcomes of such an analysis would make finding existing ontologies and vocabularies for
reuse much easier.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>CNPq 443130/2023-0 &amp; FAPES (1022/2022)</p>
      <p>Web–ISWC 2019: 18th International Semantic Web Conference, Auckland, New Zealand,
October 26–30, 2019, Proceedings, Part II 18, Springer, 2019, pp. 36–52.
[11] V. A. Carriero, P. Groth, V. Presutti, Empirical ontology design patterns and shapes from
wikidata, Semantic Web (2023).
[12] V. Presutti, E. Daga, A. Gangemi, E. Blomqvist, Extreme design with content ontology
design patterns, in: Proc. Workshop on Ontology Patterns, 2009, pp. 83–97.
[13] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak,
N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al., The fair guiding
principles for scientific data management and stewardship, Scientific data 3 (2016) 1–9.
[14] P. Groth, A. Gibson, J. Velterop, The anatomy of a nanopublication, Information services
&amp; use 30 (2010) 51–56.
[15] T. Kuhn, A. Meroño-Peñuela, A. Malic, J. H. Poelen, A. H. Hurlbert, E. C. Ortiz, L. I. Furlong,
N. Queralt-Rosinach, C. Chichester, J. M. Banda, et al., Nanopublications: a growing
resource of provenance-centric scientific linked data, in: 2018 IEEE 14th International
Conference on e-Science (e-Science), IEEE, 2018, pp. 83–92.
[16] P. Missier, K. Belhajjame, J. Cheney, The w3c prov family of specifications for modelling
provenance metadata, in: Proceedings of the 16th international conference on extending
database technology, 2013, pp. 773–776.
[17] A. Meroño-Peñuela, A. Ashkpour, C. Guéret, S. Schlobach, Cedar: the dutch historical
censuses as linked open data, Semantic Web 8 (2017) 297–310.
[18] R. Cyganiak, D. Reynolds, J. Tennison, The rdf data cube vocabulary, W3C recommendation
16 (2014).
[19] M. Doerr, The cidoc conceptual reference module: an ontological approach to semantic
interoperability of metadata, AI magazine 24 (2003) 75–75.
[20] B. Tillett, What is frbr? a conceptual model for the bibliographic universe, The Australian</p>
      <p>Library Journal 54 (2005) 24–30.
[21] R. Pearce-Moses, L. A. Baty, A Glossary of Archival and Records Terminology, volume
2013, Society of American Archivists Chicago, IL, 2005.
[22] P. Ciccarese, R. Sanderson, B. Young, Web Annotation Vocabulary, W3C Recommendation,</p>
      <p>W3C, 2017. URL: https://www.w3.org/TR/2017/REC-annotation-vocab-20170.223/
[23] M. Pasin, J. Bradley, Factoid-based prosopography and computer ontologies: Towards
an integrated approach, Digital Scholarship in the Humanities 30 (2015) 86–97. URL:
https://www.kcl.ac.uk/factoid-prosopography/overall-con.cepts
[24] A. Idrissou, V. Zamborlini, F. Van Harmelen, C. Latronico, Contextual Entity
Disambiguation in Domains with Weak Identity Criteria: Disambiguating Golden Age Amsterdamers,
in: K-CAP 2019 - Proceedings of the 10th International Conference on Knowledge Capture,
ACM, New York, NY, USA, 2019, pp. 259–262. URL:https://dl.acm.org/doi/10.1145/3360901.
3364440. doi:10.1145/3360901.3364440.
[25] L. van Wissen, V. Zamborlini, C. van den Heuvel, Toward an ontology for archival resources.
modelling persons, objects and places in the golden agents research infrastructure, 2021.
URL: https://dhistory.hypotheses.org,/d36a1ta for History Lecture 2021: Modelling Time,
Places, Agents, Online.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Comission</surname>
          </string-name>
          ,
          <source>Study on Standard-Based Archival Data Management, Exchange and Publication Final Report, Technical Report, Historical Archives Service</source>
          ,
          <year>2018</year>
          . URL: https://ec.europa.eu/isa2/sites/isa/files/isa2_action_
          <year>2017</year>
          _
          <article-title>01_standard_based_archival_ data_management_</article-title>
          <source>final_report_v1</source>
          .
          <fpage>00</fpage>
          ..pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>C. van den Heuvel</surname>
          </string-name>
          , V. Zamborlini,
          <article-title>Modeling and Visualizing Storylines of Historical Interactions. Kubler's Shape of Time and Rembrandt's Night Watch</article-title>
          , in: R. P.
          <string-name>
            <surname>Smiraglia</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Scharnhorst (Eds.),
          <article-title>Linking Knowledge: Linked Open Data for Knowledge Organization</article-title>
          and Visualization, 1 ed.,
          <source>Ergon-Verlag, Baden-Baden</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>141</lpage>
          .
          <year>1d0o</year>
          .
          <year>i5</year>
          :
          <volume>771</volume>
          /
          <fpage>9783956506611</fpage>
          -
          <lpage>99</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>L. van Wissen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zamborlini</surname>
          </string-name>
          , C. van den Heuvel,
          <article-title>Modeling provenance and uncertainties in the use of archival sources of the dutch golden age</article-title>
          ,
          <year>2022</year>
          . hUtRtLp:s://rsa.confex. com/rsa/2022/meetingapp.cgi/Paper/132,0p1resented at the 68th
          <source>Annual Meeting of the Renaissance Society of America (RSA)</source>
          ,
          <volume>30</volume>
          March - 2
          <source>April</source>
          <year>2022</year>
          , Dublin, Ireland.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Kotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Vouros</surname>
          </string-name>
          ,
          <article-title>Human-centered Ontology Engineering: The HCOME Methodology</article-title>
          ,
          <source>Knowledge and Information Systems</source>
          <volume>10</volume>
          (
          <year>2006</year>
          )
          <fpage>109</fpage>
          -
          <lpage>131</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K. I.</given-names>
            <surname>Kotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Vouros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spiliotopoulos</surname>
          </string-name>
          ,
          <article-title>Ontology engineering methodologies for the evolution of living and reused ontologies: Status, trends, findings and recommendations</article-title>
          ,
          <source>The Knowledge Engineering Review</source>
          <volume>35</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. d.</given-names>
            <surname>Moor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Leenheer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meersman</surname>
          </string-name>
          ,
          <article-title>DOGMA-MESS: A Meaning Evolution Support System for Interorganizational Ontology Engineering</article-title>
          , in: International Conference on Conceptual Structures, Springer,
          <year>2006</year>
          , pp.
          <fpage>189</fpage>
          -
          <lpage>202</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Suárez-Figueroa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gómez-Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fernández-López</surname>
          </string-name>
          ,
          <article-title>The NeOn Methodology for Ontology Engineering</article-title>
          , in: Ontology engineering in a networked world, Springer,
          <year>2012</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Vrandečić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tempich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sure</surname>
          </string-name>
          ,
          <article-title>The DILIGENT Knowledge Processes</article-title>
          ,
          <source>Journal of Knowledge Management</source>
          <volume>9</volume>
          (
          <year>2005</year>
          )
          <fpage>85</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>V. A.</given-names>
            <surname>Carriero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Daquino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gangemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Nuzzolese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Peroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Presutti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Tomasi</surname>
          </string-name>
          ,
          <source>The Landscape of Ontology Reuse Approaches</source>
          , IOS Press,
          <year>2020</year>
          . UhRtLt:p://dx.doi.org/10. 3233/SSW200033. doi:
          <volume>10</volume>
          .3233/ssw200033.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>V. A.</given-names>
            <surname>Carriero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gangemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Mancinelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Marinucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Nuzzolese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Presutti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Veninata</surname>
          </string-name>
          ,
          <article-title>Arco: The italian cultural heritage knowledge graph</article-title>
          , in: The Semantic
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>