<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>German National Socialist Injustice on the Semantic Web: from Archival Records to a Knowledge Graph</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mahsa Vafaie</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Applied Informatics and Formal Description Methods (AIFB), Karlsruhe Institute of Technology (KIT)</institution>
          ,
          <addr-line>Kaiserstraße 89, 76133 Karlsruhe</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>FIZ Karlsruhe - Leibniz Institute for Information Infrastructure</institution>
          ,
          <addr-line>Hermann-von-Helmholtz-Platz 1, 76344 Eggenstein-Leopoldshafen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Archival repositories contain vast amounts of historical data within unstructured textual documents, posing significant challenges for extracting coherent insights. This paper presents ongoing work towards an optimised workflow for constructing a knowledge graph from millions of archival records related to the “Wiedergutmachung” process in Germany. These records, documenting compensation and restitution eforts following World War II, ofer insights into the aftermath of the National Socialist regime. The proposed workflow involves converting document images to machine-readable formats, ontology design, information extraction, and entity linking. Leveraging both traditional methods and transformer-based technologies, the workflow addresses unique challenges inherent in historical documents.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Semantic Web</kwd>
        <kwd>Digital Cultural Heritage</kwd>
        <kwd>Digital Humanities</kwd>
        <kwd>Linked Open Data</kwd>
        <kwd>Optical Character Recognition</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>Wiedergutmachung</kwd>
        <kwd>Compensation for National Socialist Injustice</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Since the 1990s, researchers from various domains have been increasingly engaged with
providing access to information held within archival records, through online channels [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Archival
repositories hold invaluable historical data, often in the form of unstructured textual documents
that span decades. Extracting coherent insights from these records presents a formidable
challenge due to the lack of standardised formats and the sheer volume of the data. Digitalisation
pipelines emerge as a computational solution to this challenge, leveraging a combination of
computer vision, natural language processing, information extraction, machine learning, and
semantic analysis techniques. These pipelines, through sophisticated algorithms and
methodologies, facilitate the transformation of archival documents into structured data points. The true
power of digitalisation manifests in the construction of knowledge graphs -— a dynamic
framework that transforms discrete data points into interconnected nodes. Knowledge graphs play a
pivotal role in bridging the gap between archival records and the Semantic Web. By interlinking
the extracted information from archival records on the Semantic Web, knowledge graphs enable
researchers to discern intricate relationships, uncover latent patterns, and traverse historical
narratives that go beyond individual documents [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Furthermore, the existence of knowledge
graphs as the backbones of information systems allows for inference of previously undiscovered
knowledge from the given statements, and provides means for conducting exploratory search
and knowledge discovery. On the other hand, enhancing accessibility of archival data leads to
increased public participation and a better understanding of archival materials [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>This paper proposes an optimised workflow for constructing a knowledge graph from millions
of archival records from the “Wiedergutmachung” 1 process in Germany, for utilisation in
the semantic portal “Themenportal Wiedergutmachung” 2. The term “Wiedergutmachung
compensation records” in this paper, refers to collections of documents, records, and materials
related to the process of compensation and restitution eforts that followed World War II
and the fall of the National Socialist regime in Germany. Wiedergutmachung compensation
records which contain an estimated amount of 100 km of archival documents, originate from
the State Ofices for Compensation ( Ämter für Wiedergutmachung in German) installed by the
German government in every German Land after the war. These collections contain a wide
range of documents, including index cards, application forms, documents on legal proceedings,
correspondence, testimonies, and other materials that pertain to individuals, families, and
communities seeking compensation for forced labour, imprisonment, injuries, and other damages
caused by National Socialist Injustice.</p>
      <p>Wiedergutmachung compensation records serve as a historical account of the processes that
took place to acknowledge and address the immense human sufering caused by the Nazi regime.
They play a crucial role in documenting the complex journey of survivors and their families in
seeking justice, recognition, and support for the harm they endured. They also provide insights
into the legal, bureaucratic, and social challenges faced by those seeking compensation in the
aftermath of such a devastating period in history. The Wiedergutmachung knowledge graph
(Wiedergutmachung KG) contributes to our understanding of the impact of the totalitarian
government in Germany on individuals and society and the ongoing eforts to address historical
injustices. Construction of the Wiedergutmachung KG, illuminates the historical information
hidden within unstructured Wiedergutmachung records, which were formerly tucked away in
archival repositories, only accessible to archivists and a limited number of individuals entitled
to see them for family or scientific research purposes. For instance, Wiedergutmachung KG can
display connections between “claimants” and “compensation decisions”, elucidating trends and
disparities within the Wiedergutmachung process.</p>
      <p>The proposed digitalisation workflow starts with conversion of document images (i.e., scanned
documents) to machine-readable formats through Optical Character Recognition (OCR).
Subsequently, the workflow entails ontology design, information extraction, and linking entities with
external data sources, to construct the Wiedergutmachung KG. Due to the historical nature of
the documents, each of these stages present unique challenges that can be tackled using the more
traditional methods, while the advancements in transformer-based technologies can also be
used to address them with a more direct approach. The speed of technological advancements in
the field of AI calls for a dynamic pipeline with a modular design that allows for substitution or
1“Wiedergutmachung” is a German word that translates to “making good again” or “making amends”. In the context
of National Socialism in Germany and its aftermath, it specifically refers to the eforts made to compensate survivors
and victims for the losses they sufered during the rule of the Nazi regime.
2https://www.archivportal-d.de/themenportale/wiedergutmachung
combination of rule-based methods for accuracy with nascent AI technologies for optimisation.
Keeping this consideration in mind, the same digitalisation workflow can be extended to further
use cases and Digital Humanities research can benefit from the lessons learnt during the design
of such a pipeline.</p>
      <p>The remainder of this paper delves into details of the diferent components of the proposed
digitalisation pipeline for Wiedergutmachung compensation records and discusses the intricacies
of working with historical archival records. Section 2 outlines similar eforts for construction of
domain-specific KGs in the context of Cultural Heritage. In Section 3 the research questions are
introduced and methodologies for addressing them are discussed. Section 4 concludes the paper
and sketches the next steps for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        As archives undergo mass digitisation and the volume of digital records grows, there arises a
rich but underutilised resource for researchers in the Digital Humanities. Integration of data
from historical archival records into the Semantic Web and transformation of archival data
according to Linked Open Data (LOD) principles has been receiving significant attention by
scholars [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The Sampo series of semantic portals is a pioneering efort in applying Semantic Web
technologies to showcase Finland’s national heritage [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. These portals utilise the modular FinnONTO,
as a taxonomy of Cultural Heritage Objects [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. WarSampo for example focuses on harmonising
and publishing heterogenous datasets related to World War II in Finland as LOD [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        The Jewish Contemporary Documentation Center has created an online LOD database 3
focused on Italian Holocaust victims and persecution events. Additionally, they’ve developed
an associated application for utilising this valuable data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>In the Netherlands, “Oorlog voor de Rechter” (“War in Court”) 4 aims to unlock historical
knowledge by making The Central Archives of the Special Jurisdiction (CABR) accessible online,
leveraging advanced technologies and a user-centric design. CABR is the largest war archive
in the Netherlands, containing files of over 400,000 people suspected of collaboration with the
National Socialist regime in Germany.</p>
      <p>
        The European Holocaust Research Infrastructure (EHRI) Portal 5 serves as a valuable resource
for researchers and historians interested in Holocaust-related archival material. It provides
access to electronic finding aids, inventory information on institutions holding
Holocaustrelated records, and vocabularies related to archival descriptions [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Researchers can use these
vocabularies to improve searchability and interoperability.
      </p>
      <p>In Germany, to the best of our knowledge, this work marks the first efort to develop an
LOD-based semantic portal from archival records pertaining to World War II and National
Socialist Injustices.</p>
      <sec id="sec-2-1">
        <title>3http://dati.cdec.it/lod/shoah/website/html 4https://www.huygens.knaw.nl/en/projecten/war-in-court/ 5https://portal.ehri-project.eu/</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. An LODification workflow for Wiedergutmachung compensation records</title>
      <p>Transformation of data hidden within Wiedergutmachung compensation records into LOD
for increased accessibility, interoperability, explorability, and semantic enrichment is the main
goal of this work. Therefore, the overarching research question in this work is: What is
the most eficient pipeline for transformation of historical archival records into a
Knowledge Graph-based information system, for integration into the Semantic Web
and publication as Linked Open data? In order to accurately address the overarching
research question, it seems necessary to break it down into the diferent components of such
a pipeline, for an informed design decision. The research questions derived from a modular
design for this pipeline are as follows:
RQ1: What advanced techniques and methodologies can be developed to improve the accuracy
of text recognition for digitised archival records with challenging characteristics such as
faded ink, handwritten annotations, and non-standard fonts?
RQ2: What are the most efective Information Extraction (IE) techniques for accurately and
eficiently identifying and retrieving structured data, such as names, dates, and locations,
from digitised archival records?
RQ3: How can existing ontologies be adapted and extended to develop an ontology for
representation of archival historical records, that accurately reflects the hierarchical structure
of archival records, semantic annotation of the records, and information about the agents
involved in the creation and archiving of the records
RQ4: How can we establish reliable links between historical entities (e.g., people, places, and
events) extracted from digitised archival records and relevant external databases, authority
ifles, or reference materials?
The methodologies and initial experiments for addressing RQ1, RQ2, and RQ3 are explained
below. Solutions for RQ4 are yet to be explored as a part of future work.</p>
      <sec id="sec-3-1">
        <title>3.1. RQ1: Ontology Development</title>
        <p>The development of Wiedergutmachung KG hinges on the creation of an ontology capable of
modelling relationships among archival documents, court proceedings, individuals, and
organisations involved in the compensation application, decision-making, and document archiving.
Ensuring the validity and reliability of this model requires incorporation of domain experts’
requirements and knowledge. Archivists from the State Archives of Baden-Württemberg 6
collaborated with the author to formulate a list of competency questions 7, serving as the
foundation for the ontology’s conceptual modeling. These questions, catering to
researchers/historians and relatives/dependents of persecuted persons, provide insights into the domain’s
scope, structure, and concepts.</p>
        <sec id="sec-3-1-1">
          <title>6https://www.landesarchiv-bw.de/ 7The full list of competency questions is published on GitHub, in the Wiedergutmachung repository.</title>
          <p>
            According to the competency questions and based on the best practices in the field of Ontology
Design, the CourtDocs Ontology [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] is created to consist of three main building blocks, reusing
existing ontologies in order to avoid redundancy and thereby to also enable interoperability
with external data sources. Each of these building blocks and the ontologies that have been
reused for their creation are described below.
          </p>
          <p>
            Archival Hierarchy and Provenance. The Records in Context ontology (RiC-O8) [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] is
employed to model the hierarchical structure of Wiedergutmachung compensation records, due
to its inclusion of named individuals, which makes it adaptable across institutes with diferent
archival systems and practices. On the other hand, RiC-O’s incorporation of smaller entities
improves findability and enables a more detailed representation of archival resources, including
constituent parts like stamps [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ].
          </p>
          <p>
            Court Procedures. The PROV Ontology (PROV-O9) [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] is reused to depict the
Wiedergutmachung process within the court system. It ofers a standardised approach for modelling how
entities and activities evolve over time, making it efective for process modelling. Additionally,
PROV-O is widely recognised for representing provenance information, making it suitable for
capturing the relationships between Wiedergutmachung procedures and the records generated
or utilised at each stage of the process.
          </p>
          <p>
            Biographical Information of Persons Involved. CIDOC-CRM is employed as the
ontological foundation for representing the biographical information of individuals involved in the
compensation process 10. The use of a harmonising data model facilitates connection with other
materials and external sources. Moreover, the event-centric approach of CIDOC-CRM, in which
an individual’s existence is perceived as a series of interconnected events spanning across time
and space [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ], enables the representation of crucial life events in prosopographical research on
victims of National Socialism, including events like deportation and imprisonment.
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. RQ2: OCR Quality Enhancement</title>
        <p>
          Creation of transcripts from scanned documents using OCR systems, greatly accelerates and
streamlines the retrieval of information. However, OCR system eficacy is contingent upon
factors such as text and font styles. OCR systems are typically specialised for either machine-printed
or handwritten text due to their distinct visual characteristics [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Yet, archival documents
often feature mixed text. Traditionally, workflows dealing with a variety of text types on
scanned documents employed distinct text recognition models. In [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and [17] we propose
a pipeline for separation of machine-printed text and handwritten text on historical archival
documents that contain both text types. This OCR pre-processing step helps improve the quality
of the transcripts, by breaking down each document image into two layers, each containing
a particular text type, namely, handwritten text, or machine-printed text, and consequently,
feeding the layers into the appropriate OCR or Handwritten Text Recognition (HTR) engines.
In our preliminary work, we achieved an increase of 16% compared to the baseline systems for
separation of text types, with models trained on modern documents. In a more recent
development, the new Transformer-based OCR (TrOCR) models have demonstrated the capability to
8https://www.ica.org/standards/RiC/ontology
9https://www.w3.org/TR/prov-o/
10https://cidoc-crm.org/
adapt to variations in fonts, text types, styles, and languages [18, 19], skipping the pre-requisite
steps of dataset synthesis and model training for text type separation. With word accuracy as
an OCR evaluation metric ranging between 75% and 85%, TrOCR engines from Transkribus 11
have the potential to optimise the OCR quality improvement step. A qualitative evaluation of
the text-type separated transcripts is yet to be done, for comparison with TrOCR results.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. RQ3: Information Extraction</title>
        <p>Wiedergutmachung compensation records constitute of diferent document types, such as
application forms and index cards, and diferent layouts for each document type that vary based
on time and across diferent institutes. In a traditional information extraction pipeline, this
necessitates implementation of an automatic document identification system that classifies
the documents based on their type or layout, and feeds them into the respective information
extraction script customised for each document type. Our experiments on rule-based information
extraction using Apache Uima Ruta [20] with a subset of 75 documents with three diferent
layouts show an accuracy of 75%, for exact matches only. However, with the advance of Large
Language Models (LLMs), there is an opportunity to streamline the information extraction
process, instead of laboriously crafting multiple scripts for each document and layout type. In
the proposed LLM-based approach, a unified prompt, coupled with the appropriate context, can
facilitate the extraction of information from all the documents spanning across document types
and layouts. The quality of information extraction from archival records with LLMs is still to be
evaluated for a more accurate analysis and comparison with the rule-based methods.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion and Future Work</title>
      <p>The presented research provides a significant contribution to Digital Humanities research,
particularly on topics related to World War II and National Socialist Injustice. Furthermore, the
ifndings from implementations of diferent techniques on archival records can be extended to
similar eforts on transformation of archival records into LOD.</p>
      <p>In the next stages of the research, the focus will shift towards implementation of
transformerbased technologies in the LODification pipeline and comparing the performance of these
methods against the more traditional methods for each of the constituent pipeline components.
Moreover, RQ4 from Section 3 will be addressed to interconnect the KG with external sources
and authority files. This is a crucial step to facilitate content-based and federated semantic
search and to enrich the KG. In case of the Wiedergutmachung KG, apart from Wikidata 12,
there are several other knowledge- and databases, representing data on German figures and the
victims of National Socialism that can be interlinked with the KG. It is also necessary to map all
the extracted information to specific unique entities (e.g., persons) by means of disambiguation
and entity resolution techniques.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work is funded by the German Federal Ministry of Finance (Bundesministerium der Finanzen)
11https://www.transkribus.org/de
12https://www.wikidata.org/
and supervised by Prof. Dr. Harald Sack.
cation in historical archival documents, in: Archiving Conference, volume 19, Society for
Imaging Science and Technology, 2022, pp. 15–20.
[17] M. Vafaie, J. Waitelonis, H. Sack, Improvements in Handwritten and Printed Text Separation
in Historical Archival Documents, in: Archiving Conference, volume 20, Society for
Imaging Science and Technology, 2023, pp. 36–41.
[18] M. Li, T. Lv, et al., Trocr: Transformer-based optical character recognition with
pretrained models, in: Proc. of the AAAI Conf. on Artificial Intelligence, volume 37, 2023, pp.
13094–13102.
[19] P. B. Ströbel, T. Hodel, W. Boente, M. Volk, The Adaptability of a Transformer-Based OCR
Model for Historical Documents, in: Intl. Conf. on Document Analysis and Recognition,
Springer, 2023, pp. 34–48.
[20] P. Kluegl, M. Toepfer, P.-D. Beck, G. Fette, F. Puppe, Uima ruta: Rapid development of
rule-based information extraction applications, Natural Language Engineering 22 (2016)
1–40.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>W.</given-names>
            <surname>Duf</surname>
          </string-name>
          , Archival mediation,
          <source>Currents of archival thinking</source>
          (
          <year>2010</year>
          )
          <fpage>115</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Waitelonis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          ,
          <article-title>Towards exploratory video search using linked data</article-title>
          ,
          <source>Multimedia Tools and Applications</source>
          <volume>59</volume>
          (
          <year>2012</year>
          )
          <fpage>645</fpage>
          -
          <lpage>672</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Oomen</surname>
          </string-name>
          , M. van
          <string-name>
            <surname>Erp</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Baltussen</surname>
          </string-name>
          ,
          <article-title>Sharing cultural heritage the linked open data way: why you should sign up</article-title>
          ,
          <source>in: Museums and the Web</source>
          <year>2012</year>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hawkins</surname>
          </string-name>
          ,
          <article-title>Archives, linked data and the digital humanities: increasing access to digitised and born-digital archives via the semantic web</article-title>
          ,
          <source>Archival Science</source>
          <volume>22</volume>
          (
          <year>2022</year>
          )
          <fpage>319</fpage>
          -
          <lpage>344</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <article-title>Digital humanities on the semantic web: Sampo model</article-title>
          and portal series,
          <source>Semantic Web</source>
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>729</fpage>
          -
          <lpage>744</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Viljanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Seppälä</surname>
          </string-name>
          ,
          <article-title>Building a national semantic web ontology and ontology service infrastructure-the finnonto approach</article-title>
          ,
          <source>in: The Semantic Web: Research and Applications: 5th European Semantic Web Conference, ESWC</source>
          <year>2008</year>
          , Tenerife, Canary Islands, Spain, June 1-5,
          <source>2008 Proceedings 5</source>
          , Springer,
          <year>2008</year>
          , pp.
          <fpage>95</fpage>
          -
          <lpage>109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Koho</surname>
          </string-name>
          , E. Ikkala,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leskinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          , E. Hyvönen,
          <article-title>Warsampo knowledge graph: Finland in the second world war as linked open data</article-title>
          ,
          <source>Semantic Web - Interoperability, Usability, Applicability</source>
          <volume>12</volume>
          (
          <year>2021</year>
          )
          <fpage>265</fpage>
          -
          <lpage>278</lpage>
          . URL: https://doi.org/10.3233/SW-200392. doi:
          <volume>10</volume>
          .3233/SW-200392.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          , G. Moretti,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tonelli</surname>
          </string-name>
          , et al.,
          <article-title>Lod navigator: tracing movements of italian shoah victims</article-title>
          , Umanistica
          <string-name>
            <surname>Digitale</surname>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>N-A.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Blanke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bryant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Frankl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kristel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Speck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Daelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Horik</surname>
          </string-name>
          ,
          <article-title>The european holocaust research infrastructure portal</article-title>
          ,
          <source>Journal on Computing and Cultural Heritage (JOCCH) 10</source>
          (
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vafaie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bruns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pilz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Waitelonis</surname>
          </string-name>
          , H. Sack, CourtDocs Ontology:
          <article-title>Towards a Data Model for Representation of Historical Court Proceedings</article-title>
          ,
          <source>in: Proc. of the 12th Knowledge Capture Conference</source>
          <year>2023</year>
          ,
          <year>2023</year>
          , pp.
          <fpage>175</fpage>
          -
          <lpage>179</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Clavaud</surname>
          </string-name>
          , T. Wildi,
          <article-title>ICA records in contexts-ontology (</article-title>
          <string-name>
            <surname>RiC-O):</surname>
          </string-name>
          <article-title>a semantic framework for describing archival resources</article-title>
          ,
          <source>in: Proc. of Linked Archives Int. Workshop</source>
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vafaie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bruns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pilz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dessí</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          ,
          <article-title>Modelling Archival Hierarchies in Practice: Key Aspects and Lessons Learned</article-title>
          ,
          <source>in: 6th Intl. Workshop on Computational History (HistoInformatics</source>
          <year>2021</year>
          ),
          <article-title>Online event</article-title>
          ,
          <source>September 30-October 1</source>
          ,
          <year>2021</year>
          , volume
          <volume>2981</volume>
          , Aachen, Germany: RWTH Aachen,
          <year>2021</year>
          , p.
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Lebo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sahoo</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>PROV-O: The</surname>
            <given-names>PROV</given-names>
          </string-name>
          ontology,
          <source>W3C recommendation 30</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leskinen</surname>
          </string-name>
          ,
          <string-name>
            <surname>Bio</surname>
            <given-names>CRM</given-names>
          </string-name>
          :
          <article-title>A data model for representing biographical data for prosopographical research</article-title>
          ,
          <source>in: Proc. of the 2nd Conf. on Biographical Data in a Digital World</source>
          <year>2017</year>
          (
          <issue>BD2017</issue>
          ),
          <source>CEUR Workshop Proceedings</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Islam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Islam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Noor</surname>
          </string-name>
          ,
          <article-title>A survey on optical character recognition system</article-title>
          ,
          <source>arXiv preprint arXiv:1710.05703</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vafaie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bruns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pilz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Waitelonis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          , Handwritten and printed text identifi-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>