<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Refinement for Cultural Heritage Digital Libraries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mary Ann Tan</string-name>
          <email>ann.tan@fiz-karlsruhe.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Baltimore, Maryland, USA</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Applied Informatics and Formal Description Methods (AIFB), Karlsruhe Institute of Technology</institution>
          ,
          <addr-line>KIT</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>FIZ Karlsruhe - Leibniz Institute for Information Infrastructure</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Hermann-von-Helmholtz-Platz 1</institution>
          ,
          <addr-line>76344 Eggenstein-Leopoldshafen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Kaiserstraße 89</institution>
          ,
          <addr-line>76133 Karlsruhe</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Semantic Web, NLP, Information Extraction</institution>
          ,
          <addr-line>Knowledge Graphs, Digital Libraries, Cultural Heritage</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Digital Libraries containing metadata of diverse cultural heritage objects are meant to be accessible not only to domain experts but also to the general population. This calls for information services that can provide ease and eficiency to search, retrieval and exploration. Knowledge graphs (KGs) are essential for representation, organization, integration, and analysis of hierarchical and heterogeneous information. However, most KGs sufer from incompleteness and inaccuracies. This work intends to address various challenges arising from construction and refinement of a KG populated with historical objects, by defining domain- and application-appropriate ontologies and leveraging approaches in information extraction (IE) for improving metadata quality.</p>
      </abstract>
      <kwd-group>
        <kwd>1Deutsche Digitale Bibliothek</kwd>
        <kwd>https</kwd>
        <kwd>//www</kwd>
        <kwd>deutsche-digitale-bibliothek</kwd>
        <kwd>de</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        described by Tan et al.[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In addition, the metadata collection sufers from incompleteness and
inaccuracies as described in Tan et al[.3]. This prevents the underlying retrieval engine from
properly indexing the objects.
      </p>
      <p>To address these challenges necessitates a combination of solutions in knowledge
representation, knowledge refinement, and information extraction. Therefore, this thesis proposes i)
an ontology that enables interoperability across diferent types of CHOs while maintaining
domain-specific semantics as discussed in Section 5.1; ii) a KG refinement approach leveraging
NLP teachniques to improve metadata quality of historical objects; and iii) an Entity Linking
approach for entities in historical objects.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Importance</title>
      <p>This work will benefit not only the general population, but also the domain experts such
as librarians, curators, and archivists. Proposed solutions will empower users from diverse
backgrounds to seamlessly and eficiently search, retrieve, and explore Germany’s rich and
voluminous collection.</p>
      <p>Recent developments in AI can be leveraged to address the technical challenges facing the
DDB. This work is relevant to the researchers working at the intersection of Semantic Web
(SW), Digital Humanities (DH), and Natural Language Processing (NLP).</p>
    </sec>
    <sec id="sec-4">
      <title>3. Related Work</title>
      <p>
        There have been several notable data models or ontologies proposed for cultural heritage
representation. Liu et al.[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] provided a review of CIDOC-CRM, Sampo Model, and EDM
specific to the museum use case only. Cultural heritage data models are delineated along
two modeling paradigms: object-centric and event-centric. CIDOC-CRM follows the former,
while EDM follows a mixture of both paradigms. Object-centric modeling defines attributes
directly by describing the object, while event-centric modeling defines these attributes through
a series of events associated with the object. Object-centric modeling favors conciseness, while
event-centric modeling emphasizes completeness.
      </p>
      <p>
        A pioneer in the application of of SW technologies, theSampo series of semantic portals
showcase the national heritage of Finland. These systems make use of the modular FinnONTO
ontology infrastructure5[]. However, FinnONTO is not a full-featured ontology, but a taxonomy
of CHOs encoded as Simple Knowledge Organization System (SKOS) concepts. Following the
modular modeling approach is Italy’s ArC3O[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], where each module is intended to describe a
CHO4 in the context of cataloging activities and events.
      </p>
      <p>The core design principles of EDM, and by extension DDB-EDM, lead to definitions of general
classes that require the bare minimum of metadata properties and controlled vocabularies.
Thus, all CHOs, regardless of their sources, media types and object types, are instances of the
classedm:ProvidedCHO, while their digitized representations on the Web are instances of the
classedm:WebResource. This flexibility however results in imprecise representations and loss</p>
      <sec id="sec-4-1">
        <title>3ArCO, https://w3id.org/arco 4CHO is referred to as “Cultural Property”.</title>
        <p>of semantics inherent in the original objects 7[]. In particular, it is not possible to model the
concepts and level of abstractions widely-accepted in the bibliographical domain.</p>
        <p>The International Federation of Library Associations and Institutions (IFLA) developed the
Functional Requirements for Bibliographic Records (FRBR8)][, where a book can be represented
as several entities and the relationships that exist among these entities. A copy of a book
(frbr:Item) is a specimen or exemplification of a specific publication ( frbr:Manifestation),
which is an embodiment of an expression frbr:Expression that realizes the ideas of a creative
work (frbr:Work).</p>
        <p>
          Most Europeana users are less likely to search for specific items (11.3%) and are more inclined
to search by category (47.1%) and by subject (24.6%) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. This supports the need to align
bibliographic objects from theItem level to their respective higher-level abstractionWs o(rk,
Expression, Manifestation). Consequently, the process of alignment sets a prerequisite for objects
to possess identifiable properties and attributes, such as title, agents, dates, and subject heading.
However, due to the age of the objects, a high level of uncertainty with respect to proper author
or date attributions is apparent.
        </p>
        <p>The challenges of filling missing information and identifying erroneous information in
a knowledge graph fall under the umbrella of Knowledge Graph Refinement. In particular,
Knowledge Graph Completion (KGC) deals with the former challenge, whilEerror Detection deals
with the latter.</p>
        <p>By definition, internal methods for KGC use the content of the current KG either to determine
class membership or to predict relations between entities. These methods require the current
KG to at least possess reasonable quality in order for large scale evaluation to be feasib10le].[
On the other hand, external methods leverage other sources of knowledge for refinement, such
as other knowledge graphs or text corpora.</p>
        <p>With the rapid development in the area of Natural Language Processing (NLP), text corpora
have become an excellent source of external knowledge. The subfield of Information Extraction
(IE), an intermediate step to knowledge graph construction, can be defined as the process of
gleaning structured information from unstructured tex1t1[]. A concrete example of this task
would be to extract distinct properties and attributes identifying a literary work from the title.</p>
        <p>An IE pipeline starts with Named Entity Recognition (NER), or the detection and classification
of named entities mentioned in the text. Types of entities can be coarse-grained such asPERSON,
WORK_OF_ART, DATE, et cetera or fine-grained such as AUTHOR, PUBLISHER, ARTWORK, PUBLISHER,
LITERARY_WORK, PUBLICATION DATE, et cetera.</p>
        <p>
          Specific entity types (fine-grained) are often found in domain-specific texts, or even
timespecific texts where concept drift is quite common. In the field of digial humanities, there are
a number of studies on NER with historical text 1[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], however, coverage goes back to 17th
century B.C. at best (DROC [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]) and none belong to the domain of bibliography.
        </p>
        <p>
          Once the entities have been detected and classified, they are linked to specific entries in
reference knowledge bases or KGs. Entity Linking is particularly challenging due to the surface
form variations. In particular, names in historical texts can be multilingual, refer to aliases or
contain initials, include honorifics and designations. The names of geop-olitical entities are
also known to change through time. Pontes et al[.14] proposed an end-to-end multilingual
NER and EL (NERL) approach to address some of these challenges using some of the datasets
mentioned in Ehrmann et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Research Questions</title>
      <p>This section formulates the research questions (RQs) to address the challenges and limitations
of existing approaches described in Section3.</p>
      <p>RQ1: How can existing ontologies be adapted and extended to suit the domain and
application profile of digital libraries, such as the DDB?
Cultural heritage practitioners have been developing ontologies for specifc domains and
applications. As of this writing, only EDM is used to represent metadata from several
cultural institutions. In order to prevent data model silo1s5][ and to promote reusability,
it is beneficial to consider existing ontologies that are applicable and appropriate for the
use case of the DDB. Preliminary results are discussed in Sectio5n.1.</p>
      <p>RQ2: How can we leverage state-of-the-art NLP models to improve metadata quality
of historical objects?
Non-contemporary titles in the DDB &lt;(dc:title&gt;) encode details that can be used to
fillout missing properties, such as the title itself, author, publisher, editor, subject headings,
and dates. Hence, this calls for extractive NLP approaches. Section5.2 presents some
preliminary results.</p>
      <p>RQ2.1: How can we automatically construct an evaluation dataset from the DDB?
In order to address the succeeding RQs, an evaluation dataset for IE is required.</p>
      <p>Section 5.2 briefly describes what has been done so far.</p>
      <p>RQ2.2: How can we efectively extract fine-grained bibliographic entities from
historical texts?
The goal here is to address open challenges in the area of historical NER, such as how
to properly handle the dynamics of an evolving language, where spelling and naming
conventions change through time, and noise resulting from OCR engine. Dataset
construction, design of experiments, and model development will be accomplished.
RQ3: How can we link entities to records in the reference KG? The goal here is to
accurately disambiguate named entities and link them to entries in external KGs, while
addressing the challenges associated with historical texts. Moreover, entities that do not
exist in the reference KG can used as further contribution to increase the coverage of
authority files.</p>
      <p>The entirety of this work is envisioned to guide the construction and refinement of a
knowledge graph representing DDB’s cultural heritage objects. In addition, some open questions have
yet to be addressed concerning NERL in historical texts.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Preliminary Results</title>
      <p>The section describes preliminary work conducted to address the open questions presented in
Section 4.</p>
      <sec id="sec-6-1">
        <title>5.1. The DDB Ontology (DDB-O)</title>
        <p>
          Extensive quantitative and qualitative analysis of the entire DDB metadata collection have been
conducted in order to ascertain the applicability of existing CH ontologies. Initially, objects were
logically classified according to their originating institution, whether from libraries, archives,
museums, media libraries, or historical preservation. In addition, the media type of an object
was also taken into account. Taking up a large proportion of the entire collection, the alignment
of textual bibliographic resources to an extension of FRBR5 have been presented [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and
implemented as a SPARQL Endpoint [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Domain-specific ontologies have been adapted to
have more precise semantic representation objects (eg. components of bibliographic objects,
hierarchy of archival objects, level of representations of an image, etc.) Existing audio ontologies
intended for other domains have been extended to represent intangible audio heritage17[]. The
DDB-O Namespace6 is available online. A formal and complete specification is under review
and yet to be published.
        </p>
        <p>
          FRBR, as the upper ontology, requires that each object is looked up against a list of creative
works, such as the German Authority File oGremeinsame Normdatei (GND7). This ensures that
the relationship between diferent objects resulting from the same creative work is represented
in the KG [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Information Extraction</title>
        <p>
          The alignment of bibliographic items to their corresponding literary works proved to be a
challenging task due to incomplete object descriptions1[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Taking advantage of the greater
textual content encoded in the titles, several NLP tasks were reformulated in order to extract
contextual details present in the title. Several state-of-the-art, of-the-shelf NER and extractive
QA models, as well as LLMs were used in the experiments.
        </p>
        <p>
          As described in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], the objects in the evaluation dataset were selected according to language,
hierarchy type, existence of agent and date properties, format, and title length (&gt;30 tokens).
        </p>
        <p>
          A more forgiving evaluation measureP(recision@n) described in Section 6.2 was defined to
take into account the various naming conventions found in the text. An NER model (FLERT) 1[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
that can detect literary works and dates was initially used to test the hypothesis, and to refine
the evaluation dataset for the succeeding tasks. The results shown in Tabl1e illustrate that
these models can be leveraged but only to a lesser extent. The results were poor since the
models were not adapted to the age and domain of the texts. In addition, the results are not
indicative of the actual model performance due to evaluation dataset inaccuracie3s][.
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Evaluation</title>
      <p>The research questions enumerated in Section4 require diferent evaluation procedures, dataset,
and metrics. These are described in the succeeding subsections.
5Functional Requirements for Bibliographic Records8][
6DDB-O , https://ise-fizkarlsruhe.github.io/ddbkg/ddbo
7GND, https://www.dnb.de/DE/Professionell/Standardisierung/GND/gnd_node.html</p>
      <p>1–8</p>
      <sec id="sec-7-1">
        <title>6.1. Ontology Evaluation</title>
        <p>There several ways in evaluating ontologies. One of which is using is using competency
questions (CQs). A collection of CQs published in GitHu8bare included in the partial ontological
definitions and alignment activities. In addition, SPARQL query processing time for CQs that
can be answered with DDB-EDM will be compared with queries using the proposed ontology.</p>
      </sec>
      <sec id="sec-7-2">
        <title>6.2. Information Extraction</title>
        <p>
          Name matching for historical documents is non-trivial due to various naming conventions and
spelling variations. In a QA task, the most forgiving measure iAs ccuracy@1, which returns 1 if
there is a single token overlap between the ground truth and the answePrr.ecision@n measure is
a combination of 2 matching criteria: an exact match of the DDB object ID and an approximate
match for names using the Levenshtein edit distance [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>The evaluation measures for RQ3 will not be any diferent from those associated with EL.
A large proportion of the agents in the DDB are already linked to GND Persons. And there
already exist links between GND and Wikidata entities. This means that it is trivial to combine
naming variations and multilingual names for the evaluation dataset. Evaluating geopolitical
entities will require prior knowledge of the age of the object in question.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>7. Limitations and Future Work</title>
      <p>As discussed in Section 5.2, the lack of a gold standard evaluation dataset brings a level of
uncertainty to the experimental results. This will be addressed with the creation of a manually
annotated dataset with fine-grained entities. Consequently, this dataset will be used to address
RQ2.2. In addition, the work conducted to address RQ1 need to be finalized. Finally, entities
that already exist in GND will be linked, while non-existing ones can be used to further increase
the coverage of GND and Wikidata.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>I would like to thank my supervisors Prof. Dr. Harald Sack and Dr. Shufan Jiang for their
invaluable mentoring and support.</p>
      <sec id="sec-9-1">
        <title>8CQs for DDB-O, https://ise-fizkarlsruhe.github.io/ddbkg/docs/examples/</title>
        <p>1–8</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Purday</surname>
          </string-name>
          ,
          <article-title>Think culture: Europeana.eu from concept to construction</article-title>
          ,
          <source>Bibliothek Forschung und Praxis</source>
          <volume>33</volume>
          (
          <year>2009</year>
          )
          <fpage>170</fpage>
          -
          <lpage>180</lpage>
          . doi:
          <volume>10</volume>
          .1515/bfup.
          <year>2009</year>
          .
          <volume>018</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tietz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bruns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oppenlaender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dessì</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          ,
          <article-title>DDB-EDM to FaBiO: The Case of the German Digital Library</article-title>
          ,
          <source>in: Proc. of the 20th Int. Semantic Web Conference - Posters and Demos - ISWC</source>
          <year>2021</year>
          , volume
          <volume>2980</volume>
          , CEUR-WS.org,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          , Great Article,
          <source>in: Workshop on Deep Learning and Linguistic Linked Data</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hindmarch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hess</surname>
          </string-name>
          ,
          <article-title>A review of the cultural heritage linked open data ontologies and models, The International Archives of the Photogrammetry, Remote Sensing</article-title>
          and
          <string-name>
            <surname>Spatial Information Sciences XLVIII-M-</surname>
          </string-name>
          2-
          <fpage>2023</fpage>
          (
          <year>2023</year>
          )
          <fpage>943</fpage>
          -
          <lpage>950</lpage>
          . URL: https://isprs-archives.copernicus.org/articles/XLVIII-M-2-2023/943/202.3d/oi:10.5194/
          <article-title>isprs-archives-</article-title>
          <string-name>
            <surname>XLVIII-M-</surname>
          </string-name>
          2
          <string-name>
            <surname>-</surname>
          </string-name>
          2023-943-
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Viljanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Seppälä</surname>
          </string-name>
          ,
          <article-title>Building a national semantic web ontology and ontology service infrastructure -the finnonto approach</article-title>
          , in: S. Bechhofer,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hauswirth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          , M. Koubarakis (Eds.),
          <source>The Semantic Web: Research and Applications</source>
          , Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2008</year>
          , pp.
          <fpage>95</fpage>
          -
          <lpage>109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V. A.</given-names>
            <surname>Carriero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gangemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Mancinelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Marinucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Nuzzolese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Presutti</surname>
          </string-name>
          , C. Veninata,
          <article-title>ArCo: The Italian Cultural Heritage Knowledge Graph, The Semantic Web - ISWC</article-title>
          <year>2019</year>
          (
          <year>2019</year>
          )
          <fpage>36</fpage>
          --
          <lpage>52</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -30796-
          <issue>7</issue>
          _
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Peroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Tomasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Vitali</surname>
          </string-name>
          ,
          <article-title>Reflecting on the Europeana Data Model</article-title>
          ,
          <source>in: IRCDL</source>
          <year>2012</year>
          ,
          <year>2012</year>
          , pp.
          <fpage>228</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Tillet</surname>
          </string-name>
          , What is FRBR?
          <article-title>: A Conceptual Model for the Bibliographic Universe</article-title>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Paramita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goodale</surname>
          </string-name>
          , Europeana:
          <article-title>What users search for and why</article-title>
          , in: J.
          <string-name>
            <surname>Kamps</surname>
            , G. Tsakonas,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Manolopoulos</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Iliadis</surname>
          </string-name>
          , I. Karydis (Eds.),
          <source>Research and Advanced Technology for Digital Libraries</source>
          , Springer International Publishing, Cham,
          <year>2017</year>
          , pp.
          <fpage>207</fpage>
          -
          <lpage>219</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Knowledge graph refinement: A survey of approaches and evaluation methods</article-title>
          ,
          <source>Semant. Web</source>
          <volume>8</volume>
          (
          <year>2017</year>
          )
          <fpage>489</fpage>
          -
          <lpage>508</lpage>
          . URL: https://doi.org/10.3233/SW-160218. doi:
          <volume>10</volume>
          . 3233/SW-160218.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Okurowski</surname>
          </string-name>
          ,
          <article-title>Information extraction overview</article-title>
          ,
          <source>in: TIPSTER TEXT PROGRAM: PHASE I: Proceedings of a Workshop held</source>
          at Fredricksburg, Virginia,
          <source>September 19-23</source>
          ,
          <year>1993</year>
          , Association for Computational Linguistics, Fredericksburg, Virginia, USA,
          <year>1993</year>
          , pp.
          <fpage>117</fpage>
          -
          <lpage>121</lpage>
          . URL: https://aclanthology.org/X93-101.
          <year>2doi</year>
          :
          <fpage>10</fpage>
          .3115/1119149.1119164.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ehrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hamdi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Pontes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Romanello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doucet</surname>
          </string-name>
          ,
          <article-title>Named entity recognition and classification in historical documents: A survey</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>56</volume>
          (
          <year>2023</year>
          ). URL: https://doi.org/10.1145/3604931. doi:
          <volume>10</volume>
          .1145/3604931.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Krug</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Weimer</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Reger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Macharowsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Feldhaus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Puppe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Jannidis</surname>
          </string-name>
          ,
          <article-title>Description of a corpus of character references in german novels-droc [deutsches roman corpus]</article-title>
          ,
          <string-name>
            <surname>DARIAH-DE Working Papers</surname>
          </string-name>
          <article-title>27 (</article-title>
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Pontes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Cabrera-Diego</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Boros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Pontes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hamdi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sidère</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Coustaty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doucet</surname>
          </string-name>
          ,
          <article-title>Entity Linking for Historical Documents: Challenges and Solutions</article-title>
          ,
          <source>in: 22nd International Conference on Asia-Pacific Digital Libraries, ICADL</source>
          <year>2020</year>
          , volume
          <volume>12504</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2020</year>
          , pp.
          <fpage>215</fpage>
          -
          <lpage>231</lpage>
          . URL:https://hal. science/hal-03034492. doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -64452-9\_
          <fpage>19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>O.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Hyvönen, From MARC Silos to Linked Data Silosu</article-title>
          ,rl: https://swib.org/ swib16/slides/suominen_silos.pd,
          <year>f2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tietz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bruns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oppenlaender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dessì</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          , DDB-KG:
          <article-title>The German Bibliographic Heritage in a Knowledge Graph</article-title>
          ,
          <source>in: 6th Int. Workshop on Computational History at JCDL - Histoinformatics</source>
          , volume
          <volume>2981</volume>
          , CEUR-WS.org,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Posthumus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          ,
          <article-title>Audio Ontologies for Intangible Cultural Heritage</article-title>
          ,
          <source>in: Proc. of the 19th European Semantic Web Conference - Posters and Demos - ESWC</source>
          <year>2022</year>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          ,
          <source>The DDB Collection and the Limits of Artificial Intelligence</source>
          ,url: https: //swib.org/swib23/slides/06_Mary%
          <article-title>20Ann%20Tan_SWIB2023%20Final.pd,f2023.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Schweter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Akbik</surname>
          </string-name>
          , Flert:
          <article-title>Document-level features for named entity recognition</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>2011</year>
          .06993.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>