<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Linked Data for the Automatic Enrichment of Historical Archives</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gary Munnelly</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harshvardhan J. Pandit</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Seamus Lawless</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Adapt Centre, Trinity College Dublin</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>With the increasing scale of online cultural heritage collections, the e orts of manually adding annotations to their contents become a challenging and costly endeavour. Entity Linking is a process used to automatically apply such annotations to a text based collection, where the quality and coverage of the linking process is highly dependent on the knowledge base that informs it. In this paper, we present our ongoing e orts to annotate a corpus of 17th century Irish witness statements using Entity Linking methods that utilise Semantic Web techniques. We discuss problems faced in this process and attempts to remedy them.</p>
      </abstract>
      <kwd-group>
        <kwd>entity linking</kwd>
        <kwd>ontology creation</kwd>
        <kwd>automatic enrichment</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>di erent items and DPLA2 hosts almost 21 million resources), it becomes
increasingly worthwhile to consider how we might deal with these limitations for
collections both great and small.</p>
      <p>In this paper we present a discussion on the role of semantic web resources in
the task of automatically enriching digitised cultural heritage collections using
EL methods. Our discussion is motivated by ongoing e orts to annotate a
collection of 17th century Irish witness statements so that they may be integrated
as semantic web resources and avail of bene ts such enrichment provides. We
present some of the challenges faced and lessons learned in the course of this
endeavour. The contribution of this paper is to demonstrate and emphasise the
importance of structure in Semantic Web resources. With due consideration it
is possible to create new ontologies which may help to facilitate the EL process.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <sec id="sec-2-1">
        <title>Entity Linking</title>
        <p>Entity Linking (EL) refers to a speci c challenge in computer science whereby
a series of unknown textual mentions of entities (commonly termed \surface
forms") are provided as input to a disambiguation service. The service is tasked
with mapping each of the surface forms to an unambiguous referent entity. To
provide a concrete example, given the input sentence, \I Henry Jones Doctor in
Divinity in obedience to his majesties Commission..." and a request to identify
the entity \Henry Jones", an EL service might return a reference to the URI
http://dbpedia.org/page/Henry_Jones_(bishop), identifying the subject of
the reference as the 17th century Anglican Bishop, as opposed to the ctional
character played by Sean Connery in the 1989 lm \Indiana Jones and the Last
Crusade".</p>
        <p>In order to perform this mapping process, an EL system fundamentally
requires two components:
1. a knowledge base that stores information about all the entities of which the
system is aware, and
2. a referent selection method, which uses evidence extracted from the
knowledge base and present in any prevailing information surrounding the surface
form to arrive at a set of likely referents for each ambiguous mention.</p>
        <p>Given an ambiguous set of mentions, the EL system retrieves from the
knowledge base, a set of candidate referents to which an entity mention many be
referring. This is usually based on some fuzzy retrieval method. A variety of heuristics
are applied and the system eliminates candidate referents which are unlikely to
be the subjects of the mentions. Eventually it arrives at a set of mappings from
textual mentions to knowledge base URIs which unambiguously identi es the
referents.</p>
        <sec id="sec-2-1-1">
          <title>2 https://dp.la/</title>
          <p>Numerous di erent methods and approaches to EL may be freely found in
the literature [6{8]. Almost universally, these methods use some form of graph
based measure as one of the heuristics in the referent selection process. After
the candidates have been retrieved from the knowledge base, a graph derived
from the relationships between the entities may be constructed. The nature of
these relationships varies, but usually it is based on links between corresponding
Wikipedia pages. If a strong network exists between a number of candidate
referents, then there is a good chance that they are the correct disambiguation
choice for the given set of mentions.</p>
          <p>
            It is also common to augment the graph weights using contextual cues
derived from the words surrounding an entity mention [
            <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
            ]. We note this because
systems which consider this feature are based on the assumption that some
contextual description of each entity exists in the knowledge base.
          </p>
          <p>
            The structure and content of the knowledge base is crucial, not only for
informing the disambiguation service that an entity exists, but also for providing
information which helps the the disambiguation algorithm to distinguish good
referents from poor referents. Many modern systems make use of DBpedia3 [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]
and YAGO4 [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] for this task. These are a good choice for most problems due to
the prevalence of links between entities and the long form descriptions of
entities obtained from their corresponding Wikipedia articles. However, for cultural
heritage collections it is often the case that the range of information contained
in these Semantic Web resources are not complete enough to capture the variety
of entities we see in cultural heritage collections.
          </p>
          <p>EL systems have the potential to be extremely helpful when enriching
cultural heritage collections with semantic data. These fully automated systems are
capable of deducing suitable annotations for raw, at, textual documents based
on information that is fed to them via a knowledge base. It is easy to see how a
suitably informed EL system might dramatically ease the process of semantically
linking new cultural heritage artifacts as they are digitised.</p>
          <p>Further discussion could be had surrounding the precise point in the
digitisation process at which EL is applied. Are we linking metadata which has already
been normalised by an expert, or is the system capable of dealing with the noisy,
original, primary source content from which the digital artifact is derived? In
the case of the latter, how does the system manage archaic references, evolving
entities and other such anomalies present in the source collection?
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Automatic Enrichment in cultural heritage</title>
        <p>There have been a number of e orts to investigate the e ectiveness of EL
methods in the automatic enrichment of cultural heritage collections.</p>
        <p>A Europeana led task force produced a series of reports in 2015 which
document their experience with evaluating di erent cultural heritage enrichment
3 http://wiki.dbpedia.org/
4 https://www.mpi-inf.mpg.de/departments/databases-and-information-systems/
research/yago-naga/yago/#c10444
services and sourcing di erent descriptive vocabularies as targets for the
annotation process. Their focus was on annotating metadata for digitised artifacts.
The content of this metadata ranged from speci c elds comprised of a single
entity e.g. dc:creator, dc:publisher to more general, free-form data such as
dc:description.</p>
        <p>
          As part of the investigation, a comparative evaluation of seven cultural
heritage EL services was conducted [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Each service used a di erent vocabulary for
enrichment, however the investigators were able to normalise the annotations by
exploiting the fact that many of them made reference to corresponding DBpedia
and Geonames entities. Using an evaluation dataset that was developed based on
a combination of automatic enrichment tools and manual human investigation
the report showed that the accuracy of the targeted EL tools was extremely high
for the chosen collections.
        </p>
        <p>
          However, as a variety of previous studies have shown, while the accuracy of
EL methods may be high, quite often only a very small percentage of entities
contained in cultural heritage datasets may actually be linked with a referent. A
recent study we performed on the 1641 depositions (see Section 3.1) showed that
a human annotator could only identify referents for 33% of the people and
locations in the depositions [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. We would compare this to e orts by other scholars
such as Agirre [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], who attempted to link Europeana artifacts to Wikipedia
articles and discovered that only 22% of entities that he identi ed could be
annotated in this manner. This is an important limitation of which we must be
aware.
        </p>
        <p>One aspect of the problem is simply that cultural heritage collections are
so incredibly diverse, complicated and unique that nding a suitable Semantic
Web resource with adequate coverage for all purposes is nigh impossible. This
presents the question, how should we annotate a cultural heritage collection
when an appropriate Semantic Web resource cannot be found? Moreover (and
of particular importance to our own research) how should these new Semantic
Web resources be structured in order to aid the automatic enrichment process?</p>
        <p>
          Of particular note for this discussion is the work of Brando et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] on
the REDEN project which investigated methods of using multiple knowledge
bases for disambiguation. This is an interesting approach which may help to
ll the gaps in popular knowledge bases using the information contained in
more tailored ones. In their experiments DBpedia was used in conjunction with
the Bibliotheque Nationale de France (BnF) ontology on a collection of French
literary works.
        </p>
        <p>REDEN's candidate selection phase is based on a literal string comparison
between the surface form and entities in the knowledge base. All candidates from
all source ontologies are retrieved and a resolution step based on owl:sameAs and
skos:exactMatch properties resolves duplicate mentions into a single reference.
Once the candidates have been appropriately pruned, a degree centrality measure
is used to select the referents.</p>
        <p>REDEN demonstrated that developing EL methods which can avail of
multiple knowledge bases may help with poor coverage, but this requires that it
be possible to establish reliable, accurate mappings between ontologies. Indeed,
this property also facilitated the evaluation conducted by the Europeana task
force. This is an important consideration when developing new vocabularies for
cultural heritage collections.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Enriching the 1641 Depositions</title>
      <p>Our own research has focused on attempts to automatically annotate a
collection of 17th century manuscripts using EL methods. This has been extremely
challenging for a variety of reasons. However, we believe our experiences are a
reasonably typical example of problems faced in this eld. Below, we document
our observations and experience of working with the collection to date.
3.1</p>
      <sec id="sec-3-1">
        <title>The 1641 Depositions</title>
        <p>The 1641 depositions are a collection of letters and witness statements taken from
the people of Ireland during the 1641 Irish rebellion. The physical manuscripts
are comprised of approximately 19,000 pages bound in 31 volumes. Ireland in
1641 was a tumultuous place, and while the accuracy of some of the witness
statements may be questionable, the depositions provide an unparalleled window
into this dark chapter in Irish history.</p>
        <p>The depositions have been digitised, transcribed and annotated by a team
of historical scholars who extracted references to people and locations, tagged
depositions based on the nature of their contents, and preserved as much
information about the physical manuscripts as possible including margin notes, original
spelling etc. The resulting documents are stored in a combination of TEI
annotated les and an SQL database. This data rich digital resource presents many
interesting and exciting opportunities for computer scientists to begin
experimenting with methods of analysing and extracting new information from this
historical collection.</p>
        <p>Working with the digital versions of the depositions comes with a number
of challenges, not least of which is the inconsistent nature of the spelling and
grammar used throughout. English was still a developing language in 1641, which
means that a vast array of variant spellings for names and common words
exist across the documents. The below extract from The Deposition of Phillip
Sergeant 5 provides an example of these anomalies:
\And by those faire promisses the said tzpatrick getting possession both
of their persons &amp; goodes, they there behoulding daily cruelties &amp;
murthers vpon other English and belike suspecting the like to be exercised
against themselues, desired ed away secretly o n to to Mountrath"
From the historians' work, we nd that the depositions contain references
to more than 60,000 people and 7,000 locations. The people in question range
5 http://1641.tcd.ie/deposition.php?depID=815351r406
from individuals of great historical importance such as Sir Oliver Cromwell, Sir
Phelim O'Neill, and King Charles I, to individual servants and common folk
who were a ected by the rebellion. Locations similarly range from cities such
as Dublin which still ourish today, to small plots of land which have been lost
either as their names changed or borders shifted.</p>
        <p>We know that several of the entities extracted from the depositions are
duplicates. However, the huge range in spelling variations and naming conventions
makes it extremely di cult to determine which mentions of entities in the
depositions might be references to the same person or place. Compounding this
problem is the fact that the severity of textual noise means that standard NLP
tools can struggle with simpler tasks such as sentence chunking or Named Entity
Recognition. Performing reliable analysis based on the language of the
depositions is, to say the least, di cult.</p>
        <p>
          We are not the rst to attempt to decipher the contents of the depositions
using computational methods. The CULTURA project [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] developed and
applied a range of tools to provide a personalised experience for individuals who
are interested in exploring the collection. This project was extremely successful
and produced a number of valuable utilities for working with collections of this
nature. However, the depositions' content remains in its original SQL database,
disconnected from the Semantic Web.
        </p>
        <p>
          An earlier study conducted on a manually annotated subset of the depositions
attempted to assess the feasibility of automatically enriching the collection using
standard Entity Linking tools [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. From this study it was shown that only 33%
of entities in the annotated subset had a corresponding referent in DBpedia. It
was also observed that, of the ten Entity Linkers evaluated, no single tool could
satisfactorily annotate the test corpus. Individually some did show promise on
speci c aspects of the linking problem, but these were undermined by weaknesses
elsewhere. For example, AGDISTIS [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] often correctly abstained from annotating
where no referent existed in DBpedia, but was generally incorrect in its choice
of referent where one did exist. KEA [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] and Dexter [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] performed respectably
in choosing the correct referent where one existed, but were often overzealous
and applied labels when they should have abstained.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Challenges in Adopting Semantic Web</title>
        <p>While there are many aspects of the depositions that we may choose to model
into a potential ontology to generate linked data, we focus on people and places
due to their perceived importance in the documents. For historians there are still
a number of unanswered research questions about the motivations and in uences
behind the rebellion. A linked data solution may assist them when investigating
these questions. However, modelling the depositions is extremely challenging for
a number of reasons.</p>
        <p>First, there is no true, de nitive list of people and places on which to base the
ontology. Ideally any ontology that we create would be populated with a set of
distinct entities that can be found in the depositions. While there are resources
which can help us to determine this set of entities (as discussed below), there is
still noise present in these sources. Sometimes people are referred to by lineage
rather than their actual name, e.g. \the heirs of Mr. Gale", or even by title e.g.
\Bishop of Meath". Given these ambiguities, there is much risk of accidentally
omitting or con ating entities when the ontology is being constructed.</p>
        <p>Second, the inconsistent language of the depositions means that multiple
variant spellings for people and places can be found throughout the collection. If
a suitable ontology can be constructed to represent each entity, discovering all the
possible variant names by which it may be referenced would be a monumental
task. Sometimes these variations are minor spelling di erences e.g. \Florence
FitzPatrick" being referred to as \F orenc F tz Patrick", but some are more
severe, such as the \Barony of Fassadinin" being referred to as the \Barrony of
assa and Dyninge". Detecting such di erences is di cult through an automated
process and requires an expert to assess its correctness.</p>
        <p>Third, if we are to construct this ontology with an eye to automatic
enrichment, then the inclusion of links between entities and how to establish them is
an important consideration. We could use familial connections, but we are not
aware of any reliable sources which document these in a readily adoptable
manner. On what basis then are we to establish relationships between our entities?
Currently, there is no reliable way to specify that a relation is likely without
stating it as a fact in an ontology.</p>
        <p>We must also exercise some degree of caution in our attempts to annotate
the depositions. If the intention is to assist scholars with their research, then the
information conveyed by the proposed solution must be accurate. This can be a
subtle problem. For example, if we consider the entity \the Pope", should this be
used to describe the role of the head of the Catholic Church, or should it describe
an individual who held that role? If we assume the latter, then we must be sure
to refer to the correct pope for the source document, which involves additional
knowledge that may not be readily available in the knowledge base. Pope Urban
VIII held the position until 1644 when he passed away and was replaced by Pope
Innocent X. Modelling evolving entities such as these is a common problem in
the cultural heritage domain.</p>
        <p>In spite of these challenges, resources do exist which can help us to generate
lists of distinct entities. Three resources at our disposal are:
{ The Down Survey: A complete national survey of land in Ireland after the
rebellion. The survey was conducted in order to establish which lands should
be forfeited as penalty for crimes during the rebellion.
{ The Statute Staple: A record of transactions between individuals. The staple
documents goods bought and sold, and provides information about debts
owed between various parties before the rebellion
{ The Books of Survey and Distribution: A list of properties held by various
land owners. These documents were used to determine taxes based on land
ownership.</p>
        <p>These documents have been the subject of historical research for a number
of years and were some of the major contributing sources for the Petty Maps
project 6.</p>
        <p>
          Given the three resources above, we have begun the process of constructing
an ontology to model the entities present in the depositions. Using standard
record resolution methods based on the DICE coe cient [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and Jaro-Winkler
distance [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], we are able to resolve entities across the three collections. This work
must be checked for integrity by a historian, which is a slow process. Yet once this
resolution has been completed it will be possible to uplift the resulting records
to an RDF representation. This makes it possible to construct an ontology which
describes unique instances of land owners, geographic regions and some of the
relationships between them and can be used to enrich the process of annotating
the depositions.
        </p>
        <p>Our goal is to eventually use the information in this new ontology as a
knowledge base for an EL system. Using the information derived from the three
sources, it may be possible to automatically enrich the depositions with
semantics using standard EL methods. Furthermore, if the entities described by the
ontology are found to be recorded in other historical sources as they are digitised,
it may be possible to use this knowledge to automatically link and integrate them
with the depositions and the greater web of knowledge.</p>
        <p>In order to facilitate EL methods, we are attempting to capture features that
are commonly used by EL algorithms. The variety of surface forms which may
link to an entity will be an important consideration. As yet we are contemplating
which relationships are likely to be the most helpful. Our initial focus is on
attempting to model ownership of land and debts owed between individuals as
monetary records are often well documented based on the resources available to
us.</p>
        <p>Where possible we will attempt to reuse existing knowledge by linking
entities in our ontology to corresponding entries in DBpedia, Geonames or other
similar ontologies using properties such as owl:sameAs. Existing commonly used
vocabularies will be used to describe the entities, though it may be necessary
to extend or create additional properties for accurate representations of entities.
For example, vCard7 has a number of useful properties for describing people,
and while it does capture the concept of honori cs applied to people e.g. \Mr.",
\Mrs.", \Dr." etc., it does not quite distinguish between an honori c and a
title e.g. \Earl" is a title held that is passed on to other people, rather than an
honori c applied to an individual person's name. As yet we are assessing the
suitability of vocabularies such as FOAF, vCard and SIOC8 for describing such
properties.</p>
        <sec id="sec-3-2-1">
          <title>6 http://downsurvey.tcd.ie/index.html</title>
          <p>7 https://www.w3.org/TR/vcard-rdf/
8 https://www.w3.org/Submission/sioc-spec/</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>Performing automatic enrichment of cultural heritage collections is challenging
for a variety of reasons. As evidenced by our own experience, and the documented
experience of other researchers, nding knowledge bases with adequate coverage
for a given cultural heritage resource is extremely di cult. While developing an
entirely new ontology that does not reuse existing knowledge is a solution, if
not done properly it can lead to inaccurate or incompatible knowledge
representations that negate one of the greatest bene ts of linked data i.e. connectivity
among disparate collections. It is of far greater bene t to the community if these
new vocabularies can be integrated with existing semantic web resources in a
seamless fashion.</p>
      <p>Due to the issue with well-known knowledge bases not covering a large
percentage of the entities in specialised cultural heritage collections, it is likely that
curators of such resources will need to develop their own ontologies in order to
accurately represent the semantics of their data. While it is good to expand the
web of knowledge with this new information, we suggest that due care be given
to the structure of these resources and to how this structure may lend itself to
informing automatic enrichment processes going forward. Methods such as
REDEN may exploit owl:sameAs or similar relationships between a new ontology
and more established ones in order to knit together various knowledge bases for
the EL process. If automatic enrichment services can make use of the information
in new linked data resources, then future annotation processes may be expedited
as new collections are digitised and made available.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>The ADAPT Centre for Digital Content Technology is funded under the SFI
Research Centres Programme (Grant 13/RC/2106) and is co-funded under the
European Regional Development Fund.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hendler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lassila</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The semantic web</article-title>
          . Scienti c american
          <volume>284</volume>
          (
          <issue>5</issue>
          ) (
          <year>2001</year>
          )
          <volume>34</volume>
          {
          <fpage>43</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Van</given-names>
            <surname>Hooland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>De Wilde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Verborgh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Steiner</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          , Van de Walle, R.:
          <article-title>Exploring entity recognition and disambiguation for cultural heritage collections</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          <volume>30</volume>
          (
          <issue>2</issue>
          ) (
          <year>2015</year>
          )
          <volume>262</volume>
          {
          <fpage>279</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Wilde</surname>
          </string-name>
          , M.D.:
          <article-title>Improving Retrieval of Historical Content with Entity Linking</article-title>
          . In Morzy, T.,
          <string-name>
            <surname>Valduriez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bellatreche</surname>
          </string-name>
          , L., eds.
          <source>: New Trends in Databases and Information Systems. Communications in Computer and Information Science</source>
          , Springer International Publishing (
          <year>September 2015</year>
          )
          <volume>498</volume>
          {
          <fpage>504</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Stiller</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , Gade, M.,
          <string-name>
            <surname>Isaac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Automatic enrichments with controlled vocabularies in europeana: Challenges and consequences</article-title>
          . In:
          <string-name>
            <surname>Euro-Mediterranean</surname>
            <given-names>Conference</given-names>
          </string-name>
          , Springer (
          <year>2014</year>
          )
          <volume>238</volume>
          {
          <fpage>247</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J</given-names>
            ., Han, J
          </string-name>
          .:
          <article-title>Entity linking with a knowledge base: Issues, techniques, and solutions</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>27</volume>
          (
          <issue>2</issue>
          ) (
          <year>2015</year>
          )
          <volume>443</volume>
          {
          <fpage>460</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ganea</surname>
            ,
            <given-names>O.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganea</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lucchi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eickho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hofmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Probabilistic bagof-hyperlinks model for entity linking</article-title>
          .
          <source>In: Proceedings of the 25th International Conference on World Wide Web, International World Wide Web Conferences Steering Committee</source>
          (
          <year>2016</year>
          )
          <volume>927</volume>
          {
          <fpage>938</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Usbeck</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          , Roder,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gerber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Coelho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.A.</given-names>
            ,
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Both</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Agdistis-graph-based disambiguation of named entities using linked data</article-title>
          .
          <source>In: International Semantic Web Conference</source>
          , Springer (
          <year>2014</year>
          )
          <volume>457</volume>
          {
          <fpage>471</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Yosef</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            <given-names>art</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Bordino</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spaniol</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Aida: An online tool for accurate disambiguation of named entities in text and tables</article-title>
          .
          <source>Proceedings of the VLDB Endowment</source>
          <volume>4</volume>
          (
          <issue>12</issue>
          ) (
          <year>2011</year>
          )
          <volume>1450</volume>
          {
          <fpage>1453</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Zwicklbauer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seifert</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Granitzer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Robust and collective entity disambiguation through semantic embeddings</article-title>
          .
          <source>In: Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval. SIGIR '16</source>
          , New York, NY, USA, ACM (
          <year>2016</year>
          )
          <volume>425</volume>
          {
          <fpage>434</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isele</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakob</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jentzsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontokostas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morsey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Kleef</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Dbpedia{a large-scale, multilingual knowledge base extracted from wikipedia</article-title>
          .
          <source>Semantic Web</source>
          <volume>6</volume>
          (
          <issue>2</issue>
          ) (
          <year>2015</year>
          )
          <volume>167</volume>
          {
          <fpage>195</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago: A core of semantic knowledge</article-title>
          .
          <source>In: Proceedings of the 16th International Conference on World Wide Web. WWW '07</source>
          , New York, NY, USA, ACM (
          <year>2007</year>
          )
          <volume>697</volume>
          {
          <fpage>706</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Manguinhas</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freire</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isaac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stiller</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Charles</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soroa</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simon</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexiev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Exploring comparative evaluation of semantic enrichment tools for cultural heritage metadata</article-title>
          .
          <source>In: International Conference on Theory and Practice of Digital Libraries</source>
          , Springer (
          <year>2016</year>
          )
          <volume>266</volume>
          {
          <fpage>278</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Munnelly</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawless</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Investigating entity linking in early english legal documents</article-title>
          .
          <source>In: Digital Libraries (JCDL)</source>
          , ACM/IEEE Joint Conference on. (in press)
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Agirre</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrena</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lacalle</surname>
            ,
            <given-names>O.L.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soroa</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fern</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevenson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Matching Cultural Heritage items to Wikipedia</article-title>
          . In: LREC. (
          <year>2012</year>
          )
          <volume>1729</volume>
          {
          <fpage>1735</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Brando</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frontini</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganascia</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          :
          <article-title>REDEN: Named Entity Linking in Digital Literary Editions Using Linked Data Sets</article-title>
          .
          <source>Complex Systems Informatics and Modeling Quarterly (7) (July</source>
          <year>2016</year>
          )
          <volume>60</volume>
          {
          <fpage>80</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Steiner</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agosti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sweetnam</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hillemann</surname>
            ,
            <given-names>E.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orio</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponchia</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hampson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Munnelly</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nussbaumer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albert</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al.:
          <article-title>Evaluating a digital humanities research environment: the CULTURA approach</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          <volume>15</volume>
          (
          <issue>1</issue>
          ) (
          <year>2014</year>
          )
          <volume>53</volume>
          {
          <fpage>70</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Waitelonis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sack</surname>
          </string-name>
          , H.:
          <article-title>Named entity linking in# tweets with kea</article-title>
          .
          <source>In: # Microposts</source>
          . (
          <year>2016</year>
          )
          <volume>61</volume>
          {
          <fpage>63</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Ceccarelli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lucchese</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orlando</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perego</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trani</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Dexter: an open source framework for entity linking</article-title>
          .
          <source>In: Proceedings of the sixth international workshop on Exploiting semantic annotations in information retrieval</source>
          ,
          <source>ACM</source>
          (
          <year>2013</year>
          )
          <volume>17</volume>
          {
          <fpage>20</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Dice</surname>
            ,
            <given-names>L.R.</given-names>
          </string-name>
          :
          <article-title>Measures of the amount of ecologic association between species</article-title>
          .
          <source>Ecology</source>
          <volume>26</volume>
          (
          <issue>3</issue>
          )
          <fpage>297</fpage>
          {
          <fpage>302</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Winkler</surname>
          </string-name>
          , W.:
          <article-title>String comparator metrics and enhanced decision rules in the fellegisunter model of record linkage</article-title>
          .
          <source>In: Proceedings of the Section on Survey Research Methods</source>
          . (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>