<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Computational Humanities Research Conference, November</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Geolocation and Named Entity Recognition in Ancient Texts: A Case Study about Ghewond's Armenian History</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marcella Tambuscio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tara Lee Andrews</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Austrian Center for Digital Humanities and Cultural Heritage (ACDH-CH), Austrian Academy of Sciences</institution>
          ,
          <addr-line>1010 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Vienna</institution>
          ,
          <addr-line>1010 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>1</volume>
      <issue>4</issue>
      <fpage>7</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>We present here a discussion about diferent methods to perform Named Entity Recognition tasks in order to extract geographic entities from the English translation of an Armenian text of the eighth century. Even though many tools are available and perform quite well with modern English, in this case they are only able to detect a very low percentage of the named geographic places. We compared four existing tools: NLTK and spaCy Python libraries, among the most used for NER tasks, TagMe, an entity linking tool that provide an annotation of found entities with Wikipedia pages, and Flair, a PyTorch library. We set these tools in order to select only geographical entities and we also tried two mixed methods: the best results on our data-set have been obtained by combining Flair and TagMe outputs with geographical clustering.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;named entity recognition</kwd>
        <kwd>natural language processing</kwd>
        <kwd>clustering</kwd>
        <kwd>historical corpora</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>extract geographical names and we compared the results. F-measures appeared to exhibit low
values when confronted with the ones usually obtained with contemporary texts. We show
that, by combining the outputs of the two best-performing tools, TagMe and Flair, with a
geographical clustering, we can substantially improve their performance. We additionally modified
the output of the tool TagMe that links the entities in the text to their respective Wikipedia
pages, so that using the Wikipedia API we can select geographical places and directly extract
the coordinates (when available) to create maps. Our main goal here is twofold: on the one
hand, to compare the performance of four well-known NER toolkits on a challenging dataset
(that is part of a larger collection of texts) and discuss the possible reasons of unsatisfying
results; on the other hand, starting from the slightly more promising results of the toolkits
trained on Wikipedia data, to explore new ways to retrieve information from this knowledge
base to improve the results and lay the foundations to train a new and more performative
model in the future.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Named‐entity recognition in historical texts is known to be a challenging problem [
        <xref ref-type="bibr" rid="ref36">38</xref>
        ] and
especially for geographical places that have changed name over time or ceased to exist: empirical
evidence suggests indeed that the more recent the texts, the more entities could be detected
[13].
      </p>
      <p>
        First, researchers have tackled the problem by developing diferent NER methods for
identifying place references in a specific text or corpus, often enriched through the use of gazetteers.
Most of them are rule-based or make use of the NLP Stanford NER tool: historical
newspaper collections [
        <xref ref-type="bibr" rid="ref12 ref23 ref25 ref30">12, 25, 26, 33, 28</xref>
        ], literature [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], British parliamentary proceedings [
        <xref ref-type="bibr" rid="ref16">19</xref>
        ],
Turkish texts [
        <xref ref-type="bibr" rid="ref23">26</xref>
        ], Arabic historical texts [
        <xref ref-type="bibr" rid="ref6">6, 36</xref>
        ], Latin corpora [
        <xref ref-type="bibr" rid="ref15">14</xref>
        ], UK census data [
        <xref ref-type="bibr" rid="ref27">30</xref>
        ].
Similarly, some more general platforms have been created in order to be applied to diferent
data-sets: the VARD tool [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] pre-processes historical corpora to propose modern equivalents
alongside historical spelling variants; a digital geo-temporal gazetteer has been proposed in
[
        <xref ref-type="bibr" rid="ref26">29</xref>
        ]; the Edinburgh Geoparser [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Recogito [
        <xref ref-type="bibr" rid="ref35">37</xref>
        ] recognise mentions of place names in text
and assist in their disambiguation with respect to existing gazetteers. On the other hand, in
[
        <xref ref-type="bibr" rid="ref17">20</xref>
        ] and [18] the authors provide a comparison of some existing geoparsers that make use of
gazetteers, showing that they still have several limitations.
      </p>
      <p>
        Secondly, in the last two decades researchers have been discussing the importance of
Geographic Information Systems (GIS) technologies in the pursuit of historical research [
        <xref ref-type="bibr" rid="ref4">16, 4,
17</xref>
        ] and the necessity of introducing unsupervised methods which would allow a move from
rule-based systems toward more data-driven approaches [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref28">31</xref>
        ] the authors discuss a
combination of spatial analysis and natural language processing techniques in the field of
archaeology. A mixed method that combines NLP and geospatial clustering has been proposed
in [
        <xref ref-type="bibr" rid="ref20">23</xref>
        ] to identify places in housing advertisements: even if the inputs here are not historical
data, the challenge is similar since many local place names either have not been registered in
gazetteers or appear in abbreviated forms that do not appear.
      </p>
      <p>Thirdly, some recent work has suggested that results improve when several tools are
combined. In [39] the authors propose a method that considers five NER tools through a voting
system, which produced a better performance than any single tool. Similarly, GeoTxt [24] is
a geoparser for unstructured streaming text that supports multiple NER methods.</p>
    </sec>
    <sec id="sec-3">
      <title>3. The Dataset</title>
      <p>The History of Ghewond covers events in and around Armenia from ca. 632 to 788, focusing
on the Arab domination of the region and especially the transition from Umayyad to Abbasid
rule and its efects on Armenian politics and society. It is a short text (25K words) but the
geographical range is broad, covering the centers of power of the Caliphate, the territories
both of the former Roman Armenia and the so-called Persarmenia, as well as references to
Byzantine and Khazar places. The impetus for the study was the desire to identify and
collect the context for place names in works of Armenian history, while being fully aware that
NER tools for the Classical Armenian language are not in a particularly advanced state of
development. The solution that seemed obvious was to analyse modern English translations
of the texts. For the initial attempt we used the English translation of Ghewond’s History
made in 2006 by Robert Bedrosian, who is well known for his translations of several works of
Armenian history. Bedrosian’s translation style is to stay as close as possible to the original
text phrasings, and in particular to render all proper names in a direct transliteration from
their forms in the Armenian alphabet. While this is a welcome and helpful translation strategy
from the perspective of historians of the medieval Caucasus, the place names themselves can
be difficult both for untrained human readers and for neural networks to recognise. Adding
to the complication is the fact that, in medieval Armenian society, territories and their ruling
clans often carried the same names; this means that, for example, any occurrence of a name
such as “Rshtunik” must be examined for context to determine whether this is a mention of
a place or a group of persons. In order to compare the outputs of several NER approaches,
we manually created a gold standard to list the geographical entities and their occurrences:
we found 199 entities with 303 total occurrences in the text. The most frequent are Armenia,
Judaea, Vaspurakan, Byzantine territory, Damascus, Byzantium, Asorestan.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <p>All the codes and the data can be found on GitHub 1.</p>
      <sec id="sec-4-1">
        <title>4.1. NER tools</title>
        <sec id="sec-4-1-1">
          <title>First of all we briefly describe here the tools that we used.</title>
          <p>
            NLTK (Natural Language Toolkit)2 and spaCy3 are libraries for NLP written in Python.
They were originally developed for English and perform tasks such as tokenization,
classification and part-of-speech tagging. NLTK [
            <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
            ], developed by the University of Pennsylvania,
is intended to support research while spaCy [22], published under the MIT license, is more
application-oriented. The set of labels ofered by the standard English models of NLTK and
spaCy libraries include a GPE label for geopolitical entities as countries, cities, states and a LOC
label for physical locations as mountain ranges, rivers, seas. We selected entities detected in
our text with both labels.
          </p>
          <p>
            Flair is a relatively new library, open source and developed in Python by the Humboldt
University of Berlin and Zalando Research, that ofers support for common NLP tasks including
Named Entity Recognition [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. This service seems to be more powerful than spaCy, but it must
1https://github.com/tambu85/ancient_text_NER
2https://www.nltk.org/
3https://spacy.io/
be observed that Flair is (a bit) slower and is currently available for only a few languages. For
our purposes we selected only entities classified with the LOC label.
          </p>
          <p>
            TagMe4 is an impressive entity linking tool, developed by the University of Pisa[15], that
identifies meaningful spots in an unstructured text and links each of them to a pertinent
Wikipedia page. For this reason it has been used for disambiguation tasks [
            <xref ref-type="bibr" rid="ref32">35</xref>
            ]. TagMe usually
performs very well with short texts, but it can also be used on longer ones. Given an input text,
the TagMe API 5 provides a list of annotations, meaning a list of pairs (spot,entity), where
each spot is a substring of the input text and each entity is a reference to a unique Wikipedia
page representing the meaning of that spot, in that context. TagMe computes for each entity
a link probability lp that measures how frequently the spot text is used to link exactly that
entity page. Moreover, TagMe associates a value ρ (rho) to each annotation, which estimates
a confidence score of the annotation among the possible entities. In TagMe there is also a
parameter that can be used to fine-tune the disambiguation process, either to select the most
common topics for a spot or to take the context of each spot more into account (we selected this
second option). This parameter could be useful when annotating particularly fragmented text,
such as tweets, where it would be better to favor the most common topics because the context
is less reliable for disambiguation. Supported values are floats in the range [0,0.5], default
is 0.3. It should be noted that TagMe itself does not provide any sort of classification of the
entities and it was not designed for this task. Nevertheless, we noticed that it was able to detect
many geographical entities that the other tools missed: one reason could be that Wikipedia
often reports also the ancient names of places. Then we added our own basic classification
by filtering the results using a SPARQL query through the Wikidata Query Service 6 to select
only geographical places. Moreover, when the geographical coordinates were available in the
page, we added them to the output: in the appendix we provide all the statistics about this
extension of TagMe.
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Mixed approaches</title>
        <p>We applied the above-mentioned tools to our historical text and evaluated the results,
comparing them with a gold standard of manual annotations: in the next section we will provide
the well-known evaluation measures. None of the tools provided exceptional results, but the
best ones were obtained by TagMe and Flair, even if these tools are significantly slower than
spaCy and NLTK. We then tried to improve the results, proposing two other methods:
• M1 we simply considered the union of the results of TagMe and Flair;
• M2 we ran a geographical clustering to discard non-meaningful TagMe results and then
we considered the union of these filtered results with Flair.</p>
        <p>Note that in M2 we are considering only the entities with coordinates, and among them
we perform another selection with clustering. We chose DBSCAN, a popular density-based
clustering algorithm, and we set the parameters (ϵ = 10, minpts = 30) following the standard
procedure and computing the k-nearest neighbors (k-NN) for diferent values of k [ 21]. In
section 5 we show how this is a valid approach to remove many false positives from the TagMe
output.</p>
        <p>4https://sobigdata.d4science.org/web/tagme/tagme-help
5An official wrapper for Python is available here: https://github.com/marcocor/tagme-python
6https://query.wikidata.org/</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Validation and Evaluation</title>
        <p>We selected the LOC and GPE entities detected by NLTK, spaCy and Flair and the Wikipedia
entities labeled as geographical ones by our SPARQL query. Then we validated the results
by comparing them with a gold standard that we produced manually. We had to define some
rules for some corner cases:
• Ethnonym references. Many entities are of this form: Armenian (lords), land of the
Aghuanians or Byzantine territory. Nevertheless, we have to report that most of
the time such expressions are not detected by any of the methods described (with some
exception recognized by TagMe and Flair). We decided to consider as true positives only
expressions that include a geographical element such as the country of Armenians or
Byzantine territory.
• Temporally ambiguous results. An example is Syria, for which TagMe occasionally
provides a link to the modern republic, while in the text the author is referring to
the Roman province; these two geographical areas overlap but are not coterminous. In
the Appendix we report the percentage of the completely matching entities. In this
context, in order to compare TagMe usefully with the other tools, we did not consider
the entity linking (EL) evaluation, but only the NER task: the annotation is labeled as
correct if the spot is indeed a location in essentially the right place. Conversely, when we
had to deal with geographical clustering (mixed method M2) only the entities with the
correct coordinates were considered valid. Another example is the entity Iberia that
TagMe associates (in diferent occurrences) to two diferent Wikipedia pages: Iberian
Peninsula and Kingdom of Iberia. For NER task, both are considered correct since
Iberia is a geographical entity in the text, while for EL task in M2 only the second is
considered correct.</p>
        <p>
          To evaluate the results we used the three common measures in classification tasks: precision,
recall and F1-measure [
          <xref ref-type="bibr" rid="ref3 ref31">3, 34</xref>
          ]. Here we briefly review their formula and meaning in our context.
Our task is to find geographical entities and we have a gold standard (the correct list) to
compare them. For each tool we can count:
• TP (true positives), number of entities correctly labeled as locations;
• FP (false positives), number of entities incorrectly labeled as locations;
• FN (false negatives), number of entities which were not labeled as locations but should
have been.
        </p>
        <p>From this we can compute precision and recall:
Tool</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In Table 1 we give the evaluation measures that were computed for the diferent methods that
we used to detect entities defining locations in Ghewond’s History.</p>
      <p>The values are surprisingly low if compared with the average results on known datasets,
where these NER tools usually give F-measures that far exceed 0.9. These low values represent a
clear example of how the NER task can still present research challenges, especially on historical
texts. Some comments about the comparison:
• NLTK is the worst-performing tool, with very low values for each measure;
• spaCy performs a little better than NLTK in terms of precision, meaning that there is
a smaller fraction of FP (entities incorrectly labeled as locations) but the recall is very
low due to a high number of missed entities (FN);
• TagMe performs slightly better than the previous two, with similar values for precision
and recall;
• Flair is the best among the four tools and exhibits the best precision score among all the
proposed methods.</p>
      <p>It is important to note that low recall scores are partially a consequence of our choice to
consider as TP expressions like country / land of the Armenians since they appear many
times in the text but are only rarely captured by the tools. This is not necessarily a problem,
since we could manually tune many of these tools by adding some specific context-based pattern
matching rules to detect such expressions.</p>
      <p>While Flair clearly exhibited the best performance, it must be observed that TagMe detected
many entities that were not captured by the other three, particularly locations of places referred
to using obsolete names (for instance the Roman Province Judaea or the medieval name of
Istanbul, Constantinople). This is due to the fact that TagMe links entities in the text to
their Wikipedia entries, which are often reachable under the several names by which these
places were known during diferent eras. Moreover, by adding the geographical coordinates we
were able to enrich the output and, since Wikipedia is constantly growing, when repeating the
experiment on Ghewond’s History or other similar texts we can expect to obtain ever better
results.</p>
      <p>Since there was a significant subset of entities detected only by TagMe, our next experiment
was to use a mixed approach M1, where we considered the union of the outputs given by Flair
longilotnude
longlointude
and TagMe. This gave the best recall score but a lower precision with respect to Flair by
itself: this means that (as would be expected) Flair and TagMe together are less likely to miss
locations, but still produce a large number of entities incorrectly labeled as places. Since Flair
had a high value for precision, the imprecision of the M1 approach clearly originates in TagMe.
To partially mitigate this problem, we made use of the coordinates returned by TagMe: we
selected only those returned entities with geographical coordinates and we ran a clustering
algorithm to detect noise.</p>
      <p>In Figure 1 (top) we plot the validated entities with coordinates that were found within the
text using TagMe and Wikipedia queries: blue circles are the true positives, while red crosses
are the false positives. The validation used in this case is a strict one: if an entity appears as a
location, but the coordinates are wrong, we label it as a false positive. We can observe that the
true positives are quite clustered so we ran DBSCAN, an unsupervised clustering algorithm:
we remark that (as yet) we are not interested in how many or which clusters the algorithm
detected; we have merely used the clustering to remove many false positives, labeled as noise.
In Figure 1 (bottom) we show that this method works quite well, which led us to consider only
the entities within the clusters for the M2 method.</p>
      <p>Finally we tried the second mixed approach M2, in which we consider (as in M1) the union
of the outputs given by Flair and TagMe with a clustering step. As shown in the last row of
Table 1, M2 was the best among the six approaches: even if it is not optimal, we obtained the
highest F-score and a good balance among precision and recall. In the appendix we discuss
the results in detail, considering the advantages and disadvantages of the various approaches.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>In this paper we propose a case study on automatic detection of geographical entities in a
corpus originally written in a medieval minority language. Although an English translation
of the text was available for use, the fact that the place names in the text refer to a totally
diferent geopolitical system is actually an obstacle for a system trained on modern English,
meaning that “out-of-the-box” NER tools fail most of the time. Indeed, NLTK and spaCy,
the most well-known tools for NER tasks, obtained very low F-measures on our corpus. Our
attempts with TagMe (designed for Entity Linking) and Flair, two diferent tools both trained
on Wikipedia data, do provide better results. Although these two tools are significantly slower
than NLTK and spaCy, and the execution time can be very important for NLP tasks on
streaming and real-time data, we do not consider this a major problem for a NER task run on
a historical text, since it is likely to be ran only once. Moreover, the modified version of TagMe
that we have devised is even slower due to the use of the Wikipedia API to get coordinates,
which is also impacted by a rate limitation on requests. Concerning the quality of performance
of these tools, we must bear in mind that, since Wikipedia is an online encyclopedia maintained
by a community of volunteer editors, it continuously changes over time: this means that the
results of a repeated analysis could vary (hopefully for the better), or even that larger common
knowledge databases could ofer alternative solutions in the future. Finally, we tried two mixed
approaches and found that by combining Flair and TagMe results with clustering techniques
we were able to significantly improve their performance. It is, however, important to note
that this approach can also depend on the type of data: since we knew that most of the
events described in the text happened in a circumscribed area, clustering was helpful in that
it allowed us to discard some entities that were wrongly classified as places. This could also be
the case for other historical texts, but more detailed research on larger datasets could provide
new insights about the usefulness of geographical clustering for entities.</p>
      <p>We see diferent future steps for this research line: (i) this is an initial a case study, so more
tests are needed on other corpora which could also include some comparison with other tools
that gave similar F1 scores on other datasets [24]; (ii) the detection process using TagMe and
geographical clustering could still be improved; (iii) the mixed approach of Flair and TagMe
(maybe perhaps with additional metadata from Wikipedia) could also be used for other types
of entities such as person and organisation names.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <sec id="sec-7-1">
        <title>Thanks to the developers of NLTK, spaCy, Flair and TagMe.</title>
        <p>[15] P. Ferragina and U. Scaiella. “Tagme: on-the-fly annotation of short text fragments (by
wikipedia entities)”. In: Proceedings of the 19th ACM international conference on
Information and knowledge management. 2010, pp. 1625–1628.
[16]</p>
        <p>M. F. Goodchild and L. L. Hill. “Introduction to digital gazetteer research”. In:
International Journal of Geographical Information Science 22.10 (2008), pp. 1039–1044.
[17] I. N. Gregory and A. Hardie. “Visual GISting: bringing together corpus linguistics and
Geographical Information Systems”. In: Literary and linguistic computing 26.3 (2011),
pp. 297–314.
[18]</p>
        <p>M. Gritta, M. T. Pilehvar, N. Limsopatham, and N. Collier. “What’s missing in
geographical parsing?” In: Language Resources and Evaluation 52.2 (2018), pp. 603–623.</p>
        <p>M. Won, P. Murrieta-Flores, and B. Martins. “ensemble named entity recognition (ner):
evaluating ner Tools in the identification of Place names in historical corpora”. In:
Frontiers in Digital Humanities 5 (2018), p. 2.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>A. Detailed analysis of TagMe results</title>
      <p>In the course of our work we also explored whether we can use the values of the parameters ρ
and lp to better tune the TagMe tool. These values can be used to discard annotations that
are below a given threshold. First of all we visualize the True Positive (TP) and False Positive
(FP) in a scatter plot (see Fig. 2). Even if there are not well-separated clusters, we can see
that there is a higher density of FP for low values of ρ and lp.</p>
      <p>
        We then computed how precision, recall and F1-measures vary when moving the thresholds
of ρ and lp in [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]. When we fix a threshold τ all the TP obtained for values lower than τ
become FN (missed entities). Results are shown in Fig.3 (varying one parameter at a time)
and Fig.4 (varying both parameters). As we can see, the F-measure slowly decreases at the
beginning and then falls. This accords with the recommendation of the TagMe authors, who
indicate values between 0.1 and 0.3 as reasonable threshold for ρ. In our case, we set 0.1 as
threshold for ρ and 0.05 for lp (the recommended standard), but such an approach should be
repeated in other data sets to explore the role of these parameters.
      </p>
      <p>Finally in Tab.2 we report the rates of TP associated with the correct Wikipedia entity
and coordinates, obtained with a double manual validation: TagMe identified the right entity
81% of the times over all detected geographical entities and provided coordinates for 71% of
them. By combining both we obtained the result that 59% of the detected entities are linked
to the right Wikipedia pages that exhibit coordinates. This could surely help in automatically
providing a map of the diferent entities.</p>
      <p>Right Entity Linking</p>
      <p>Coordinates</p>
      <p>Right Entity Linking and Coordinates
0.81
0.72</p>
      <p>0.61</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Akbik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Blythe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Vollgraf</surname>
          </string-name>
          . “
          <article-title>Contextual String Embeddings for Sequence Labeling”</article-title>
          .
          <source>In: COLING</source>
          <year>2018</year>
          , 27th International Conference on Computational Linguistics.
          <year>2018</year>
          , pp.
          <fpage>1638</fpage>
          -
          <lpage>1649</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Alex</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Byrne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Grover</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Tobin</surname>
          </string-name>
          . “
          <article-title>Adapting the Edinburgh geoparser for historical georeferencing”</article-title>
          .
          <source>In: International Journal of Humanities and Arts Computing</source>
          <volume>9</volume>
          .1 (
          <issue>2015</issue>
          ), pp.
          <fpage>15</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ribeiro-Neto</surname>
          </string-name>
          , et al.
          <article-title>Modern information retrieval</article-title>
          . Vol.
          <volume>463</volume>
          . ACM press New York,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T. J.</given-names>
            <surname>Bailey</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Schick</surname>
          </string-name>
          . “
          <string-name>
            <surname>Historical</surname>
            <given-names>GIS</given-names>
          </string-name>
          <article-title>: enabling the collision of history and geography”</article-title>
          .
          <source>In: Social Science Computer Review 27.3</source>
          (
          <issue>2009</issue>
          ), pp.
          <fpage>291</fpage>
          -
          <lpage>296</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baron</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Rayson</surname>
          </string-name>
          . “
          <article-title>VARD2: A tool for dealing with spelling variation in historical corpora”</article-title>
          .
          <source>In: Postgraduate conference in corpus linguistics</source>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Bidhendi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Minaei-Bidgoli</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>H.</given-names>
            <surname>Jouzi</surname>
          </string-name>
          . “
          <article-title>Extracting person names from ancient Islamic Arabic texts”</article-title>
          .
          <source>In: Proceedings of Language Resources and Evaluation for Religious Texts (LRE-Rel) Workshop Programme, Eight International Conference on Language Resources and Evaluation (LREC</source>
          <year>2012</year>
          ).
          <year>2012</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          . “
          <article-title>NLTK: the natural language toolkit”</article-title>
          .
          <source>In: Proceedings of the COLING/ACL 2006 Interactive Presentation Sessions</source>
          .
          <year>2006</year>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          , E. Klein, and
          <string-name>
            <given-names>E.</given-names>
            <surname>Loper</surname>
          </string-name>
          .
          <source>Natural Language Processing with Python. 1st.</source>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          , Inc.,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Borin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kokkinakis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Olsson</surname>
          </string-name>
          . “
          <article-title>Naming the past: Named entity and animacy recognition in 19th century Swedish literature”</article-title>
          .
          <source>In: Proceedings of the Workshop on Language Technology for Cultural Heritage Data (LaTeCH</source>
          <year>2007</year>
          ).
          <year>2007</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Brooke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hammond</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. Hirst. “</surname>
          </string-name>
          <article-title>GutenTag: an NLP-driven tool for digital humanities research in the Project Gutenberg corpus”</article-title>
          .
          <source>In: Proceedings of the Fourth Workshop on Computational Linguistics for Literature</source>
          .
          <year>2015</year>
          , pp.
          <fpage>42</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          , J. Baldridge,
          <string-name>
            <given-names>M.</given-names>
            <surname>Esteva</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          . “
          <article-title>The substantial words are in the ground and sea: computationally linking text and geography”</article-title>
          .
          <source>In: Texas Studies in Literature and Language 54.3</source>
          (
          <issue>2012</issue>
          ), pp.
          <fpage>324</fpage>
          -
          <lpage>339</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12] [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Crane</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Jones</surname>
          </string-name>
          . “
          <article-title>The challenge of virginia banks: an evaluation of named entity analysis in a 19th-century newspaper collection”</article-title>
          .
          <source>In: Proceedings of the 6th ACM/IEEECS joint conference on Digital libraries. 2006</source>
          , pp.
          <fpage>31</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Ehrmann</surname>
          </string-name>
          , G. Colavizza,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rochat</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          . “
          <article-title>Diachronic evaluation of NER systems on old newspapers”</article-title>
          .
          <source>In: Proceedings of the 13th Conference on Natural Language Processing (KONVENS</source>
          <year>2016</year>
          ).
          <source>Conf. Bochumer Linguistische Arbeitsberichte.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <year>2016</year>
          , pp.
          <fpage>97</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Erdmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Joseph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Janse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ajaka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Elsner</surname>
          </string-name>
          , and M.-C. de Marnefe.
          <article-title>“Challenges and solutions for Latin named entity recognition”</article-title>
          .
          <source>In: COLING 2016: 26th International Conference on Computational Linguistics. Association for Computational Linguistics</source>
          .
          <year>2016</year>
          , pp.
          <fpage>85</fpage>
          -
          <lpage>93</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Grover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Givon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tobin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ball</surname>
          </string-name>
          . “
          <article-title>Named Entity Recognition for Digitised Historical Texts</article-title>
          .” In: Lrec.
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Grover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tobin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Byrne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Woollard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Reid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dunn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ball</surname>
          </string-name>
          . “
          <article-title>Use of the Edinburgh geoparser for georeferencing digitized historical collections”</article-title>
          .
          <source>In: Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences</source>
          <volume>368</volume>
          .
          <year>1925</year>
          (
          <year>2010</year>
          ), pp.
          <fpage>3875</fpage>
          -
          <lpage>3889</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Hahsler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Piekenbrock</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Doran</surname>
          </string-name>
          .
          <article-title>“dbscan: Fast density-based clustering with R”</article-title>
          .
          <source>In: Journal of Statistical Software 91.1</source>
          (
          <issue>2019</issue>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Honnibal</surname>
          </string-name>
          , I. Montani,
          <string-name>
            <surname>S. Van Landeghem</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Boyd</surname>
          </string-name>
          . spaCy:
          <string-name>
            <surname>Industrial-strength Natural</surname>
          </string-name>
          Language Processing in Python.
          <year>2020</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.1212303.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mao</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. McKenzie.</surname>
          </string-name>
          “
          <article-title>A natural language processing and geospatial clustering framework for harvesting local place names from geotagged housing advertisements”</article-title>
          .
          <source>In: International Journal of Geographical Information Science 33.4</source>
          (
          <issue>2019</issue>
          ), pp.
          <fpage>714</fpage>
          -
          <lpage>738</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Karimzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pezanowski</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. M. MacEachren</surname>
            , and
            <given-names>J. O.</given-names>
          </string-name>
          <string-name>
            <surname>Wallgrün</surname>
          </string-name>
          . “
          <article-title>GeoTxt: A scalable geoparsing system for unstructured text geolocation”</article-title>
          .
          <source>In: Transactions in GIS 23.1</source>
          (
          <issue>2019</issue>
          ), pp.
          <fpage>118</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Kettunen</surname>
          </string-name>
          , E. Mäkelä,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ruokolainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kuokkala</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Löfberg</surname>
          </string-name>
          . “
          <article-title>Old content and modern tools-searching named entities in a Finnish OCRed historical newspaper collection 1771-1910”</article-title>
          . In: arXiv preprint arXiv:
          <volume>1611</volume>
          .02839 (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>D.</given-names>
            <surname>Küçük</surname>
          </string-name>
          et al. “
          <article-title>Named entity recognition experiments on Turkish texts”</article-title>
          .
          <source>In: International Conference on Flexible Query Answering Systems</source>
          . Springer.
          <year>2009</year>
          , pp.
          <fpage>524</fpage>
          -
          <lpage>535</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sun</surname>
          </string-name>
          , J. Han, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>“A survey on deep learning for named entity recognition”</article-title>
          .
          <source>In: IEEE Transactions on Knowledge and Data Engineering</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S. Mac</given-names>
            <surname>Kim</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Cassidy</surname>
          </string-name>
          . “
          <article-title>Finding names in trove: named entity recognition for Australian historical newspapers”</article-title>
          .
          <source>In: Proceedings of the Australasian Language Technology Association Workshop 2015</source>
          .
          <year>2015</year>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>65</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>H.</given-names>
            <surname>Manguinhas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Martins</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Borbinha</surname>
          </string-name>
          . “
          <article-title>A geo-temporal web gazetteer integrating data from multiple sources”</article-title>
          .
          <source>In: 2008 Third international conference on digital information management. Ieee</source>
          .
          <year>2008</year>
          , pp.
          <fpage>146</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>P.</given-names>
            <surname>Murrieta-Flores</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baron</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gregory</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hardie</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Rayson</surname>
          </string-name>
          . “
          <article-title>Automatically analyzing large texts in a GIS environment: The registrar general's reports and cholera in the 19th century”</article-title>
          .
          <source>In: Transactions in GIS 19.2</source>
          (
          <issue>2015</issue>
          ), pp.
          <fpage>296</fpage>
          -
          <lpage>320</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>P.</given-names>
            <surname>Murrieta-Flores</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Gregory.</surname>
          </string-name>
          “
          <article-title>Further frontiers in GIS: Extending spatial analysis to textual sources in archaeology”</article-title>
          .
          <source>In: Open Archaeology</source>
          <volume>1</volume>
          .open-issue (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nadeau</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Sekine</surname>
          </string-name>
          . “
          <article-title>A survey of named entity recognition and classification”</article-title>
          .
          <source>In: Lingvisticae Investigationes 30.1</source>
          (
          <issue>2007</issue>
          ), pp.
          <fpage>3</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>C.</given-names>
            <surname>Neudecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wilms</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Faber</surname>
          </string-name>
          , and
          <string-name>
            <surname>T. van Veen. “</surname>
          </string-name>
          <article-title>Large-scale refinement of digital historic newspapers with named entity recognition”</article-title>
          .
          <source>In: Proc IFLA Newspapers/GENLOC Pre-Conference Satellite Meeting</source>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>D.</given-names>
            <surname>Powers</surname>
          </string-name>
          . “
          <article-title>Evaluation: From Precision, Recall and</article-title>
          <string-name>
            <surname>F-Measure to</surname>
            <given-names>ROC</given-names>
          </string-name>
          , Informedness, Markedness &amp; Correlation”.
          <source>In: Journal of Machine Learning Technologies 2.1</source>
          (
          <issue>2011</issue>
          ), pp.
          <fpage>37</fpage>
          -
          <lpage>63</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [35] [36]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rovera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nanni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Goy</surname>
          </string-name>
          . “
          <article-title>Domain-specific named entity disambiguation in historical memoirs”</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          . Vol.
          <year>2006</year>
          . Rwth.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          2017, Paper-
          <volume>20</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Shaalan.</surname>
          </string-name>
          “
          <article-title>A survey of arabic named entity recognition and classification”</article-title>
          .
          <source>In: Computational Linguistics 40.2</source>
          (
          <issue>2014</issue>
          ), pp.
          <fpage>469</fpage>
          -
          <lpage>510</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>R.</given-names>
            <surname>Simon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Barker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Isaksen</surname>
          </string-name>
          , and P. de Soto Cañamares. “
          <article-title>Linking early geospatial documents, one place at a time: annotation of geographic documents with Recogito”</article-title>
          .
          <source>In: e-Perimetron 10.2</source>
          (
          <issue>2015</issue>
          ), pp.
          <fpage>49</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>S.</given-names>
            <surname>Van Hooland</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. De Wilde</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Verborgh</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Steiner</surname>
          </string-name>
          , and R. Van de Walle. “
          <article-title>Exploring entity recognition and disambiguation for cultural heritage collections”</article-title>
          .
          <source>In: Digital Scholarship in the Humanities 30.2</source>
          (
          <issue>2015</issue>
          ), pp.
          <fpage>262</fpage>
          -
          <lpage>279</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>