<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Plague Dot Text: Text mining and annotation of outbreak reports of the Third Plague Pandemic (1894-1952)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arlene Casey</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mike Bennett</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Richard Tobin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claire Grover</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lukas Engelmann</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beatrice Alex</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Edinburgh Futures Institute, School of Literatures, Languages and Cultures University of Edinburgh</institution>
          ,
          <addr-line>Edinburgh</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Language</institution>
          ,
          <addr-line>Cognition and Computation</addr-line>
          ,
          <institution>School of Informatics</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Science Technology and Innovation Studies, School of Social and Political Science</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Edinburgh Library</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The design of models that govern diseases in their relation to population is built on information and data gathered from past outbreaks. However, epidemic outbreaks are never captured in statistical data alone but are communicated by narratives, built on empirical observations. Outbreak reports discuss correlations between populations, locations and the disease to infer insights into causes, vectors and potential interventions. The problem with these narratives is usually the lack of consistent structure that allows for exploration of their collection as a whole. Our interdisciplinary research investigates more than 100 reports from the third plague pandemic (1894-1952) evaluating ways of building a corpus to extract and structure information through text mining and manual annotation. In this paper we discuss the progress of our exploratory project, how we enhance optical character recognition (OCR) methods to improve text capture, our approach to structure the narratives and identify relevant entities in the reports. The structured corpus is made available via Solr enabling search and analysis across the whole collection for future research dedicated e.g. to the identi cation of concepts. The corpus will enable researchers to analyse the reports collectively and allows for deep insights into the global epidemiological consideration of plague in the early 20th century.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The Third Plague Pandemic (1894 - 1950) spread along sea trade routes a ecting almost every
port in the world and almost all inhabited countries, killing millions of people in the late
nineteenth and early twentieth centuries [
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ]. However, as outbreaks di ered in severity,
mortality and longevity, questions emerged at the time of how to identify the common drivers of
the epidemic. After the Pasteurian Alexandre Yersin had successfully identi ed the epidemic's
pathogen in 1894, yersinia pestis [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], the attention of epidemiologists and medical o cers
turned to the speci c local conditions to understand the conditions under which the presence
of plague bacteria turned into an epidemic. These observations were regularly transferred
into reports, written to deliver a comprehensive account of the aspects deemed important by
the respective author. These reports are the underlying data set for ongoing work in the
Plague.TXT project which is conducted by an interdisciplinary team of medical historians,
computer scientists and computer linguists. While each historical report was written as a
standalone document relating to the spread of disease in a particular city, the goal of our work is to
bring these reports together as one systematically structured collection of knowledge, preparing
an annotated corpus made accessible to the wider research community and for follow-on analysis
of narrative structures and concepts.
      </p>
      <p>
        The pandemic reports o er deep insights into the ways in which epidemiological knowledge
about plague was articulated at the time of the pandemic. While the reports contain a wealth
of statistics and tabulated data their main value is found in articulated viewpoints about the
causes for a plague epidemic, about the attribution of responsibility to populations, locations
or climate conditions as well as about evaluating various measurements of control. The corpus
thus constitutes an archive, from which future analysis will discern concepts, with which plague
has been shaped into an object of knowledge in modern epidemiology. Further, the corpus will
allow inferences to be made on the history of narrative epidemiology, a genre that has been
widely overlooked in the historiography of `formal epidemiology' [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Analysis of reports produced during the third plague has been done before, involving mainly
manual collation of data such as collecting statistics across reports for mortality rates. The
derived data has been used to reconstruct transmission trees from localised outbreaks [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and to
study potential sources and transmission across Europe [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In this work we are focused on
collating the knowledge available across our report collection to consider comparable conventions
within the entire corpus. This approach is challenging though as each report di ers in the
presentation, ordering and style of content. Despite their di erences in style, these communicative
reports are intended for the same audience of government o cials and fellow epidemiologists and
will present their arguments comparably. We build a schema structure to label the information
contained in individual reports such that similar discourse segments can be linked and studied
across the collection to enable for example comparative analysis across time and places. In our
structuring and annotation e orts, we apply genre analysis as our methodological approach,
treating each report as a communicative event about a speci c plague outbreak. Whilst these
arguments may di er in style, our work has shown that they can be isolated, annotated and
mapped under one structure.
      </p>
      <p>Our contribution is the development of a systematically structured corpus, which we capture
through annotation, to assimilate similar discourse segments such as causes or treatments across
the reports. In addition, we develop an interactive search interface to our collection eliminating
the need for manual read and search activities. This search tool in combination with our
structured schema allows follow-on research to conduct automated exploration of a rich source
about the conceptual thinking at the time on the plague pandemic, to better understand the
historical epistemology of epidemiology and to thus provide valuable lessons about dealing with
contemporary global spread of disease.</p>
      <p>In the following sections we give an overview of our pilot study describing the data collection,
the challenges presented by OCR and improvements we have made to the original digitised
reports. Following this we describe our annotation process including our schema to structure
the reports to extract information. We discuss our combination of manual annotation and
automated text mining techniques that support the retrieval and structuring of information
from the reports. We discuss aspects of our search interface, enabled through Solr ,which we
use to make the collection available online. Finally we give some examples of potential use of
this interface.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>The Third Plague Pandemic was documented in over 100 outbreak reports for most major cities
around the world. Many of them have been digitised, converted to text via optical character
recognition (OCR) and are available via the Internet Archive and the UK Medical Heritage
Library. See Figure 1 for an example of such a report covering the Hong Kong outbreak which
was published in 1895 and is accessible with open access on Internet Archive.1</p>
      <p>We treat all relevant reports for which we have a scan as one collection. While the majority
of reports in this set (102) are written in English, there are further reports in French, Spanish,
Portuguese and other languages which we excluded from the analysis at this stage.</p>
      <p>Table 1 provides an overview of the data set in terms of counts of sentences and words
in the collection and illustrates the variety of documents in this collection. To derive these
counts we used automatic tokenisation and sentence detection over the raw OCR output which
is part of the text mining pipeline described in section 3. While the smallest document is
only 32 sentences long containing 1,091 word tokens, the largest report contains almost 400,000
word tokens. The collection contains 38 documents with up to 5K words each, 15 reports with
between 5K and 10K words each, 32 documents with between 10K and 100K words each and
17 documents with 100K or more words each. In total, the reports amount to over 4.4 million
word tokens and over 229K sentences.
2.1</p>
      <sec id="sec-2-1">
        <title>OCR Improvements</title>
        <p>When initially inspecting this digitised historical data, we realised that some of the OCR was
of inadequate quality. We therefore spent time during the rst part of the project on improving
the OCR quality of the reports.</p>
        <p>
          Using computer vision techniques, we processed the report images to remove warping
artefacts [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. We then identi ed likely textual areas in report images, and produced an e ective
crop, to provide the OCR engine with less extraneous data.2 OCR is then performed using
Tesseract3, trained speci cally for typeface styles and document layouts common to the time
period of the reports.
        </p>
        <p>Training was done across a range of truth data, covering period documents obtained from the
IMPACT Project Datasets4, documents from Project Gutenberg prepared for OCR training5,
internal ground-truth data compiled as part of the Scottish Session Papers project at the</p>
        <sec id="sec-2-1-1">
          <title>1https://archive.org/details/b24398287</title>
          <p>2http://libraryblogs.is.ed.ac.uk/librarylabs/2017/06/23/automated-item-data-extraction/
3https://opensource.google.com/projects/tesseract
4https://www.digitisation.eu/tools-resources/image-and-ground-truth-resources/
5https://github.com/PedroBarcha/old-books-dataset
Zones
Title-matter
Preface
Content-page
Introduction
Disease history
Outbreak history
Local conditions
Causes
Measures
Clinical appearances
Laboratory
Treatment
Cases
Statistics
Epizootics
Appendix
Conclusion</p>
          <p>Description
Title page
Preface information
Content page information
State of the epidemic at the time of the production of the report,
summary of key features, evaluation of signi cance of the epidemic
General points on the history of the epidemic, origin of outbreak
Geographical and chronological overview of local outbreak. What
happened this place this year
Descriptions of key elements that are considered noteworthy,
something that has contributed or impacted the outbreak
Causes identi ed by the author e.g. usually points of origin, speci c
local conditions or descriptions of import
List of the measures e.g. undertaken to curb the outbreak, sanitary
improvements, quarantines, disinfection or fumigation and rat
catching
Description of the disease appearance, its usual course and its
mortality
Description of bacteriological analysis, human lab work
Description of the treatment given to patients
List of individual cases, usually with age, gender, occupation, course
of disease, and time and dates of infection and death
Contains many lists or tables of statistics such as deaths
Contains information solely about animals, experiments or discussions
Labelled appendix
Conclusion
University of Edinburgh6 and typeface datasets designed for Digital Humanities collections.7</p>
          <p>
            While we have not yet formally evaluated the improvements made to the OCR, observation
of the new OCR output shows clear improvements in text quality. This is important as it a ects
the quality of downstream text mining steps. Previous research and experiments have found
that errors in OCRed text have a negative cascading e ect on natural language processing or
information retrieval tasks [
            <xref ref-type="bibr" rid="ref1 ref10 ref12 ref14">12, 14, 10, 1</xref>
            ]. In future work, we would like to conduct a formal
evaluation comparing the two versions of OCRed text to quantify the quality improvement.
3
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Annotation</title>
      <p>This section describes the schema we implemented to structure the information contained within
the reports and automatic and manual annotation applied to our collection of plague reports.
We rst processed them using a text mining pipeline which we adapted and enhanced speci
cally for this data. The text mining output annotations are then corrected and enriched during
a manual annotation phase which is still ongoing. Each report that has undergone manual
annotation then undergoes further automatic geo-resolution and date normalisation to normalise
and disambiguate di erent mentions of location names and dates in the text.</p>
      <sec id="sec-3-1">
        <title>6https://www.projects.ed.ac.uk/project/luc020/brief/overview 7https://github.com/jbest/typeface-corpus</title>
        <p>Entity Type
person
location
geographic-feature
plague-ontology-term
date
date-range
time
duration
distance
population/group of people
percent
Entity Mentions
Professor Zabolotny, Professor Kitasato, Dr. Yersin, M. Ha kine
India, Bombay, City of Bombay, San Francisco, Venice
house, hospital, port, store, street
plague, bubo, bacilli, pneumonia, hemorrhages, vomiting
1898, March 1897, 4th February 1897, the beginning of June, next day
1900-1907, July 1898 to March 1899, since September 1896
midnight, noon, 8 a.m., 4:30 p.m.
ten days, months, a week, 48 hours, winter, a long time
20 miles, 100 yards, six miles, 30 feet
Chinese, Europeans, Indian, Russian, Asiatics, coolies, villagers
8%, 25 per cent, ten per cent</p>
        <p>
          Developing a Schematic Structure for the Reports
Whilst our corpus provides a rich source of material the collation of report narratives into
a structured format for information retrieval is not straightforward. The authors approach
the narrative with di erent styles making the application of a schema to support extraction
challenging. As discussed in the Introduction our methodological approach is to treat each
report as a communicative event and we hypothesise that the reports - despite their variation
of styles - as they are intended for the same audience of fellow epidemiologists and government
o cials will present and structure their arguments comparably. This hypothesis is based on
the works of Swales and Bhatia [
          <xref ref-type="bibr" rid="ref18 ref4">18, 4</xref>
          ] who propose that authors achieve their argumentative
structure through steps and moves. Whilst the sequence of arguments may di er in order
of presentation, nonetheless they can identi ed and mapped to a schematic structure. Our
schematic structure is represented by the zoning schema presented in Table 2.
        </p>
        <p>Annotating text with zones creates structure within our reports, allowing us to collate similar
discourse text segments, such as those that discuss treatments or local conditions. Identifying
similar discourse segments allows for targeted analysis which can reveal knowledge and thinking
on these zones and how this knowledge developed over time. Our zoning schema was created
from studying a subsection of reports, section titles and three rounds of pilot annotation on a
subsection of documents.
3.2</p>
        <sec id="sec-3-1-1">
          <title>Automatic Annotation and Text Mining</title>
          <p>
            To process the plague reports, we used the Edinburgh Geoparser [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ], a text mining pipeline
which has been previously applied to other types of historical text [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]. This tool is made
up of a series of processing components. It takes an input raw text and performs standard
text pre-processing on documents in XML format, including tokenisation, sentence detection,
lemmatisation, part-of-speech tagging and chunking as well as named entity recognition. Before
tokenising the text, we also applied a script to repair broken words which were split in the input
text as a result of end-of-line hyphenation [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ].
          </p>
          <p>We adapted the Edinburgh Geoparser by expanding the list of types of entities it recognises
in text, including geographic-feature, plague-ontology-term and population/group etc.8
8Note that our goal was to emulate descriptions used by the report authors at the time, mirroring concepts of
race and ethnicity that were often implicated in the construction of epidemiological arguments. Some examples
of the population/group entities show that these are often derogatory and considered o ensive today. They are
The full list of entity types extracted from the plague reports and examples are presented
in Table 3. Date entity normalisation and geo-resolution are also applied once the manual
annotation (described in the next section) for a document is completed. This is to ensure that
the corrected text mining output is disambiguated, including manual corrections of spelling
mistakes occurring in entity mentions.
3.3</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Manual Annotation</title>
          <p>Manual annotation was necessary for a number of reasons. Whilst some zones could be identi ed
automatically from section titles we found that this was often hampered by spelling errors due to
OCR issues arising from typeface styles and title placements in margins. In addition, depending
on author narrative styles some zones could be found nested within sections with no titles. This
created the need to manually annotate zones. The automatic recognition of named entities (see
Section 3.2) was partially successful but also su ered from spelling errors and OCR issues. In
addition, as more reports are being annotated new entity mentions are identi ed. Thus manual
correction of erroneous and correction of spurious entity mentions was required.</p>
          <p>Manual annotation is conducted using Brat9, a web-based text annotation tool. After the
text was processed automatically as described above, it was converted from XML into Brat
format to be able to correct the text mining output and add zone annotations.10 Figure 2
shows an excerpt of an example report being annotated in Brat. Entities such as date, location
or geographic feature listed in Table 3 can be seen highlighted in the text. The start of an
outbreak history zone is also marked at the beginning of the excerpt.
3.3.1</p>
          <p>Zone Annotation
Zone annotation, as de ned by our schema shown in Figure 2, is applied inclusive of a section
title and can be nested. For example, zones of cases are often found inside treatment or
clinical appearances zones. Footnote zones were added as these often break the ow of the
text and make downstream natural language processing challenging. In addition, we added
Header/Footer markup to be able to exclude headers and footers, e.g. the publisher name or
report name repeated on each page, from further analysis or search.</p>
          <p>Tables were a challenge for the OCR and unusable for the most part. When marking up
tables, we also record its page number. Text within tables is currently ignored when ingesting
the structured data to Solr (see Section 4). However, tables include a lot of valuable statistical
information. In the next phase of the project we will investigate whether this information can
be successfully extracted or whether it will need to be manually collated.
3.3.2</p>
          <p>Entity Annotation
During manual annotation we instruct our annotators to correct any wrongly automated entities
and add those that were missed. Any mis-spellings of entity mentions, mostly caused by the
OCR process, are also corrected in the note eld, as shown in Figure 2. The mis-spellings are
used as part of our text cleaning process. The corrected forms are also used to geo-resolve place
names and normalise dates. These nal two processing steps of the Edinburgh Geoparser are
carried out on each report once it has been manually annotated and converted back to XML.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data Search Interface</title>
      <p>
        Historians and humanities researchers are often faced with the daunting task of manual search
through document collections to nd information pertinent to their research interest.
Additionally, the challenges of working with such text digitally require interdisciplinary collaboration.
HistSearch [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], an on-line tool applied to historical texts, demonstrates how computational
linguists and historians can work together to automate access to information extraction. One goal
of the Plague.TXT project is to make our digital collection available as an on-line search and
retrieval resource but in addition this collection should be accessible. This means being
available to computational linguists for in-depth analysis as well as interactive search for humanities
researchers. We do this with Apache Solr11.
      </p>
      <p>Solr is an open-source enterprise-search platform, widely used for digital collections. The
features available through the Solr search interface make our collection accessible to a wide
audience. For example, it o ers faceting features that support grouping and organising data in
multiple ways, whilst data interrogation can be achieved through its simple interface with term,
query, range and data faceting. Solr also supports rich document handling with text analytic
features and direct access to data in a variety of formats.12</p>
      <p>We are currently customising and improving the ltering of the data for downstream
analysis in Solr. Below we describe on-going ltering steps with Solr and provide examples to
demonstrate a search interface customisation and to show analysis that can be done from data
retrieved via the search interface.</p>
      <p>11https://lucene.apache.org/solr/
12See the Solr website for further description of features.
4.1</p>
      <sec id="sec-4-1">
        <title>Data Preparation and Filtering in Solr</title>
        <p>
          The annotated data is prepared and imported to Solr using Python, with annotations created
both automatically by the Geoparser and manually by the annotators mapped to appropriate
data elds (e.g. date-range entities are mapped to a Date Range eld13), enabling complex
queries across the values expressed in the document text. Additionally, manual spelling
corrections are used to replace the corresponding text in the OCR rendering prior to Solr ingestion,
thus improving the accuracy of language-based queries and further textual analysis. We also
implement lexicon-based entity recognition for entities that have been missed during the
annotation and for additional entity types, e.g animals. Solr allows for storing and searching
by geo-spatial coordinates and we import geo-coordinates associated with entities identi ed
by the Geoparser. Geo-coordinates can be used to support interactive visualisations, as
developed in the Trading and Consequences project [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] which visualises commodities through
their geo-spatial history. In addition, this location information can be used in analysis such as
transmission and spread, e.g. geo-referenced plague outbreak records have been used to show
how major trade routes contributed to the spread of the plague [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Using `Case Zones' we are
currently assessing NLP techniques to extract case information into a more structured format
for direct access to statistical information from hundreds of individual case descriptions.
4.2
        </p>
        <p>Use Case Example: Illustration of Interactive Search
Our search interface facilitates queries across the collection on content and meta-data as well as
queries based on zones or entity types using facets such as date range. We are interested to see
topic discussed in causes zones and if these di er between report time periods. Using Solr we
search for cause zones published during the pandemic 1894-6 comparing these to cause zones in
reports 1904 and beyond. We apply stop-word removal and removal of all non dictionary terms
to the cause zone text. In future, we will make indexed versions of cleaned data in this format
directly accessible from Solr. We use topic modelling (LDA with Gensim Python library18) to
compare the cause zone text at the di erent time points, selecting two topics.</p>
        <p>Results are presented in Table 4. The earlier reports show the discussion centering around
environment aspects with focus on populations, conditions of living and buildings and how this
might cause the spread. The second topic is linked to the concepts at that time period, about
how the diseases may spread through the water system, with studies of ordinance maps of
13https://lucene.apache.org/solr/guide/8_1/working-with-dates.html#date-range-formatting
14http://www.loc.gov/standards/alto/
15https://github.com/mbennett-uoe/whiiif
16https://iiif.io/api/search/1.0/
17http://libraryblogs.is.ed.ac.uk/librarylabs/2019/07/03/introducing-whiiif/
18https://radimrehurek.com/gensim/index.html
sewerage and water ways. Looking at the later reports we now see rats and eas and infection
are more prominent as a discussion topic but also season, temperature and weather form a topic
being discussed as a causal factor.
5</p>
        <p>Discussion, Conclusions and Future Work
In this paper we have presented the work done in the pilot stage of our Plague.TXT project.
The work is the outcome of an interdisciplinary team working together to understand the
nature and complexities of a historical text collection and the needs of the potential di erent
types of users of this collection. A major contribution of this project is the development of an
annotation schema to bring individual reports together as one corpus. This enables streamlined
and e cient linking of knowledge and concepts used in the comprehension of the third plague
pandemic covering the time period of the collection and provides the ability to analyse these
reports as a one coherent corpus. By making this data accessible through the Solr search
interface, we can share it with the research community in ways that cater for the needs of
di erent eld experts.</p>
        <p>
          As a next step we will explore spelling normalisation. Diachronic and synchronic spelling
variance is a known issue in historical documents [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] but in addition the OCR process also
introduces mis-spellings. We will use existing methods for spelling normalisation and fuzzy
string matching capabilities within Solr to correct for spelling variation introduced by OCR.
        </p>
        <p>Manual annotation is time consuming and can be an error prone process. As we increase the
number of reports annotated with zone markup, we also intend to investigate if text similarity
measures can be used for automatic zone identi cation. We are also developing our methods
to directly access the statistical information contained within case zones and within tables.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work was funded by the Challenge Investment Fund 2018-19 from the College of Arts,
Humanities and Social Sciences, University of Edinburgh.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Alex</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Burns</surname>
          </string-name>
          .
          <article-title>Estimating and rating the quality of optically character recognised text</article-title>
          .
          <source>In In Proceedings 1st DATeCH</source>
          , pages
          <volume>97</volume>
          {
          <fpage>102</fpage>
          , New York, NY, USA,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Alex</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Byrne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Grover</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Tobin</surname>
          </string-name>
          .
          <article-title>Adapting the Edinburgh Geoparser for Historical Georeferencing</article-title>
          .
          <source>International Journal for Humanities and Arts Computing</source>
          ,
          <volume>9</volume>
          (
          <issue>1</issue>
          ):
          <volume>15</volume>
          {
          <fpage>35</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Alex</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Grover</surname>
          </string-name>
          , E. Klein, and
          <string-name>
            <given-names>R.</given-names>
            <surname>Tobin</surname>
          </string-name>
          . Digitised Historical Text: Does it have to be mediOCRe?
          <source>In Proceedings of KONVENS 2012 (LThist 2012 workshop)</source>
          , pages
          <fpage>401</fpage>
          {
          <fpage>409</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V. K.</given-names>
            <surname>Bhatia</surname>
          </string-name>
          . Analysing Genre:
          <article-title>Language use in Professional Settings</article-title>
          . New Youk:Routledge,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bramanti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.R.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Walle</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.C.</given-names>
            <surname>Stenseth</surname>
          </string-name>
          .
          <article-title>The third plague pandemic in europe</article-title>
          .
          <source>Proceedings of the Royal Society B: Biological Sciences</source>
          ,
          <volume>286</volume>
          (
          <year>1901</year>
          ):
          <fpage>20182429</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.R.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Krauer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.V.</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <article-title>Epidemiology of a bubonic plague outbreak in glasgow, scotland in 1900</article-title>
          . Royal Society Open Science,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>181695</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Echenberg</surname>
          </string-name>
          .
          <source>Plague Ports: The Global Urban Impact of Bubonic Plague</source>
          ,
          <fpage>1894</fpage>
          -
          <lpage>1901</lpage>
          . New York University Press, New York,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Engelmann</surname>
          </string-name>
          . Mapping Early Epidemiology:
          <article-title>Concepts of Causality in Reports of the Third Plague Pandemic 18941950</article-title>
          . In E. T.
          <article-title>Ewing and</article-title>
          K. Randall, editors,
          <source>Viral Networks: Connecting Digital Humanities and Medical History</source>
          , pages
          <volume>89</volume>
          {
          <fpage>118</fpage>
          . VT Publishing,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>A model-based book dewarping method using text line detection</article-title>
          .
          <source>In Proc. CBDAR</source>
          <year>2007</year>
          , pages
          <fpage>63</fpage>
          {
          <fpage>70</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gotscharek</surname>
          </string-name>
          , U. Re e, C. Ringlstetter,
          <string-name>
            <given-names>K. U.</given-names>
            <surname>Schulz</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Neumann</surname>
          </string-name>
          .
          <article-title>Towards information retrieval on historical document collections: the role of matching procedures and special lexica</article-title>
          .
          <source>IJDAR</source>
          ,
          <volume>14</volume>
          (
          <issue>2</issue>
          ):
          <volume>159</volume>
          {
          <fpage>171</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Grover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tobin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Byrne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Woollard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Reid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dunn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ball</surname>
          </string-name>
          .
          <article-title>Use of the Edinburgh Geoparser for georeferencing digitized historical collections</article-title>
          .
          <source>Philosophical Transactions of the Royal Society A</source>
          ,
          <volume>368</volume>
          (
          <year>1925</year>
          ):
          <volume>3875</volume>
          {
          <fpage>3889</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hauser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Heller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Leiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. U.</given-names>
            <surname>Schulz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Wanzeck</surname>
          </string-name>
          .
          <article-title>Information access to historical documents from the Early New High German period</article-title>
          . In L. Burnard,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dobreva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Fuhr</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A</surname>
          </string-name>
          . Ludeling, editors,
          <source>Digital Historical Corpora- Architecture</source>
          , Annotation, and
          <string-name>
            <surname>Retrieval</surname>
          </string-name>
          , Dagstuhl, Germany,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>U.</given-names>
            <surname>Hinrichs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Alex</surname>
          </string-name>
          , J. Cli ord, A.
          <string-name>
            <surname>Watson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Quigley</surname>
            , E. Klein, and
            <given-names>C.M.</given-names>
          </string-name>
          <string-name>
            <surname>Coates</surname>
          </string-name>
          .
          <article-title>Trading Consequences: A Case Study of Combining Text Mining and Visualization to Facilitate Document Exploration</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          ,
          <volume>30</volume>
          (
          <issue>suppl 1</issue>
          ):i50{
          <fpage>i75</fpage>
          ,
          <year>10 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lopresti</surname>
          </string-name>
          .
          <article-title>Measuring the impact of character recognition errors on downstream text analysis</article-title>
          .
          <source>In B.A. Yanikoglu and K</source>
          . Berkner, editors,
          <source>Document Recognition and Retrieval</source>
          , volume
          <volume>6815</volume>
          . SPIE,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Morabia</surname>
          </string-name>
          , editor.
          <source>A history of epidemiologic methods and concepts</source>
          .
          <source>Birkhauser Verlag</source>
          , Basel ; Boston,
          <year>2004</year>
          . OCLC:
          <volume>55534998</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pettersson</surname>
          </string-name>
          , J. Lindstrom,
          <string-name>
            <given-names>B.</given-names>
            <surname>Jacobsson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Fiebranz</surname>
          </string-name>
          .
          <article-title>Histsearch { implementation and evaluation of a web-based tool for automatic information extraction from historical text</article-title>
          .
          <source>In 3rd HistoInformatics Workshop</source>
          , Krakow, Poland,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Piotrowski</surname>
          </string-name>
          .
          <article-title>Natural language processing for historical texts</article-title>
          .
          <source>Synthesis lectures on human language technologies</source>
          ,
          <volume>5</volume>
          (
          <issue>2</issue>
          ):1{
          <fpage>157</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Swales</surname>
          </string-name>
          .
          <article-title>Aspects of article introductions</article-title>
          .
          <source>Language Studies Unit</source>
          . University of Aston in Birmingham,
          <year>1981</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Yersin</surname>
          </string-name>
          .
          <article-title>La Peste Bubonique a Hong Kong</article-title>
          . Annales de l'
          <source>Institut Pasteur</source>
          , pages
          <volume>662</volume>
          {
          <fpage>667</fpage>
          ,
          <year>1894</year>
          . f667,
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Yue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.F.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.Y.H.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Trade routes and plague transmission in pre-industrial Europe</article-title>
          . Scienti c reports,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>12973</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>