<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The HistCorp Collection of Historical Corpora and Resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eva Pettersson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beata Megyesi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Uppsala University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present the HistCorp collection, a freely available open platform aiming at the distribution of a wide range of historical corpora and other useful resources and tools for researchers and scholars interested in the study of historical texts. The platform contains a monitoring corpus of historical texts from various time periods and genres for 14 European languages. The collection is taken from well-documented historical corpora, and distributed in a uniform, standardised format. The texts are downloadable as plaintext, and in a tokenised format. Furthermore, a subset of the corpus contains information on the modern spelling variant, and some of the texts are also annotated with part-of-speech and syntactic structure. In addition, precon gured n-gram language models and spelling normalisation tools are provided to allow the study of historical languages.</p>
      </abstract>
      <kwd-group>
        <kwd>digital humanities</kwd>
        <kwd>historical corpora</kwd>
        <kwd>language models</kwd>
        <kwd>spelling normalisation</kwd>
        <kwd>HistCorp</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Historical text, and tools for processing historical text, are of great interest
to historians, literary scholars and other researchers in humanities, as well as
for researchers in computational linguistics and digital humanities. However,
corpora containing historical text are not easily found. Furthermore, since there
is no well-established standard for the format of historical corpora, the corpora
that do exist are often available in di erent formats, that are not always well
documented. Thus, it might be time-consuming and hard for the user to extract
the adequate information from these corpora. In addition, copyright issues and
terms of use are often unclear. There are also many cases where the corpus at
hand is only available via a web-based search interface, giving no possibility
to download the actual corpus in a machine-readable format, which might be
desirable for the research topic in question.</p>
      <p>Another useful resource in corpus studies is language models, representing the
probability for certain combinations of words (or letters) to occur in a sequence
in a speci c language (or sublanguage). For example, the sequence he is would
have a signi cantly higher probability to occur in the English language than
the sequence he are. These statistics are computed automatically from corpora
of any kind. Within the domain of historical text, language models could be
created for texts from di erent time periods, providing clues on how language
has changed over time; information that could be of interest to for example
historical linguists.</p>
      <p>Furthermore, searching historical text for particular words and/or structures
is trickier than searching modern text, mainly due to the peculiar and variable
spelling in historical text. Therefore, natural language processing tools especially
designed for handling historical text might be useful as an aid in the information
extraction phase, for example spelling normalisation tools (as described in more
detail in Section 5).</p>
      <p>In this paper, we present the HistCorp collection of historical corpora and
resources: http://stp.lingfil.uu.se/histcorp/. The aim of our work is to
provide a platform for researchers working with historical text, where we gather a
wide range of historical corpora and other useful resources and tools in one place,
available for download in a well-de ned and uniform format. The HistCorp
platform consists of three entries:</p>
    </sec>
    <sec id="sec-2">
      <title>1. Historical Corpora</title>
      <p>Historical corpora are available for download in a plain text format and in
a tokenised format (separated into words and sentences). For some corpora
there are also other formats available for download, containing information
on for example modernised spelling, or morphological and syntactic
annotation. See further Section 3.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Language Models</title>
      <p>The user may download precon gured n-gram language models built on the
corpora available on the HistCorp platform. There is also a possibility for
the user to upload his/her own text les to create a language model based
on these speci c les. See further Section 4.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Tools</title>
      <p>The Tools part of the HistCorp platform provides tools for facilitating the
process of analysing historical text. See further Section 5.
2</p>
      <sec id="sec-4-1">
        <title>The Decode Project</title>
        <p>The HistCorp platform was created within the Decode project1, a project
aiming at the development of computer-aided tools for (semi-)automatic decoding of
enciphered historical manuscripts, so called ciphertexts. Ciphertexts are
important historical records, encrypted to keep the content of the message hidden from
others than the intended receiver(s). Examples of such material include
diplomatic correspondence, intelligence reports, alchemistic and scienti c writings,
1 http://stp.lingfil.uu.se/~bea/decode/
private letters, diaries, and sources related to secret societies. There are
thousands of such historical texts in archives and libraries, waiting to be discovered
and decrypted.</p>
        <p>Within the rather young and highly interdisciplinary eld of historical
cryptology, researchers from areas such as language technology, computer science,
history, linguistics and philology make e orts to develop an infrastructure to
be able to systematically decode ciphertexts. In order to develop algorithms for
(semi-)automatic decryption of historical ciphers, there is a need for all three
aspects of the HistCorp platform. Historical corpora re ecting the language of
the time, as well as language models derived from these resources, are useful for
language identi cation and more well-informed hypotheses and guesses in the
decryption phase. In addition, natural language processing tools based on
methods such as spelling normalisation techniques could be used for mapping several
word forms with di erent spellings to the same normalised form, facilitating the
decoding process, as well as the analysis of the decoded text.
3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Historical Corpora</title>
        <p>The corpus part of the HistCorp platform is to be considered as a monitoring
corpus, intended to grow over time, as more texts and more languages are added.
The aim is to collect diplomatic transcriptions of historical text, that is texts
with minimal or no editorial intervention in the digitization process. Another
quality aspect is that the included texts are taken from well-documented corpus
collections, rather than just sampling historical texts from the Internet. In
addition, the texts included on the HistCorp platform should be free for download.2
Additional terms of use (e.g., regarding redistribution, commercial use etc.) are
clearly stated in the licence information for each corpus. As mentioned above,
the corpus collection is not static, meaning that more corpora will be added over
time. At the time of writing, there are corpora for 14 languages available on the
HistCorp platform, as presented in Table 1.
2 There are a few exceptions to this, but then the procedure for getting access to the
corpus is clearly stated in the readme le for that particular corpus.</p>
        <p>Czech
1) Diakorp: the diachronic section of the Czech National Corpus</p>
        <p>
          <xref ref-type="bibr" rid="ref9">(Kucera and Stluka 2011)</xref>
          2) Gutenberg texts (http://www.gutenberg.org/)
        </p>
        <p>
          Dutch
1) Compilation Corpus Historical Dutch
          <xref ref-type="bibr" rid="ref2">(Cousse 2010)</xref>
          2) Gutenberg texts (http://www.gutenberg.org/)
        </p>
        <p>English
Lampeter Corpus of Early Modern English Tracts
(http://ota.ox.ac.uk/desc/3193)</p>
        <p>French
Paris Speech in the Past (http://ota.ox.ac.uk/desc/2423)</p>
        <p>
          German
1) Deutsches TextArchiv
          <xref ref-type="bibr" rid="ref7">(Geyken et al. 2010)</xref>
          2) GerManC
          <xref ref-type="bibr" rid="ref4">(Durrell et al. 2007)</xref>
          3) Reference Corpus of Middle High German (Klein and Dipper 2016)
        </p>
        <p>Greek
Ancient Greek Dependency Treebank</p>
        <sec id="sec-4-2-1">
          <title>Corpus Format</title>
          <p>Since the corpora on the HistCorp platform are collected from many di erent
sources, the source format and annotation level di er between the corpora. First,
the text encoding for representing language speci c characters is not the same for
all source corpora. In addition, some corpora are only available in a plain text
format, whereas others are annotated with for example morphological and/or
syntactic information. For annotated corpora, the linguistic information may be
represented in a tab-separated format, in an XML-based format, in a
parenthesisbased format etc. Apart from the representation of the text content and the
linguistic annotation, metadata (such as time period, author, genre etc.) may
also be given in many shifting formats.</p>
          <p>One aim with the HistCorp platform is to provide corpora in a uniform
and well-documented format. Therefore, every corpus is converted to the
Unicode UTF-8 encoding scheme3, and standardised into a plain text format with
metadata stated in a TEI-compatible format at the top of each le. Furthermore,
each corpus is also provided in a uniform tokenised format, with one token on
each line, and blank lines separating each sentence. Since the metadata
information, as well as the plain text format and the tokenisation format, are the
same for all corpora, it will be possible for the user to apply the same tools
for metadata extraction and linguistic processing of all corpora. Regarding
morphological and syntactic annotation, it is a much trickier task to standardise
this information over all corpora. Thus, the source format for representing these
features are currently kept unchanged.</p>
          <p>
            All in all, each corpus on the HistCorp platform is possibly available in ve
di erent download formats:
1. A plain text format, with metadata stated at the top of each le, in a
TEIcompatible format (see further Section 3.3). For corpora that were originally
only available in an XML format, the text elements are automatically
extracted from the XML le, to create a plain text le.
2. A tokenised format, with one token on each line, and a blank line separating
sentences, as illustrated in the following example taken from the Gender and
Work corpus
            <xref ref-type="bibr" rid="ref18">(Agren et al. 2011)</xref>
            :
          </p>
          <p>
            Asen
stambd
till
tingz
For corpora originally lacking tokenisation, UDPipe
            <xref ref-type="bibr" rid="ref16 ref17">(Straka and Strakova
2017)</xref>
            is used with the CoNLL 2017 Shared Task baseline models
            <xref ref-type="bibr" rid="ref16 ref17">(Straka
2017)</xref>
            , to create the tokenised version of the corpus.
3. A format containing information on historical and modern spelling of the
text. This applies to corpora where all or some of the words are available
both in their original spelling, and in a manually modernised spelling. This
kind of annotation is presented in a format with one token pair on each line,
with a tab separating the historical spelling from the modernised spelling,
as illustrated in the following example taken from the the Gender and Work
corpus
            <xref ref-type="bibr" rid="ref18">(Agren et al. 2011)</xref>
            :
4. A morphologically annotated format. This applies only to corpora where
morphological information is available in the original source corpus. The
original annotation format is then kept in the HistCorp download as well,
as illustrated in the following example taken from the Tycho Brahe Parsed
Corpus of Historical Portuguese
            <xref ref-type="bibr" rid="ref6">(Galves and Faria 2010)</xref>
            :
O/D amor/N puro/ADJ ,/, bel ssima/ADJ-S-F Genoveva/NPR ,/,
e/SR-P muito/Q raro/ADJ ./.
5. A syntactically annotated format. This applies only to corpora where
syntactic information is available in the original source corpus. The original
annotation format is then kept in the HistCorp download as well, as
illustrated in the following example taken from the Tycho Brahe Parsed Corpus
of Historical Portuguese
            <xref ref-type="bibr" rid="ref6">(Galves and Faria 2010)</xref>
            :
((IP-MAT (NP-SBJ (D O)
(N amor)
(ADJP (ADJ puro)))
(, ,)
(NP-VOC (ADJ-S-F bel ssima) (NPR Genoveva))
(, ,)
(SR-P e)
(ADJP (Q muito) (ADJ raro))
(. .))
3.2
          </p>
        </sec>
        <sec id="sec-4-2-2">
          <title>Presentation Format</title>
          <p>On the HistCorp platform, each corpus is presented in 12 columns, as
illustrated in Figure 1. The rst column states the name of the corpus, whereas
the second and third column gives information on the time period and genres
contained in the corpus. Furthermore, there are ve download columns,
corresponding to the di erent annotation levels described in Section 3.1, that is: a)
plain text format, b) tokenised format, c) spelling normalisation, d)
morphological annotation, and e) syntactic annotation. The ninth column contains a link
to a short readme le for the corpus, with information on the contents and
format of the corpus. The tenth column provides the possibility to download all
the corpora les in one go, instead of downloading only the plain text format or
the morphologically annotated le etc. The eleventh column contains a link to
the source page, from which the corpus was originally downloaded, whereas the
last column states the terms of use for each corpus, typically in the form of a
link to the speci c licence adhering to the corpus in question.</p>
          <p>In the speci c example of the German language, illustrated in Figure 1, it could
be noted that the Deutsches TextArchiv corpus contains texts from the time
period 1600{1899, and that the texts are available in a plain text format and in
a tokenised format. The GerManC corpus on the other hand contains texts from
the time period 1654{1799, and is available on all annotation levels, except for
the spelling normalisation level. The third corpus, Reference Corpus of Middle
High German, contains older texts (1050{1350) and is available in a plain text
format, in a tokenised format, and in a morphologically analysed format.
3.3</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>Metadata</title>
          <p>
            Many corpora contain extra-textual information (in the following referred to as
metadata), such as the title of the text, the name of the author, the year in
which the text was written and so on. This is very useful information, but the
way this information is structured often di ers between di erent source corpora,
making it hard for the user to extract the relevant information. Therefore, in the
HistCorp les, the metadata information has been converted to a consistent
format that is identical for all corpora, regardless of the original source corpus
format. The metadata information is stored at the top of each plain text le, with
a hash sign (#) preceding each new piece of metadata information, as illustrated
in Figure 2, showing the metadata information available for one of the texts in
the diachronic section of the Czech National Corpus
            <xref ref-type="bibr" rid="ref9">(Kucera and Stluka 2011)</xref>
            .
Some metadata are extracted from the metadata information given in the source
corpora, whereas some metadata have been added during the creation of the
HistCorp les (for example the number of tokens).
#title: Obrazy ze zivota meho, Marinka
#author: Karel Hynek Macha
#distributor: Distributed within the Diakorp project, see further the
project webpage at: https://wiki.korpus.cz/doku.php/en:cnk:diakorp.
#availability: The data are licenced under the CC BY-NC-SA license,
http://creativecommons.org/licenses/by-nc-sa/4.0/.
#sourceDesc: Part of the diachronic section of the Czech National Corpus.
#extent tokens: 5,253
#extent documents: 1
#normalization: diplomatic
#language: Czech
#date: 1834--1835
#domain: prose
          </p>
          <p>To make the metadata information compatible with metadata contained in other
corpora, we use the Text Encoding Initiative (TEI) standard for naming the
metadata elements4, but in the simpli ed format illustrated in Figure 2 instead
4 http://www.tei-c.org/
of the full XML format usually associated with the TEI standard. This way, the
metadata information is more comprehensible for the human user, while at the
same time o ering a format that is straightforward to convert to the traditional
TEI XML format if needed. The most common metadata elements occurring in
the HistCorp les are:
{ author</p>
          <p>Name of the author of the text.
{ availability</p>
          <p>Terms of use for the corpus, typically with a link to the actual licence.
{ date</p>
          <p>The year or time period during which the text was written.
{ distributor</p>
          <p>Information on the person or organisation providing the source corpus.
{ domain</p>
          <p>Information on the genre of the text.
{ extent</p>
          <p>Divided into extent documents, for the number of subtexts within the
text, and extent tokens, for the number of tokens (words and
punctuations) in the text. The number of tokens are calculated using the Unix
command wc -w.
{ language</p>
          <p>The language in which the text is written.
{ sourceDesc</p>
          <p>A short description of the contents of the text.
{ title</p>
          <p>The title of the text, possibly divided into title main for the main title,
and title sub for the subheading.
4</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Language Models</title>
        <p>On the Language Models part of the HistCorp platform, the user may
download precon gured n-gram language models based on the corpora in the
HistCorp collection of texts. The language models contain statistical
information on the occurrence of sequences of words and characters in a language or a
sublanguage, and are useful for many purposes, such as language identi cation
and research on language change.</p>
        <p>
          The language models provided on the HistCorp platform were created using
the irstlm toolkit5
          <xref ref-type="bibr" rid="ref5">(Federico et al. 2008)</xref>
          with n-gram size 5 for
characterbased language models, and n-gram size 3 for word-based language models. The
HistCorp language models are typically divided into centuries, as illustrated
in Figure 3, showing the download format for the Swedish language models.
5 https://github.com/irstlm-team/irstlm
        </p>
        <p>As seen from the Figure, the user may choose to download a word-based or a
character-based language model for the time period of interest. The table also
provides information on the number of words or characters covered by each
language model. The last column of the download table contains a link to a
short readme le, describing the contents of the language model, and how it was
created. Below the table is a link for downloading a language model containing all
the texts for the language in question, instead of downloading separate language
models for each time period.</p>
        <p>Apart from downloading precon gured language models, the user may also
upload one or more text les to the HistCorp platform, to generate his/her
own language model, as illustrated in Figure 4.</p>
        <p>This gives the user the possibility to upload les not contained in the
HistCorp collection of texts, and to de ne his/her own criteria for time periods
and genres included in the resulting language model. The default values for how
many sequential units to include in the language models are three words for the
word-based models, and ve characters for the character-based model, but the
user has the possibility to choose a value between one and ten for both types of
language models.
5</p>
      </sec>
      <sec id="sec-4-4">
        <title>Tools</title>
        <p>
          The Tools section of the HistCorp platform provides the possibility to
download tools especially designed for processing historical text. Here, the HistNorm
package for automatic spelling normalisation of historical text is available for
download
          <xref ref-type="bibr" rid="ref10">(Pettersson et al. 2014)</xref>
          . Spelling normalisation is the process of
automatic modernisation of the spelling in historical text, a method that could be
used for several purposes. One is to facilitate search in historical text, since the
user then can search for the standardised spelling of a word, and nd all
occurrences of that word form in the text, regardless of spelling variation. Searching
for the same word form in the text in its original spelling on the other hand,
would mean that the user would have to guess what di erent spelling variants to
enter for that particular word form. Another bene t of spelling normalisation,
is that natural language processing (NLP) tools such as taggers for
morphological analysis and parsers for syntactic analysis are generally trained on modern
language texts, and do not perform well on texts with the variable historical
spelling. However, studies have shown that spelling normalisation of historical
text prior to the application of taggers and parsers trained on modern language
data signi cantly improves the performance of the NLP tools
          <xref ref-type="bibr" rid="ref11">(Pettersson 2016)</xref>
          .
        </p>
        <p>In the Tools section, the user also has the possibility to download
predened datasets suitable for training spelling normalisation systems for di erent
languages, as illustrated in Figure 5. These datasets contain both a set of
historical spellings mapped to their corresponding modern spellings, and references
to modern language resources that could also be useful in training a spelling
normalisation tool (depending on the normalisation method chosen).</p>
        <p>We plan to add more tools for spelling normalisation, as well as other useful
tools for processing historical text, in the near future.
6</p>
      </sec>
      <sec id="sec-4-5">
        <title>Conclusion</title>
        <p>In this paper, we have presented the HistCorp collection consisting of
historical corpora and resources for 14 languages from various time periods, as a
freely available open-access resource. The historical corpora are taken from
welldocumented corpus collections with minimal or no editorial intervention. The
corpora for various languages are provided in a uniform and well-documented
format, converted to the Unicode UTF-8 encoding scheme, and standardised into
plaintext. TEI-compatible metadata is used for describing extra-textual
information about the texts.</p>
        <p>Each corpus is presented in several downloadable formats. Apart from
plaintext, the texts are also available in a tokenised format. In addition, where
applicable, tokens are annotated with their modern spelling as well as part-of-speech
and morphological features, and the sentences are marked-up with syntactic
dependencies.</p>
        <p>The platform also provides several precon gured n-gram language models
generated from the various texts in the collection. The language models contain
statistical information on co-occurrence sequences of characters and words in a
particular dataset, which might be useful for historical linguistic studies of
language change. The user can also create his/her own language model by uploading
texts of their interest.</p>
        <p>In addition, HistCorp contains downloadable tools for the automatic
processing of historical texts including spelling normalisation to create standardised
spelling across texts to facilitate search, and prede ned datasets for training
spelling normalisation systems for various languages. In the near future, we plan
to add more historical datasets and tools to the platform.</p>
        <p>XIV</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Buchholz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Marsi</surname>
          </string-name>
          , E.:
          <article-title>CoNLL-X shared task on Multilingual Dependency Parsing</article-title>
          .
          <source>Proceedings of the 10th Conference on Computational Natural Language Learning</source>
          (CoNLL-X)
          <volume>149</volume>
          {
          <fpage>164</fpage>
          ,
          <string-name>
            <surname>Association for Computational Linguistic</surname>
          </string-name>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Cousse</surname>
          </string-name>
          , Evie:
          <article-title>Een digitaal compilatiecorpus historisch Nederlands</article-title>
          .
          <source>Lexikos</source>
          <volume>20</volume>
          :
          <fpage>123</fpage>
          {
          <fpage>142</fpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Delsing</surname>
          </string-name>
          , Lars-Olof:
          <article-title>Forsvenska textbanken</article-title>
          . Lagman,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Olsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. O.</given-names>
            and
            <surname>Voodla</surname>
          </string-name>
          , V. (ed.):
          <source>Nordistica Tartuensia 7</source>
          <volume>149</volume>
          {
          <issue>156</issue>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Durrell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ensslin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Bennett</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The GerManC project</article-title>
          .
          <source>Sprache und Datenverarbeitung</source>
          <volume>31</volume>
          :
          <fpage>71</fpage>
          {
          <fpage>80</fpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Federico</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bertoldi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Cettolo</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>IRSTLM: an open source toolkit for handling large scale language models</article-title>
          .
          <source>Proceedings of Interspeech</source>
          <volume>1618</volume>
          {
          <fpage>1621</fpage>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Galves</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Faria</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Tycho Brahe Parsed Corpus of Historical Portuguese (</article-title>
          <year>2010</year>
          ). url: http://www.tycho.iel.unicamp.br/ tycho/corpus/en/index.html.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Geyken</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haaf</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurish</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinmann</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Wiegand</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Das Deutsche Textarchiv: Vom historischen Korpus zum aktiven Archiv</article-title>
          .
          <source>Digitale Wissenschaft</source>
          <volume>157</volume>
          {
          <fpage>161</fpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Klein</surname>
            <given-names>T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Dipper</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <source>Handbuch zum Referenzkorpus Mittelhochdeutsch Bochumer Linguistische Arbeitsberichte</source>
          <volume>19</volume>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Kucera</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Stluka</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>DIAKORP: Diachronn korpus, version 5</article-title>
          .
          <article-title>Ustav Ceskeho narodn ho korpusu FF UK</article-title>
          ,
          <string-name>
            <surname>Praha</surname>
          </string-name>
          (
          <year>2011</year>
          ). url: http://www.korpus.cz.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Pettersson</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Megyesi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Nivre</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A Multilingual Evaluation of Three Spelling Normalisation Methods for Historical Text</article-title>
          .
          <source>Proceedings of the 8th Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities (LaTeCH)</source>
          <volume>32</volume>
          {
          <fpage>41</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Pettersson</surname>
          </string-name>
          , Eva.:
          <article-title>Spelling Normalisation and Linguistic Analysis of Historical Text for Information Extraction</article-title>
          .
          <source>Doctoral thesis</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Ro</surname>
          </string-name>
          gnvaldsson, E.,
          <string-name>
            <surname>Ingason</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sigurdsson</surname>
            ,
            <given-names>E. F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Wallenberg</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The Icelandic Parsed Historical Corpus (IcePaHC)</article-title>
          .
          <source>Proceedings of the Eighth International Conference on Language, Resources and Evaluation (LREC)</source>
          <year>1977</year>
          {
          <year>1984</year>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Sallander</surname>
          </string-name>
          , Hans:
          <article-title>Akademiska konsistoriets protokoll</article-title>
          .
          <source>Acta Universitatis Upsaliensis</source>
          ,
          <string-name>
            <surname>Uppsala</surname>
          </string-name>
          (
          <year>1968</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Sanchez-Mart nez</surname>
          </string-name>
          , F.,
          <string-name>
            <surname>Mart</surname>
            nez-Sempere,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ivars-Ribes</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Carrasco</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          :
          <article-title>An open diachronic corpus of historical Spanish</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>47</volume>
          (
          <issue>4</issue>
          ):
          <volume>1327</volume>
          {
          <fpage>1342</fpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Scherrer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Erjavec</surname>
          </string-name>
          , T.:
          <article-title>Modernising historical Slovene words</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>22</volume>
          (
          <issue>6</issue>
          ):
          <volume>881</volume>
          |
          <fpage>905</fpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Straka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Strakova</surname>
          </string-name>
          , J.: Tokenizing, POS Tagging,
          <article-title>Lemmatizing and Parsing UD 2.0 with UDPipe</article-title>
          .
          <source>Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies</source>
          <volume>88</volume>
          {
          <fpage>99</fpage>
          ,
          <string-name>
            <surname>Association for Computational Linguistics</surname>
          </string-name>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Straka</surname>
          </string-name>
          ,
          <string-name>
            <surname>Milan: CoNLL 2017 Shared Task - UDPipe Baseline</surname>
          </string-name>
          Models and
          <article-title>Supplementary Materials LINDAT/CLARIN digital library at the Institute of Formal and Applied Linguistics</article-title>
          , Charles University (
          <year>2017</year>
          ). url: http://hdl.handle.net/11234/1- 1990.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Agren</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fiebranz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindberg</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          and Lindstrom, J.: Making Verbs Count.
          <article-title>The research project `Gender and Work' and its methodology</article-title>
          .
          <source>Scandinavian Economic History Review</source>
          <volume>59</volume>
          (
          <issue>3</issue>
          ):
          <volume>271</volume>
          |
          <fpage>291</fpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>