<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Literary Studies Meet Corpus Linguistics: Estonian Pilot Pro ject of Private Letters in KORP?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marin Laak</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kaarel Veskis</string-name>
          <email>kaarel.veskisg@kirmus.ee</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olga Gerassimenko</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Neeme Kahusk</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kadri Vider</string-name>
          <email>kadri.viderg@ut.ee</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>0134) under the activity \Support for Research Infrastructures of National Importance</institution>
          ,
          <addr-line>Roadmap"</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Estonian Literary Museum Vanemuise 42</institution>
          ,
          <addr-line>51003 Tartu</addr-line>
          ,
          <country country="EE">Estonia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Computer Science, University of Tartu</institution>
          ,
          <addr-line>J. Liivi 2, 50409 Tartu</addr-line>
          ,
          <country country="EE">Estonia</country>
        </aff>
      </contrib-group>
      <fpage>283</fpage>
      <lpage>294</lpage>
      <abstract>
        <p>Digitalisation of cultural heritage in Estonia has been in progress during recent years, and we will see expansive mass digitalisation of printed books and handwritten documents in the very near future. This situation and potential actualises questions of the usage of the literary heritage. In our paper, we consider bene ts for digital literary research that arise from representing a literary text collection as an annotated language resource. We discuss the pilot project of creating a text corpus based on private letters between two Estonian avantgarde writers in the beginning of the 20th century. The advantages and possibilities of corpus query system KORP that we have chosen for representing and searching literary heritage DH corpora as a language resource are described. Challenges that the application of Natural Language Processing and Text and Data Mining imposes on the preparation and representation of texts are discussed along with bene ts for the research.</p>
      </abstract>
      <kwd-group>
        <kwd>literary studies</kwd>
        <kwd>private letters</kwd>
        <kwd>digital heritage collections corpus linguistics</kwd>
        <kwd>natural language processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Application of Natural Language Processing (NLP) and Text and Data Mining
(TDM) methods and tools is gradually increasing and becoming more relevant in
DH research in Estonia. Similarly to the Nordic and Baltic countries, Estonia can
expect an explosive growth in digital heritage and text resources: preparations
for massive digitisation of cultural heritage started as a national programme
in 2019. The creation of these new digital resources will be the priority for all
Estonian leading memory institutions and the scope of this huge project includes
di erent types of cultural heritage (printed books, archival documents, photo and
lm heritage, ethnographic and ne art objects) [
        <xref ref-type="bibr" rid="ref10 ref6">6,10</xref>
        ] Digitised resources will
be made accessible on the internet as open data, but it is still an open question
how and for what purpose can such digital resources be used [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        What are the new challenges and the new knowledge o ered by using the
tools of NLP in analysing the digital literary heritage? Our interdisciplinary
project \Literary Studies Meet Corpus Linguistics" concentrates on using
cultural heritage, especially literary history sources in research with the help of the
NLP&amp;TDM methods. The primary question is how to bridge the gap between
the research possibilities o ered by the contemporary NLP&amp;TDM and the ever
increasing amount of digitised texts and other digital data, produced by memory
institutions. This has proved to be a complicated task and international
practice has shown that literary scholars are slower to embrace new practices than
linguists for whom corpus-based research is already a professional standard. In
general, literary scholars are used to working with texts using traditional
methods, analysing them as undivided poetical and semantic entities [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Digitising
of texts often preserves them as undivided and unannotated entities with
occasionally added metadata - as a result, the digitised materials are human-readable
rather than machine-readable.
      </p>
      <p>
        NLP&amp;TDM methods require treating literary works or texts as data, which
can be analysed and processed with computer programmes. To process data for
statistics and for nding trends, a literary scholar needs either to do a close
reading of literary texts as huge data amounts, or to develop new professional skills
in data analysis methods. This imposes rethinking the approach to empirical
object in literary studies in general and posing new and di erent research
questions. The possibility to compare text strategies, rhetorical and stylistic patterns
in literary, religious and political text corpora might give us new insights into
intertwining ideology, rhetoric and identity presentations. One of the important
trends of automated textual analysis in DH is Sentiment Analysis which, for
instance, allows to measure emotion in parliamentary debates [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Private letters as a literary source</title>
      <p>The empirical foundation of our pilot project of private letters in KORP is
based on the archival ego-documents. This is a rare collection of handwritten
private letters: the correspondence between two Estonian writers Johannes Semper
(1892{1970) and Johannes Barbarus (1890{1946). The authors were well-known
avant-garde writers in Estonian literature, best friends and school classmates.
They shared the same intellectual attitudes, views, and values. For example, they
were Francophiles and during several decades they both reviewed and translated
French literature into Estonian.</p>
      <p>During the years covered by the correspondence, in 1911{1939, Barbarus
published the total of 17 collections of poetry, and Semper published four
collections of poetry and four books of short stories and novels. The letters were
written mainly at di erent locations in Estonia, and also during travels: they
both travelled extensively in Europe (and other parts of the world) and in their
letters they described to each other the sights they saw as well as the literary
events at home. They also wrote about their work, everyday life, health and
prescriptions, hobbies (hunting, swimming, skating, etc.), and guests.</p>
      <p>The temporal and contextual frames of this correspondence are the period in
Europe between two World Wars and, according to the French philosopher Pierre
Bourdieu, the creation of the Estonian cultural eld in the Estonian Republic
in years 1918{1940. As such, the correspondence o ers semantically rich and
multidimensional content for literary scholars.</p>
      <p>
        Generally, we have been inspired by the question how this kind of sources
(private letters, correspondences, manuscripts, etc.) in di erent kinds of
machinereadable formats (pdf, docx, txt etc.) can be used for creating new knowledge
in the interdisciplinary eld of literary studies, including biographies and life
writing studies [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The private correspondence of two Estonian writers during 28 years, as
empirical material, is extremely rich in the themes and possibilities to pose the
traditional as well as the new research questions: the factors of subjectivity and
emotionality, verbal and poetic creativeness of the both authors, etc.</p>
      <p>The original letters are held at di erent archives: the letters of Barbarus to
Semper belong to the Estonian Cultural History Archives of the Estonian
Literary Museum and the letters of Semper to Barbarus belong to the Estonian
National Archives due to the unique literary and historical value of this
correspondence. Their correspondence consists of 670 letters with, all in all, more
than 1,100 pages and more than 310 000 tokens or about 250 000 words. The
number of tokens and words is approximate, as we will discuss later.
3</p>
    </sec>
    <sec id="sec-3">
      <title>KORP as a tool for corpus linguistics</title>
      <p>
        KORP is a corpus query analysis system that allows to nd concordances and
build various statistics from di erently annotated corpora using text metadata
(author, year of publishing, text type etc) and linguistic annotation (splitting
into sentences and words, punctuation, morphology, syntax and semantics).
Technically, KORP is a frontend tool that uses much of IMS Open Corpus
Workbench [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] as backend. It was created and maintained at Swedish Language Bank
Sprakbanken [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and is being developed and customised by language technology
networks in several countries: Sweden3 , Finland4, Norway5, Estonia6, Denmark7
and Island8.
      </p>
      <p>Estonian KORP is hosted by the Centre of Estonian Language Resources9.
The Estonian data available in KORP currently consists of more than 850 million
tokens. As Estonian is a highly in ective language with a rich morphology and
free word order, almost all texts are automatically annotated and disambiguated
on the morphological level to enable form- and baseform-based search.</p>
      <p>As a pilot project for the needs of literary studies in addition to the
corpora designed for linguistic research, we have created Correspondence Corpus
of private letters (Semper and Barbarus, 1911{1940). The original handwritten
manuscripts where already typed in. They were to transform to machine-readable
text les by manually adding metadata and automatically analysing and
disambiguating morphological categories.</p>
      <p>KORP as a corpus query system was chosen because of its open source value,
search exibility and easiness, graphical overview of search results in subcorpora,
easy switching between concordances and broader context (as well as between
frequency statistics and concordances), broad possibilities to group absolute
frequencies of tokens or occurrences and automatically calculated relative
frequencies (occurrences per million tokens). Search results and statistics from KORP
are exportable into Comma Separated Values (CSV) and JSon format.</p>
      <p>Text fragments cited in KORP are limited to a sentence or a passage, so that
KORP does not infringe copyrights. Metadata cited in the query results allow
exact pinpointing of the source of the sentence, and it is possible to link whole
texts hosted elsewhere to the search metadata for close reading of a broader
context when needed.</p>
      <p>KORP helps to analyse and objectively verify or refute di erent research
arguments, for example, such as presented by Estonian literary scholars in the
1980s: (1) The correspondence of Semper and Barbarus is subjective and
emotional, the letters reveal the character and state of mind of the authors at the
moment of writing them. (2) The letters demonstrate the authors awareness of
topical problems in Estonia and in Europe. (3) The subject range of the letters
includes everyday life, health, hobbies, visits and visitors, literary work, books
and reading, and the literary, economic and political life in Estonia and Europe.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Challenges for annotation in manuscript corpus</title>
      <p>Which distinctive features of older correspondences need to be considered in
preparing them to be used as a corpus? Several steps are necessary to convert the
3 https://spraakbanken.gu.se/korp/
4 http://korp.csc.
5 http://gtweb.uit.no/korp/
6 http://korp.keeleressursid.ee
7 http://alf.hum.ku.dk/korp/
8 http://malheildir.arnastofnun.is/
9 https://www.keeleressursid.ee
digitised text collection into a morphologically analysed corpus, and to provide
it with essential metadata about the sources.
1. The rst step in digitising di erent kinds of manuscripts, incl. private
letters, is retyping the handwritten originals in order to transform them into
machine-readable format. Digitised typed texts can be subjected to the
character recognition (OCR).
2. In order to enable statistics and to nd di erent linguistic phenomena in
the texts, it is important to tokenise at least sentences and words, and to
preserve meaningful units (chapters, articles, verses, letters).
3. The process of annotation was carried out semi-automatically. At the rst
stage, sentence boundaries were determined, then words were tokenised.
Tokenised words were lemmatised and tagged by part-of-speech and other
grammatical features.</p>
      <p>
        Di erent levels of annotation make it possible to search for di erent
linguistic or other phenomena. In a tokenised text we can search by word forms (or
their parts), but not by lemmas. In a lemmatised text we can search also by
lemmas, but not by morphological characteristics (e.g., the case, the number).
When the text has been morphologically analysed and disambiguated, we can
search also by morphological characteristics and parts of speech. The automatic
morphological analysis and automatic disambiguation that enable searching for
di erent morphological forms of a word based on the lemma, seriously improve
the usability and quality of the corpus. The accuracy of automatic
morphological disambiguation in the Estonian standard language corpora reaches up to
93{98% [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The accuracy for these texts is probably much lower.
      </p>
      <p>Meaningful units, e.g., temporal expressions or named entities, found in the
text can also be used in searches. If all the meanings of words were tagged, it
would be possible to search for a speci c meaning of a word and discard all other
meanings.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Annotation types speci c for private letters</title>
      <p>The private letters of writers, the correspondence as whole was not addressed
for a wider audience. Creators of the corpus met serious di culties because
often, both correspondents were very familiar with the subjects mentioned in
the letters, so that sometimes they only hinted at the events or persons and
used plenty of abbreviations known only to themselves. Authors of letters re ect
both the historical period and the personal context and idiolects. Being
avantgarde poets, Semper &amp; Barbarus both experimented with the language a lot.</p>
      <p>In the following sections we are going to present the distinctive features of
the texts that would be of great interest for both linguists and literary scholars.
See Fig. 1 for a passage from a letter by Barbarus to Semper with annotations
and English translation.</p>
      <sec id="sec-5-1">
        <title>Abbreviations</title>
        <p>In this corpus, there are a lot of nonstandard abbreviations that were
comprehensible to the authors of the letters. The abbreviations pose a challenge to
the morphological analysis and disambiguation of abbreviations itself, but also
of the neighbouring words, the disambiguated analysis of which often relies on
the context (for instance, \is. Linde" should be analysed as \isand (i.e master)
Linde" and not as a sentence ending and a starting of a new sentence although
there is no common abbreviation \is." in standard Estonian).
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Date and other temporal expressions</title>
        <p>
          The date format of the manuscript letters varies greatly: \11. jaanuar 1927",
\19/XII.37", \2. 2. 1934", \J~oulu 3. puhal 1934" (on the 3rd day of Christmas
1934). The date in the format chosen by the letter authors needs to be saved
as a part of the text, but in addition to that, all the date formats need to be
normalised so that they could be present in the letter metadata in a standardised
format (i.e., YYYY-MM-DD) whenever possible. When it is not possible, we can
use the category \unde ned" for the analysis transparency. Temporal expressions
in the text of letters can be automatically recognised with the help of EstNLTK
library for Python [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and a software tool developed by Siim Orasmaa [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
5.3
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>Named Entity Recognition</title>
        <p>Special interest for the literary scholars lies in the proper names used in the
letters: their usage combined with the metadata (time period, author) might
lead to recognising the important patterns and tendencies.</p>
        <p>Named Entity Recognition (NER) annotations can be produced with the
EstNLTK toolkit and include the types of entities: person (PER), organisation
(ORG), and location (LOC).
5.4</p>
      </sec>
      <sec id="sec-5-4">
        <title>Detect other languages</title>
        <p>Semper &amp; Barbarus were both promoters of French literature and culture in
Estonia. They both worked as translators and travelled a lot. They scattered
their letters with words, phrases, sentences and longer citations in many other
languages than Estonian that were familiar to them (Latin, French, Russian,
German). This is a challenge for the automatic language recognition and annotation
tools which are currently oriented only to the analysis of standard Estonian.
These tools do not help much in case of excerpts from other languages, but if
such code-switching pieces were precisely recognised, lemmatised and analysed,
that information would be useful both to linguists and literary scholars.
\ Bifur "i kiri saabus Sinu kirjaga uhel paeval . Palutakse kohe luule ule
artikkel ara saata ja kedagi paluda, kes teisi kusimusi ( sur la vie en general )
kasitaks. Neil olla to~lkija, nii siis vo~ivat eesti keeleski kirjutada. Et mul
artikkel juba valmis oli, siis saatsin ta tana minema. Kui Sul lusti
midagi saata, siis lakita kohe , | ehk novelli to~lge (maksavad 50 fr.
lehekuljest), ehk siis mahutavad neljandamasse nr -isse; ehk viskad proosa
&amp; teaatri ulegi artikli. Aadress : \ Bifur " , Editions du Carrefour ,
199 , boul . St.-Germain , Paris
(VIe)
( M{eur le redacteur en chef
Ribemont Dessaignes ). Kusisin kirjas, kas nende to~lkija luuletisi to~lkida
vo~iks, siis vo~iksime valiku teha, ehk Suitsi ilmuva antoloogia neile saata.
\ Bifur " harrastab kull rohkem proosat &amp; informatsioonilaadilisi ulevaateid,
nii siis vaevalt nad luule liimile lahevad, aga eks ole ju veel teisi zurnaale paale
\ Bifur "i, kus avaldada saaks, kui aga to~lkija leiduks. Mis teeb see E.K.L.
propagandakomitee? Kas peab viimaks nende liikmete seas propagandeerima
hakkama? Visnapuu kirjutab \ V.-M aas" Igori potserduse puhul, et vaja
propagandeerida, aga, kui ma ei eksi, oli ta ise selles komitees?
Legend:
green</p>
        <p>| Named Entity: Person, Organisation, Title
yellow</p>
        <p>| Named Entity: Location, Address
blue
orange
| Abbreviations</p>
        <p>| other languages than Estonian
magenta</p>
        <p>| temporal expression
\The letter from Bifur arrived on the same day as yours. They asked to send
immediately the article about poetry, and to nd somebody who could treat
other questions (sur la vie en general). They are supposed to have a translator,
thus it could be written in Estonian. Since I had already nished my article, I
sent it to them today. If you think youd like to send something, do it at once,
| maybe a translation of some short story (they pay 50 francs a page), and
perhaps they can t it into the fourth issue; or maybe you will even rush an
article about prose and theatre. Address: Bifur, Editions du Carrefour, 199, boul.
St.-Germain, Paris (VIe) (M|eur le redacteur en chef Ribemont Dessaignes).
I asked them in my letter whether their translator can translate poems, in
this case we could select something, or we can send them the anthology by
Suits, which will come out soon. Bifur is more into prose and information-like
reviews, so it's unlikely that they can be cajoled into accepting poetry, but
there are other magazines besides Bifur where we could publish, if only we
found a translator. What is the E.W.U.'s [Estonian Writers' Union]
propaganda committee up to now? Should we by any chance start propagandising
among their members? Visnapuu writes in the V.-Maa about Igor's handiwork
that it is necessary to propagandise, but if I'm not mistaken, wasn't he on the
committee himself?"
Fig. 1. Excerpt from letter (from Barbarus to Semper, Nov 8., 1929) with di erent
annotation types. See Table 1 for metadata associated to the letter the passage comes
from.</p>
        <p>Text Attribute Value
author Barbarus
recipient Semper
catalogue no. 333
original date 8. november 1929
category [empty]
date 1929-11-08
year 1929
location Parnu
notes [empty]</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Using private letters as text corpus</title>
      <p>The creation of the text corpus of the Semper &amp; Barbarus correspondence had
two wider objectives.</p>
      <p>From the perspective of corpus linguistics, to study the e ect of the
distinctive characteristics of private letters in creating a text corpus: what can and
what cannot be achieved? What di culties may arise in the course of such work?
How can a digitised collection be converted into a language resource to study
the linguistic phenomena in their literary and cultural contexts?</p>
      <p>From the perspective of literary studies, to test the suitability of the methods
of corpus linguistics in solving the problems that literary scholars face.</p>
      <p>One of the corpus usage examples is veri cation of the index of proper names
mentioned in the correspondence. The index was created manually more than
30 years ago and it lists the foreign writers mentioned in the correspondence.
From the visualisation of index statistics (Fig 2) we can only see that Gide is
the most discussed French writer in the whole correspondence.</p>
      <p>Query of KORP allows us to see the concordances of the Gide mentions,
their absolute and relative frequency in the corpus, but not only. KORP allows
us to organise the statistics by all the categories used in the corpus, including
metadata categories. To know who of the authors and when mentioned Andre
Gide, we only have to add those metadata categories (\sender" and \date")
to the statistics criteria; other metadata and linguistic categories can also be
used for more sophisticated statistics that is connected to the concordances and
broader context text blocks.</p>
      <p>From the statistics we see that it was mainly Semper who talked about Gide
and the time interval is 1926 to 1930. It aligns well with the literary history: in
1928 Semper graduated as Master of Arts (magister artium ) from the
University of Tartu (with a masters diploma in literary studies, \The structure of the
literary style of Andre Gide") and continued his work at the university as an
aesthetics and stylistics lecturer.</p>
    </sec>
    <sec id="sec-7">
      <title>Current work ow and future work</title>
      <p>At present stage we have applied the standard procedures for converting text
into KORP-compatible corpus to the Correspondence corpus as well with some
minor additions. The standard procedure is a pipeline consisting of tokeniser,
sentence splitter, morphological analyser (including POS-tagger and lemmatiser)
and morphological disambiguator. Sentence splitter was modi ed according to
non-standard abbreviations found from the Correspondence texts, but
morphological analyser was used with standard options: guesser and forced proper-name
nder.</p>
      <p>Sentence splitter and tokeniser are tools that determine the quality of the
corpus. They can be adjusted by properties of corpus and annotation aims, e.g.
we can annotate a date as one token, or as a series of several tokens. This would
give di erent results of number of tokens and number of sentences.</p>
      <p>Guesser makes the analyser nd answers for words not in analyser's lexicon,
and proper-name nder gives precedence to proper names, if analysis is
ambiguous and word starts with a capital letter. The corpus is not properly evaluated,
but quick overview revealed at least two problems in POS tagging and
lemmatisation. First, the POS tagging seem to be too optimistic about proper names,
and second, foreign words and phrases are analysed as Estonian.</p>
      <p>We compared the number of proper names with manually disambiguated
corpus, and found the general number to be rather normal. Although being far
more than in ction, it fell between relative frequency of proper nouns in
informational texts and journal texts (newspapers). The over-estimation of proper
nouns showed up at the introductory parts of the letters, where capitalised word
\Armas" stands before name of the recipient. The English equivalent would be
\Dear" at the very beginning of letter. Well, \Armas" is a rst name used in
Estonia, but it is a extremely rare one, and it was really never used as such in this
correspondence. We adjusted the work ow accordingly by switching o forced
proper-name nder, and with some decline of the number of proper nouns, got
rid of analysis of \Armas" as name. How much did it a ect the recall of proper
names, needs some further examination.</p>
      <p>As it was mentioned before, both of the authors used foreign language in
their letters as well. Just now it is di cult to tell, how much, because the
Estonian morphological analyser was too eager to nd suitable analyses. It would be
possible to apply some language detecting tool before morphological analysis,
but it is possible that foreign phrases are not so numerous and this strategy
would generate too much noise.
8</p>
    </sec>
    <sec id="sec-8">
      <title>For conclusion</title>
      <p>The number of words and sentences in this corpus is rather big to annotate it
entirely manually, but at least by some extent it would be necessary. Without
manually annotated data it would be di cult, if not impossible, to evaluate
automatic annotation that is trained on nowadays texts. We are going to annotate
at least some parts of the texts manually to nd better methods to improve
automatic analysis.</p>
      <p>Such archival materials as a private correspondences of writers have a great
cultural value, but promise an incredible linguistic importance as well. The
elaborated search possibilities of KORP may be used to verify objectively, via
linguistic analysis, the conclusions which have so far been drawn by using traditional
methods of literary research, mainly using the well known method of slow close
reading of the texts.</p>
      <p>Application of corpus linguistic methods in literary studies based on archival
sources requires meticulous preparation of the material and transforming it into
a text corpus. The text corpus of the correspondence of the writers, translators
and friends Johannes Semper and Johannes Barbarus allows to go further with
principally new type of research questions, about the networks of the writers,
or the ideas, that have in uenced the original works of both authors, leading
more abstract semantic research. This could be the next phrase of the work with
KORP.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forsberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roxendal</surname>
          </string-name>
          , J.:
          <article-title>Korp the corpus infrastructure of sprakbanken</article-title>
          . In: Calzolari,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Dogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.U.</given-names>
            ,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)</source>
          .
          <source>European Language Resources Association (ELRA)</source>
          , Istanbul, Turkey (may
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hardie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>CQPweb combining power, exibility and usability in a corpus analysis tool</article-title>
          .
          <source>International Journal of Corpus Linguistics</source>
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <volume>380</volume>
          {
          <fpage>409</fpage>
          (
          <year>2012</year>
          ). https://doi.org/doi:10.1075/ijcl.17.3.04har
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Laak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viires</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Digital culture as part of estonian cultural space in 20042014: current state and forecasts (</article-title>
          <year>2016</year>
          ), https://www.kogu.ee/vana/ wp-content/uploads/2015/06/EIA-ENG OK
          <article-title>-1</article-title>
          .pdf,
          <source>last accessed 2018-10-31</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Orasmaa</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Towards an integration of syntactic and temporal annotations in estonian</article-title>
          . In: Calzolari,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Loftsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          .
          <source>European Language Resources Association (ELRA)</source>
          , Reykjavik, Iceland (may
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Orasmaa</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petmanson</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tkachenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laur</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaalep</surname>
            ,
            <given-names>H.J.:</given-names>
          </string-name>
          <article-title>Estnltk | nlp toolkit for estonian</article-title>
          . In: Calzolari,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Goggi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Grobelnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Mazo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ).
          <source>European Language Resources Association (ELRA)</source>
          , Paris, France (may
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Rand</surname>
          </string-name>
          , E.:
          <article-title>A third of Estonias cultural heritage to be available digitally in ve years</article-title>
          . Press release (
          <year>2018</year>
          ), https://www.kul.ee/en/news/ third-estonias
          <article-title>-cultural-heritage-be-available-</article-title>
          <string-name>
            <surname>digitally-</surname>
          </string-name>
          ve-years
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Rheault</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beelen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cochrane</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirst</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Measuring emotion in parliamentary debates with automated textual analysis</article-title>
          .
          <source>PLOS ONE</source>
          <volume>11</volume>
          (
          <issue>12</issue>
          ), e0168843 (dec
          <year>2016</year>
          ). https://doi.org/10.1371/journal.pone.0168843
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Schreibman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siemens</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Unsworth</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          . (eds.):
          <article-title>A New Companion to Digital Humanities</article-title>
          . John Wiley &amp; Sons, Ltd (dec
          <year>2015</year>
          ). https://doi.org/10.1002/9781118680605
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Veskis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liba</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Automatic tagger evaluation</article-title>
          .
          <source>NLP course assignment report</source>
          (
          <year>2008</year>
          ), https://entu.keeleressursid.ee/public-document/entity-7052, last accessed 2018-
          <volume>10</volume>
          -31
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Viires</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laak</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Digital humanities meet literary studies: Challenges facing estonian scholarship</article-title>
          . In: Mkel,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Tolonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (eds.) DHN Helsinki 2018. Book of Abstracts (
          <year>2018</year>
          ), https://www.helsinki. /sites/default/ les/atoms/ les/dhn2018-book
          <string-name>
            <surname>-</surname>
          </string-name>
          of-abstracts.pdf,
          <source>last accessed 2018-10-31</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>