<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>as part of the Text Corpus Creation Pipeline at the Language Bank of Finland</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tommi Jauhiainen</string-name>
          <email>tommi.jauhiainen@helsinki.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jussi Piitulainen</string-name>
          <email>jussi.piitulainen@helsinki.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erik Axelson</string-name>
          <email>erik.axelson@helsinki.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Krister Lindén</string-name>
          <email>krister.linden@helsinki.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Helsinki</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <fpage>251</fpage>
      <lpage>259</lpage>
      <abstract>
        <p>The Language Bank of Finland hosts text corpora originating from Finland. Two of the most used ones are the Newspaper and Periodical Corpus of the National Library of Finland and the Suomi24 Corpus. The Language Bank has received considerable additions to both corpora and is currently creating new versions of the corpora. We are debuting language identification as part of the corpus creation pipeline. As a language identifier, we are using our recently published HeLI-OTS software. This paper investigates the results and the quality of language identification.</p>
      </abstract>
      <kwd-group>
        <kwd>text corpus</kwd>
        <kwd>language identification</kwd>
        <kwd>OCR</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        FIN-CLARIN cooperation. The LBF hosts text corpora in Finnish as well as in other languages if
they originate from or are used in Finland. Two of LBFs most used corpora are the Finnish
Subcorpus of the Newspaper and Periodical Corpus of the National Library of Finland (KLK-fi) 2 and
LGOBE
CEUR
Workshop
Proceedings
the Suomi24 Sentences Corpus 2001-2017 (suomi24-2001-2017)3. The over 5 billion tokens of the
KLK-fi corpus originate from the Finnish magazines and newspapers starting from 1820, which
the National Library of Finland (NLF) has digitized. The suomi24-2001-2017 corpus contains
over 4 billion tokens from the various discussion forums of the Suomi24 social networking
website from 2001 to 2017. LBF has received considerable additions to both corpora: c. 0.5
billion tokens for Suomi24 and c. 18 billion tokens for NLF. The first version of the new Suomi24
corpus is already available4, and new versions of the KLK corpora are planned to be published
during summer 2022.5 We are debuting language identification (LI) as part of this process. The
LBF has previously participated in using LI in creating text corpora for underresourced Uralic
languages within the Finno-Ugric Languages and the Internet -project [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        As a language identifier, we are using the recently published HeLI-OTS software. 6
HeLIOTS is an of-the-shelf language identifier equipped with language models for 200 languages.
In HeLI-OTS, the ISO 639-3 standard is used to determine the language division, and when
determining the language of a text, the program returns an ISO 639-3 identifier as a result. As
a method for LI, the program uses the HeLI method developed at the University of Helsinki
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This method is currently state-of-the-art for LI between many languages, as evidenced by
its rfist place on the ULI-178 track of the Uralic Language Identification (ULI) 2020 and 2021
shared tasks [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We are identifying the language of each sentence, and it will be possible to use
the resulting information as part of queries in the Korp user interface.7 This paper investigates
the results and the quality of the LI process.
      </p>
      <p>In Section 2, we explain how sentences are extracted from the lfies received from the NLF.
Section 3 describes the evaluation dataset created from the new NLF and Suomi24 material. The
following Section provides the numerical results of the LI on this dataset. In Section 5, we take
a closer look at the errors the best models made on the development sets of both materials. The
last Section discusses some improvement ideas and concludes the paper’s nfidings.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Tokenization and sentence detection</title>
      <p>Many of the “sentences” problematic for LI in our corpora are not well-formed sentences in
the rfist place. Therefore, we thought it essential to shed some light on the process of how
these sentence-like objects were constructed. The rest of the paper will refer to these objects as
sentences even if they could not be considered sentences by any linguist.</p>
      <p>Each page of the NLF magazines and newspapers of the corpus has been OCR’d into an
XML lfie in ALTO/METS format8. The XML file consists of one &lt;Page&gt; element consisting of
&lt;TextBlock&gt; elements that consist of &lt;TextLine&gt; elements consisting of &lt;String&gt; elements. One
Page comprises one page. One TextBlock comprises one block of text, i.e., paragraph, title, table,
legend, and so on. One TextLine usually contains one line separated by a line/page break. String
is basically anything separated by space. Each TextBlock is retokenized with the command-line
3suomi24-2001-2017-korp-v1-1, http://urn.fi/urn:nbn:fi:lb-2020021803
4http://urn.fi/urn:nbn:fi:lb-2021101525
5http://urn.fi/urn:nbn:fi:lb-202009152
6http://urn.fi/urn:nbn:fi:lb-2022021303
7https://korp.csc.fi/korp/
8https://www.loc.gov/standards/alto; https://www.loc.gov/METS/
tool finnish-tokenize (part of Finnish Tagtools v1.5.19). The input to finnish-tokenize is one
string that contains all TextLine and String elements in the given TextBlock, separated by spaces.
The result from finnish-tokenize is a list of sentences, where a sentence is a list of tokens.
Tokenizations from the XML file and finnish-tokenize are aligned, preferring the result from
ifnnish-tokenize . Alignment is needed as the XML file contains attributes such as OCR character
confidence and word confidence values for each token. We used ifnnish-tokenize for the texts
written in other languages as well, which is not optimal, but remedying this is out of the scope
of this article.</p>
      <p>
        For the Suomi24 material, we used UDPipe version 1.2.0 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] with the finnish-tdt (Turku
Dependency Treebank) model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to segment paragraphs into sentences and tokens.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Evaluation environment</title>
      <p>In order to improve and test the eficiency of the language identifier, we manually annotated
random samples from the corpora. As both of the corpora contain fairly recent material, we
scrambled all the numbers within the sentences to anonymize any possible personal information
they might contain, such as street addresses and phone numbers. The scrambled texts contain
the same amount of number characters as the originals, and each of the original number
characters has been replaced by a random number character. For example, “Salpausseläntie 7 E
23 , 00710 Helsinki 71 , puh. 225 932 .” could become “Salpausseläntie 2 E 57 , 55667 Helsinki 06
, puh. 116 071 .”. The development and the test data are available via LBF10 with a CC-BY-NC
license for the Suomi24 part and CC-BY license for the NLF part.</p>
      <sec id="sec-3-1">
        <title>3.1. NLF Random evaluation set</title>
        <p>The LBF has published four diferent types of corpora originating from the material collected by
the NLF.11 Two of these are language-based, and two are based on the date of the publication of
the material in question. So far, the material for the language-based, Swedish or Finnish, has
been selected using the language tags provided by the NLF. In 2021, LBF received a new batch
of material from the NLF.</p>
        <p>The random development and test sets for the NLF data were generated from publications
OCRd by the NLF during 1.1.2014-30.6.2014 and 1.1.2015-30.6.2015.12 The publications contained
“texts” tagged as written in Finnish, Swedish, English, and German. The number of texts in each
language can be seen in Table 1. We randomly selected 500 for both Finnish and Swedish, 200
for English, and all six German texts from those texts. After this, from each text, we randomly
sampled 20 “sentences” generated as described in Section 2. The number of sentences for each
language is also seen in Table 1.</p>
        <p>We manually identified the language of each sentence, except those of the Russian collection,
due to our low ability to read Cyrillic. Cyrillic sentences in the collections of other languages
were annotated as Russian if they contained more than individual Cyrillic letters. This means
9http://urn.fi/urn:nbn:fi:lb-2021042101
10http://urn.fi/urn:nbn:fi:lb-2022021301
11http://urn.fi/urn:nbn:fi:lb-2021092404
12They contain publications from the early 1800s to the 2010s.
that sentences annotated as “rus” could be written in any language using Cyrillic characters. If
the sentence seemed to consist solely of one or more proper names, it was annotated as “xxx”.
The annotation “xxx” was also used if it was impossible to determine the language because
the sentence contained only non-linguistic contents such as severe OCR errors or numbers.
The number of sentences in each detected language can be seen on the left side of Table 2. 26
sentences were also deemed multilingual.</p>
        <p>From these sentence collections, we selected the first 50% of the sentences from each of the
four NLF tagged languages for the development set and the last 50% for the test set. The language
distribution within the two sets was not equal, as can be seen on the right side of Table 2. For
example, all the Cyrillic (rus) texts were in the latter half of the original collection and thus
ended up in the test set. When LI methods are developed using these sentence collections, it
is essential that the information about the language distribution in the test set is not used to
improve the results.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Suomi24 Random evaluation set</title>
        <p>The new material for the Suomi24 corpus was already imported to Korp, so we used the Korp
interface concordance function to sample 3,000 sentences from 2018–2020 randomly. The total
number of sentences for each year ranged from 12.5 to 14.5 million. Unlike in the NLF material,
there are no existing language tags available for the texts. The supposition was that most of
the texts would be in Finnish, but texts or sentences in other languages could also be found.
The language of these 9,000 sentences was annotated manually, and it supported the prior
supposition as only c. 0.6% of the sentences were written in languages other than Finnish.
Again, we divided the dataset into equal portions for development and testing. The dataset
contained some Swedish and English sentences, but over 96% of the dataset was clearly in
Finnish, as can be seen in Table 3.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>
        In this Section, we use the HeLI-OTS language identifier to identify the language of the sentences
in the evaluation set. The 1.1 version of the identifier includes only one language model for
Finnish trained from one million sentences of 2002 Web crawl data from Leipzig Corpora
Collection [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The 1.2 and 1.3 versions have several additional language models for Finnish,
which were created from data gathered during the Finno-Ugric Languages and the Internet
-project to overcome the problem of dialectal Finnish being identified as some of the close
relatives of Finnish such as Kven Finnish or Karelian. This same problem had been noticed
by early users of the 1.1 version, which is why the 1.2 version was developed. The eficacy
of adding the additional language models for Finnish was never adequately evaluated during
the Finno-Ugric Languages and the Internet -project, and the current study describes the first
proper evaluation of this strategy.
      </p>
      <p>In these first experiments, we set a baseline for LI results for this dataset, focusing on the
results attained on the development set. We evaluate the HeLI-OTS 1.3 version, with and
without the additional Finnish language models. The results for the NLF development set are in
Table 4 and the results for the Suomi24 development set are in Table 5. Using the additional
models for Finnish, the overall micro F1 score for the NLF test set was 87.11 and 97.60 for the
Suomi24 test set. The results on the test set could be improved easily by using the information
in the development set in various ways, such as using the development set as additional training
material. However, we wanted to see how the of-the-shelf version of the language identifier
fares with both sets and leave further experiments to future work.
Language
Finnish (fin)
Swedish (swe)
English (eng)
German (deu)
Unknown (xxx)
French (fra)
Latin (lat)
Italian (ita)
Dutch (nld)
Multilingual
All
Language
Finnish (fin)
Swedish (swe)
English (eng)
Unknown (xxx)
Multilingual
All</p>
    </sec>
    <sec id="sec-5">
      <title>5. Error analysis</title>
      <p>In this Section, we focus on the errors made by the HeLI-OTS language identifier with additional
dialectal language models for Finnish. In this study, we refrain from studying the errors made
on the test set in order to be able to continue developing the LI methods using the collection.
Therefore we focus on analyzing some of the errors made on the development set in detail.
5.1. KLK
From the four languages of interest, the lowest F1 score was attained for German with both
relatively low recall and low precision. From the 43 sentences annotated as German in the
development set, only 31 were identified as such. The recall errors seem to be primarily due to
the combined efect of OCR errors and the HeLI language repertoire, including several closely
related languages: Swiss German (4 errors), Bavarian (3), and Pfaelzisch (2). An example of a
sentence identified as Swiss German is “Die Ver n i chtung der boi sch ewistischen Ge f ahr
und der plutokratischen Ausbeutung wird die Möglichkeit einer fr i edlichen harmonischen
und f r uchtbaren Z us amm enar b eit aller Volker des europäischen Kontinents so w ohi auf p
oli tisch em als auf wirtschaftlichem und kultur eli em Geb iet scha f f e n. » ]”. Most of the
precision errors for German came from short sentences tagged as Swedish. An example of such
behaviour is: “Hummern hade galt öfwer bord .”, where the misclassification is probably due to
the combination of OCR errors, “galt” should be “gått”, and not using current spelling, “öfver”
is “över” in standard spelling.</p>
      <p>The other three languages all had F1 scores over 85, with Swedish having the lowest score
due to a very low recall of 77.78%. Of the 902 erroneously identified Swedish sentences, 87 were
identified as Danish, 47 as Norwegian Bokmål, 38 as Nynorsk, 36 as Lushai, 35 as Swiss German,
33 as Finnish, and 30 as Kölsch. The rest were divided between 103 diferent languages. The
sentences incorrectly identified as Danish seem to be primarily due to OCR errors such as “6
anlebning af OfebaftionenS od ) målaren ©parfS i fenafte nummer af SÖ .”, which should be a
Swedish sentence: “I anledning af Redaktionens och målaren Sparss i senaste nummer af B.”.
OCR errors are to blame also for both of the Norwegian languages, Swiss-German, and Kölsch.
The samples mistaken as Finnish are primarily due to Finnish proper names. The samples
mistaken as Lushai are due to “tel” being a word in Lushai and the abbreviation for telephone
number “tel.” being the only word in 30 Swedish sentences.</p>
      <p>The version with extra language models for Finnish made significantly more precision errors
for Finnish, the number of errors rising from 103 to 193. The gold-standard language for 142
of these erroneous identifications was tagged as “xxx”, meaning that they were either
nonlingual or consisted of names or numbers. The 142 sentences included Finnish personal names
“Heikki Jussila .”, addresses “Iso Uusikatu N:o 29 .”, place names “HELSINKI-TAMPERE”, and
names of companies “KONE O.Y .”, as well as their combinations. They also included some
incomprehensible sentences, probably due to OCR problems, which might have originated
from Finnish such as “§nraästi !” and “Miwälista .”. Examples of non-lingual sentences in the
development set are “m . , y .”, “g .”, and “1 § .”.</p>
      <p>Of the 152 incorrectly identified English lines, only two were actual sentences containing
more than one word. Both “Cornelius was obedient .” and “In Jesus ’ name .” were identified as
Afrikaans.
5.2. Suomi24
On the Suomi24 development set, the worst-performing language was English, with an F score
of only 28.57. Only four of the nine sentences tagged as English were well-formed sentences, all
identified as English. For example, two sentences, “wtf” and “WTF?” were identified as Lushai.
These kinds of errors lowered the recall for English. The precision for English was down due to
15 sentences tagged as something else identified as English. Five of the fiteen were HTTPS
addresses containing English words, two were Finnish English multilingual, one was incorrectly
labelled as Finnish, and most of the others contained proper names.</p>
      <p>The Finnish sentences were most often identified as one of the close languages within the
repertoire. Of the 75 incorrectly identified Finnish sentences, eight were identified as Karelian,
six as Ingrian, and five as Estonian. Some of the sentences identified as Karelian were reasonably
well-formed Finnish sentences, such as “En kieliä vieroksu.” and “Ei eletä.”. Also, some of the
sentences identified as Ingrian were well-formed Finnish sentences such as “Kolmestihan se
sinne kaislikkoon lensi.”. As the Ingrian training corpus contained only 582 words, any word
seen in the Ingrian training corpus is also heavily scored as Ingrian in the mystery text. For
example, the word “sinne” is used in well-formed Ingrian sentences in the training corpus.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and Conclusions</title>
      <p>In many cases where the language of a sentence was incorrectly predicted, the correct language
had almost as good a score as the predicted language. It could be beneficial to give some of the
commonly expected languages a slight increase to the prior probability. These probabilities
could be determined using the development set for this particular evaluation setting. However,
they would not be suitable as ready-to-use general-purpose probabilities for the HeLI-OTS
software.</p>
      <p>Another way to improve the results on the test set would be to restrict the HeLI-OTS language
repertoire based on errors seen in the development set. For example, none of the sentences
identified as written in one of the languages related to Finnish in the Suomi24 development set
seemed to be written in them. Leaving them out of the repertoire when processing the test set
would undoubtedly increase the accuracy of the identifications.</p>
      <p>Many sentence identification errors were due to the presence of named entities using common
words from other languages. An example of such a sentence is “Allmänna Svenska Elektriska
Aktiebolaget, Vesteräs, Ruotsi : laite aikalisiin akkumulaattoreihin.”, which contains a name of a
Swedish company in a Finnish sentence.</p>
      <p>
        In addition to using the development material as in-domain training material when processing
the test set, there are ways to adapt the language models to the material being identified. Several
international LI competitions have been won using adaptive language models with the HeLI
and Naive Bayes methods [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. This functionality will be incorporated into HeLI-OTS in the
future, and we will also evaluate its eficacy on this dataset.
      </p>
      <p>Especially the NLF corpus contains many non-lingual sentences, which could be identified as
such if an unseen language detection capability were incorporated into the software. In many
cases, the OCR process had created extra spaces inside words. HeLI-OTS uses spaces in word
tokenization and gives each word equal value, which leads to parts of words being treated as
words and thus into erroneous identifications. It would be worthwhile to see whether it would
be beneficial to ignore spacing completely in the NLF material or try another word tokenization
strategy.</p>
      <p>Our results indicate that implementing some of the previously mentioned functionalities will
significantly improve the accuracy of language identification.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>The research presented in this article has been funded by the Finnish Research Impact
Foundation13 from its Tandem Industry Academy 2020 call in cooperation with Lingsoft Oy 14, the
Academy of Finland, and the University of Helsinki.</p>
      <p>13https://www.vaikuttavuussaatio.fi/en/
14https://www.lingsoft.fi/en</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lindén</surname>
          </string-name>
          ,
          <article-title>Building Web Corpora for Minority Languages</article-title>
          ,
          <source>in: Proceedings of the 12th Web as Corpus Workshop</source>
          , European Language Resources Association, Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>23</fpage>
          -
          <lpage>32</lpage>
          . URL: https://www.aclweb.org/anthology/ 2020.wac-
          <volume>1</volume>
          .4.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lindén</surname>
          </string-name>
          , H. Jauhiainen,
          <article-title>HeLI, a word-based backof method for language identification</article-title>
          ,
          <source>in: Proceedings of the Third Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial3)</source>
          ,
          <source>The COLING 2016 Organizing Committee</source>
          , Osaka, Japan,
          <year>2016</year>
          , pp.
          <fpage>153</fpage>
          -
          <lpage>162</lpage>
          . URL: https://www.aclweb.org/anthology/W16-4820.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gaman</surname>
          </string-name>
          , R. T. Ionescu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lindén</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ljubešić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Partanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priyadharshini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Purschke</surname>
          </string-name>
          , E. Rajagopal,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Scherrer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <article-title>Findings of the VarDial evaluation campaign 2021</article-title>
          , in: Proceedings of the Eighth Workshop on NLP for
          <article-title>Similar Languages, Varieties and Dialects, Association for Computational Linguistics</article-title>
          , Kyiv, Ukraine,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          . URL: https://www.aclweb.org/ anthology/2021.vardial-
          <volume>1</volume>
          .1.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Straka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Straková</surname>
          </string-name>
          , Tokenizing, POS Tagging,
          <article-title>Lemmatizing and Parsing UD 2.0 with UDPipe, in: Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, Association for Computational Linguistics</article-title>
          , Vancouver, Canada,
          <year>2017</year>
          , pp.
          <fpage>88</fpage>
          -
          <lpage>99</lpage>
          . URL: http://www.aclweb.org/anthology/K/K17/K17-3009.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Haverinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nyblom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Viljanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Laippala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kohonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Missilä</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ojala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salakoski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          ,
          <article-title>Building the essential resources for Finnish: the Turku Dependency Treebank</article-title>
          ,
          <source>Language Resources and Evaluation</source>
          <volume>48</volume>
          (
          <year>2014</year>
          )
          <fpage>493</fpage>
          -
          <lpage>531</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Goldhahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Eckart</surname>
          </string-name>
          , U. Quasthof,
          <article-title>Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages</article-title>
          , in
          <source>: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Istanbul, Turkey,
          <year>2012</year>
          , pp.
          <fpage>759</fpage>
          -
          <lpage>765</lpage>
          . URL: http: //www.lrec-conf.org/proceedings/lrec2012/pdf/327_Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          , K. Lindén,
          <article-title>HeLI-based experiments in Swiss German dialect identification</article-title>
          ,
          <source>in: Proceedings of the Fifth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial</source>
          <year>2018</year>
          ),
          <article-title>Association for Computational Linguistics</article-title>
          , Santa Fe, New Mexico, USA,
          <year>2018</year>
          , pp.
          <fpage>254</fpage>
          -
          <lpage>262</lpage>
          . URL: https://www.aclweb.org/anthology/W18-3929.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jauhiainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lindén</surname>
          </string-name>
          ,
          <article-title>Naive Bayes-based experiments in Romanian dialect identification</article-title>
          ,
          <source>in: Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects</source>
          , Association for Computational Linguistics, Kyiv, Ukraine,
          <year>2021</year>
          , pp.
          <fpage>76</fpage>
          -
          <lpage>83</lpage>
          . URL: https://www.aclweb.org/anthology/2021.vardial-
          <volume>1</volume>
          .9.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>