<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Suoidne-varra-bleahkka-mála-bihkka-senet-dielku 'hay-blood-ink-paint-tar-mustard-stain' - Should compounds be lexicalized in NLP?</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Chiara Argese</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Linda Wiechetek</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Tommi A Pirinen</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Trond Trosterud</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>35</fpage>
      <lpage>44</lpage>
      <abstract>
        <p>English. Lexicalizing compounds, in addition to treating them dynamically, is a key element in giving us idiomatic translations and detecting compound errors. We present and evaluate an e-dictionary (NDS) and a grammar checker (GramDivvun) for North Sámi. We achieve a coverage of 98% for NDSqueries and of 96% for compound error detection in GramDivvun.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>In this paper1, we discuss the use and necessity
of the lexicalization of compounds – in addition
to the dynamic approach to compounding – in two
rule-based Natural Language Processing (NLP)
applications, a grammar checker GramDivvun and
an electronic dictionary NDS (short for
Neahttadigisánit). We argue for a dual approach and
support this view with an evaluation of these tools.
For comparison, we also look at a third application,
a corpus tool (Korp) for the North Sámi corpus
SIKOR. SIKOR, the Sámi International KORpus,
is the collection of texts in different Sámi languages
compiled by UiT The Arctic University of Norway
and the Norwegian Sámi Parliament.</p>
      <p>In the past, we have mostly focussed on the
dynamic approach to morphological analysis. This
means that we have a lexicon with lemmata and
stems, which in a finite-state manner are combined
1Copyright ©2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
with inflectional and derivational affixes and other
stems and modified when morpho-phonological
processes apply. In this way the linguistic
processes inflection, derivation and compounding are
modelled in a dynamic way, i.e. by means of
concatenation and composition as opposed to listing of
all forms. Lexicalization, i.e. listing compounds
or inflected word forms as such, is the
alternative approach to the dynamic one. In addition to
these two approaches we also use guessers for
certain tasks, i.e. proper name guessing in
morphosyntactic parsing. Our approach is entirely
rulebased and open source. Within our 20 year
experience with language tools for the Sámi languages
and other languages with complex morphology, we
have achieved good results and produced reliable
tools.</p>
      <p>
        There are a number of approaches to error
detection of a few errortypes for morphologically
complex - although less complex than North Sámi
languages like Latvian
        <xref ref-type="bibr" rid="ref5">(Deksne, 2019)</xref>
        and
Russian
        <xref ref-type="bibr" rid="ref11 ref15">(Rozovskaya and Roth, 2019)</xref>
        . The
Latvian neural network grammar checker focusses
on preposition-postposition confusion,
adjectivenoun agreement, mood errors in verb forms,
number and case in noun forms, definiteness of
adjectives and missing commata. All of these error
types have a good performance with precisions
between 78% and 98.5%. Judging from their regular
expressions to insert artificial errors, most of their
error types seem to be fairly local errors that can
be resolved based on bigrams.
      </p>
      <p>The Russian system focusses on more advanced
error types - case, number agreement, gender
agreement, preposition and aspect. However, the
results show that the system is still in its initial
phase with low precision and recall for most error
types (precision is between 22% and 56%, only
gender agreement reaches 68%, and recall is
significantly lower, between 9% and 36%). None of
these approaches deals with compound error
detection.</p>
      <p>For neural network approaches, large corpora
with error mark-up are necessary, which are not
available for North Sámi. The error marked-up
corpus contains 120 459 words, and when
looking at specific error types – as in this case
compound errors – the corpus is even smaller. The
Russian system is based on an error-marked corpus
of 200k words (deemed too small by its authors),
the Latvian system works with artificial errors, an
approach that can be problematic as it does not
reflect real text errors.</p>
      <p>In compounding, two or several words are
combined to form a new word. In Sámi, Finnic and
Germanic languages, compounding is a
productive process and new compounds like in (1) can
be made on the fly.2 In Romance languages, these
compounds typically correspond to prepositional
constructions (ital. ‘la federa del cuscino del
divano’).3
(1)
soffájguoddájolggoža (North Sámi)
sofajputejtrekk (Norwegian)
‘sofa pillow cover (English)’</p>
      <p>The initial motivation for extensive
lexicalization of compounds of North Sámi goes back to
adapting the spellchecker to users’ needs, i.e.
avoiding false alarms in Ávvir newspaper’s texts.</p>
      <p>
        North Sámi is a Uralic language spoken in
Norway, Sweden and Finland by approximately 25 700
speakers
        <xref ref-type="bibr" rid="ref14">(Simons and Fennig, 2018)</xref>
        . It is a
synthetic language, where the open parts of speech
(PoS) – nouns, adjectives, etc. – inflect for case,
person and number. The grammatical categories
are expressed by a combination of suffixes and
stem-internal processes affecting root vowels and
consonants alike, making it perhaps the most
fusional of all Uralic languages. In addition to
compounding, inflection and derivation are common
morphological processes in North Sámi.
      </p>
      <p>North Sámi has seven morpho-syntactic cases,
i.e. nominative (Nom.), genitive (Gen.), accusative
(Acc.), illative (Ill.), locative (Loc.), comitative
(Com.), and essive (Ess.). Case plays a more
central role in Sámi than in preposition-based case
languages, since here syntactic functions are
identified based on case only. In addition, nouns
can bear possessive suffixes. Verbs are inflected
2To avoid confusion with hyphenated compounds, “j” is
used to mark word boundaries in compounds</p>
      <p>3Although there are a number of real compounds in Italian,
such as fruttivendolo, as well.
for person, number (singular, dual, plural), tense
(present and past tense) and mood (indicative,
conditional, and potential). Derivational processes
(passive, causative, inchoative, diminutive,
reflexive, to name only some of them) enhance the
combinatory possibilities of each verb.</p>
      <p>Table 1 illustrates that compounding in North
Sámi is by no means restricted to noun noun
combinations, but includes a number of other
parts-ofspeech (PoS) as well, also as heads.4</p>
      <p>Type
N N
A.Attr N
Adv N
Pron A
Pron N
Adv
Pcle
Adv V
PrfPrc N
Num
Num
Num N
Num A
Num A</p>
      <p>Example
láhkajrievdadusat
boahttejáigi
dáppejolmmoš
iešguđetjlágan
eanetjlohku
duššejfal</p>
      <p>Gloss and
translation
lawjchange.pl ‘law
changes’
comingjtime ‘future’
herejperson ‘person
from here’
eachjalike ‘different
kinds of’
morejnumber
‘majority’
onlyjreally ‘just’
vuostáijváldojuvvo againstjtake.pass.3sg</p>
      <p>‘received’
mearridanjfápmu decide.prfprcjpower</p>
      <p>‘authority’
oktajnuppejlohkái onejsecondjten.ill</p>
      <p>‘eleven’
1978j-láhka 1978j-law ‘1978 law’
3j-ivnnat 3j-colored ‘3-colored’
golmmajivnnat threejcolored ‘three
colored’</p>
      <p>
        In North Sámi, compounds are formed without
a hyphen, except for those involving a proper noun,
a digit, or an acronym like Davvi-Norgii
‘Northern Norway (Ill.)’, 3-juvllatsykkel ‘tricycle’, and
ILO-álgoálbmotsoahpamuš ‘ILO-indigenous
people agreement’ (
        <xref ref-type="bibr" rid="ref10">Riektačállinrávvagat, 2015</xref>
        , p.46).
There are a number of multiwords where a space
is obligatory (albma ládje ‘properly’ and duollet
dálle ‘sometimes’). Also genitive first compounds
have an alternative interpretation when written
apart, which makes error detection more difficult.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        The North Sámi tools described in this
article – NDS, Korp for SIKOR and
GramDivvun
        <xref ref-type="bibr" rid="ref16">(Wiechetek, 2012)</xref>
        – all rely on the
Giel4The following abbreviations are used: N=noun, V=verb,
A=adjective, Attr=attributive, Adv=adverb, Pron=pronoun,
Pcle=particle, PrfPrc=past participle, Num=numeral,
Prop=propernoun.
laLT infrastructure
        <xref ref-type="bibr" rid="ref8">(Moshagen et al., 2013)</xref>
        , a
technological framework for managing lexical data and
building it into language technology applications
including e-dictionaries and grammar checkers.
All of them make use of a morphological
analyzer, an FST (Finite-State Transducer) described
in Pirinen (2014), where word formation processes
are moduled. Additionally, SIKOR and
GramDivvun include a Constraint Grammar-based
syntactic analysis. The full modular structure of the
latter is described in Wiechetek (2019b).
      </p>
      <p>
        The computational modeling of the language is
done using finite-state morphology
        <xref ref-type="bibr" rid="ref2">(Beesley and
Karttunen, 2003)</xref>
        . The method of recognizing
grammatical words as well as querying their
grammatical information is based on looking up the
words in an FST that contains the morphological
dictionary of the language. There are two types of
compounds in the language model: the ones that
are stored in the lexicon as lexicalized units and the
ones generated dynamically using a compounding
model. Table 2 gives the statistics over the length
of lexicalized compounds.5
      </p>
      <p>Lexicalized four-element compounds are quite
common in the noun lexicon, e.g.
davvisámegielterminologiĳa ‘North Sámi language
terminology’. Even six-element compounds
(sáivačáhceguollevuostáiváldindilli ‘fresh water fish receive
situation’) can be found.</p>
      <p>The different types of North Sámi compounds in
Table 1 are not treated equally in the morphological
analyzer. Only the compounds in the first two lines
can be derived dynamically. All others need to be
lexicalized, i.e. listed in the lexicon, to receive a
compound analysis. Numeral compounding is not
treated dynamically in the FST. The dynamic
compounds are generated from the dictionary by
concatenating word forms (such as a genitive or
nominative noun followed by other noun) and adding
a compound tag +Cmp. The main dynamic
compounds are (derived and non-derived) noun + noun
pairs. One feature of the underlying technology
is that the compounding mechanism is capable of
modeling infinitely long compounds: for
example nouns of any magnitude are compounds and
modeled by the finite-state automaton. Since the
compounding mechanism of an FST is very
powerful, it also leads to ambiguity. When we allow
arbitrary lexemes to combine to form compounds,
5The table is based on the dictionary size at the time of the
writing (September 2020); it is actively developed daily.
Further abbreviations are Adp=adposition, Conj=conjunction.
some will overlap other existing lexemes, cf. ex.
(2).
(2)</p>
      <sec id="sec-2-1">
        <title>Davvi regiuvdna</title>
        <p>North region;direction.oven
‘The northern region’
Here, regiuvdna ‘region’ has a typical spelling
error, o&gt;u. The FST analyzes it as a misspelling of
regiovdna ‘region’, but also as a compound with
the elements regi, a common wrong form of regiĳa
‘direction’, and uvdna ‘oven’. While this example
has only two possible analyses, twenty or more
different analyses are not uncommon.</p>
        <sec id="sec-2-1-1">
          <title>Roots PoS</title>
          <p>N
Num</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Prop</title>
        <p>A
V
Adv
Adp
Conj
2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Compounds in three NLP applications</title>
      <p>
        We present three applications, an e-dictionary, a
corpus tool, and a grammar checker tool.
3.1
The North Sámi – Norwegian dictionary contains
25 000 lemmata and uses an FST. The e-dictionary
was first implemented in 2013 with no use of
relational databases (all linguistic resources are
contained within static files and external
commandline tools)
        <xref ref-type="bibr" rid="ref12">(Ryan Johnson, 2013)</xref>
        . It is an intelligent
dictionary in the sense that is able to look up North
Sámi word forms and find lemmas via the FST. It
also allows a tolerant mode, which accepts the
letters acdnstz for áčđŋš-tž in addition to their usual
values. The e-dictionary can split compounds to
provide the user with its elements as well as the
whole compound if a translation is available. The
lexicalization of compounds is important since the
translation of the compound cannot necessarily be
derived from the translation of its parts
        <xref ref-type="bibr" rid="ref1">(Antonsen,
2018, p.54)</xref>
        .
      </p>
      <p>In the FST 90% of the 100 000 nouns, and in
the dictionary 75% of the 25 000 nouns are
compounds.
3.2</p>
      <sec id="sec-3-1">
        <title>A corpus tool</title>
        <p>
          The web application and corpus search tool
Korp
          <xref ref-type="bibr" rid="ref4">(Borin et al., 2012)</xref>
          does not show the internal
structure of compounds in SIKOR. Neither
lexicalized, nor dynamic compounds are searchable as
either the lexicalized analysis is picked instead of
the dynamic one or – in the case of compounds
that are not listed in the lexicon – a lexicalized
compound is made by the preprocessor. This is
a problem inherent in the implementation of the
tool. However, when searching for the compound
tag used in the FST (+Cmp), there are 94 658
results. The reason for that is that the first element in
split compounds in coordination receives a specific
compound tag (+Cmp/SplitR) as well.
        </p>
        <p>Table 3 shows the statistics for compounds in
SIKOR.6 The results are obtained using the scripts
that can be found in GiellaLT.7 According to our
analyses 8.6% of the tokens in corpus are
compounds, and 86% are lexicalized. The rest is mainly
composed of 2-elements compounds (13.4%) and
a very small part of 4-7 elements (0.5%).</p>
        <p>Many of the longer compounds in SIKOR are
quite creative and are hyphenated as the one in
ex. (3).
(3)
suoidne-varra-bleahkka-mála-bihkka-senet-dielku
hay-blood-ink-paint-tar-mustard-stain
mu báiddis lei dušše lihkohisvuohta.
my shirt.loc was only mishap
‘The hay-blood-ink-paint-tar-mustard-stain on my
shirt was only a mishap.’</p>
        <sec id="sec-3-1-1">
          <title>Parts PoS</title>
          <p>N
Prop</p>
          <p>2
96.2
3.8</p>
          <p>3
98.9
1.1</p>
          <p>4
89.2
10.8</p>
          <p>5</p>
          <p>
            The current public version of the Sámi corpus
SIKOR
            <xref ref-type="bibr" rid="ref13">(SIKOR, 2018)</xref>
            (in Korp) consists of 32.2
million words. It was analyzed with a preprocessor
          </p>
        </sec>
        <sec id="sec-3-1-2">
          <title>6The search was done on 2020-09-07.</title>
          <p>7https://github.com/giellalt/
conf-clicit2021
that does not distinguish between lexicalized and
dynamic compounds. The (non-public) version of
SIKOR used in this article makes this distinction,
though, as will future versions in Korp.</p>
          <p>A search for compound tags only returns split
compounds, i.e. the first coordinated hyphenated
nominal element, cf. in ex. (4), i.e. riddo- ‘coast-’.
(4)
riddo- ja vuotnaguovlluin
coast- and fjordregion.loc.pl
‘in coastal and fjord regions’</p>
          <p>GiellaLT has already produced a solution, i.e.
a tag for cohorts with a dynamic compound
(&lt;with-dynamic-compound&gt;) added by a
Constraint Grammar module. However, this tag does
not provide any information about the number of
elements and the beginning and ending of each
element.
3.3</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>A grammar checker (GramDivvun)</title>
        <p>
          GramDivvun, the North Sámi grammar
checker
          <xref ref-type="bibr" rid="ref15 ref16">(Wiechetek et al., 2019b)</xref>
          takes
input from the FST to a number of other modules,
the core of which are several Constraint Grammar
modules. Constraint Grammar is a rule-based
formalism for writing disambiguation and syntactic
annotation grammars
          <xref ref-type="bibr" rid="ref6 ref7">(Karlsson, 1990; Karlsson et
al., 1995)</xref>
          . In our work, we use the free open source
implementation VISLCG-3
          <xref ref-type="bibr" rid="ref3">(Bick and Didriksen,
2015)</xref>
          . All components are compiled and built
using the GiellaLT infrastructure
          <xref ref-type="bibr" rid="ref8">(Moshagen et
al., 2013)</xref>
          .
        </p>
        <p>Lexicalization of compounds is relevant for
grammar checking within compound error
detection. One common error that cannot be resolved by
a spellchecker is the spelling of compounds as two
or more words. GramDivvun performs this type
of error detection as part of the tokenization. The
tokenization is done in two steps. In the first step
potential compounds are tokenized ambiguously
(either as one or as two words, the first of which is
accompanied by an errortag). In the second step,
a Constraint Grammar module8 selects or removes
the error reading. Two conditions need to be met to
find the compound error: 1. the compound needs
to be lexicalized, and 2. the syntactic context needs
to support the compound reading.</p>
        <p>The syntactic context is specified in
handwritten Constraint Grammar rules. The
8https://github.com/giellalt/lang-sme/blob/
3a43911929458fd39da309ed23178bf5dbd04bcd/
tools/tokenisers/mwe-dis.cg3
REMOVE-rule below removes the compound
error reading (identified by the tag Err/SpaceCmp)
if the head is a 3rd person singular verb (cf. l.2)
and the first element of the potential compound is
a noun in nominative case (cf. l.3). The context
condition further specifies that there should be a
finite verb (VFIN) somewhere in the sentence (cf.
l.4) for the rule to apply.</p>
        <p>All possible compounds written apart are
considered to be errors by default, unless the lexicon
specifies a two or several word compound or a
syntactic rule removes the error reading. There are
numerous syntactic contexts where the potential
parts of compounds make perfectly sense. In the
case of noun-noun compounds, the second element
can for example be a simple adverbial, as in ex. (5).
The second element can be homonymous with
another PoS, it can be a finite verb or an infinitive.
(5)
son lea boarráseamus mánná joavkkus.
s/he is oldest child group.loc
‘s/he is the oldest child in the group.’
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>We evaluate the e-dictionary (coverage) and the
grammar checker (precision, recall) for
compounding (errors). The corpus search tool does not
exhibit compounding information and is therefore not
evaluated.
4.1
We analyzed the logs for NDS (Neahttadigisánit)
for 2019, and found that 12.6% of the types in
the user queries are compounds. The results are
obtained using the scripts that can be found in
GiellaLT 7. The amount of lexicalized compounds in
the logs (72.1%) is approximately the same as in
the dictionary, where it is 75% (cf. Section 3.1
above). As much as 98% of the compound queries
get a translation, either a lexicalized one or of its
parts. Thus dynamic compounding contributes
with a substantial improvement to dictionary
coverage. If the alternatives are “getting no help from
the dictionary” and “getting help to translate the
parts” then the latter is to be preferred, even though
the correct translation would be different from just
joining the parts. For example, the compound
word ruhtahearrá ‘rich man’ is not lexicalized in
NDS but it does get a translation of its parts ruhta
‘money’ and hearrá ‘man’, which can help the user
to understand the meaning of the compound word
itself.</p>
      <p>Most of the non lexicalized compounds are
composed of 2 elements (96% in the logs and 93% in
the entries). When analyzing the entries in the
dictionary, we found that 24.8% are compounds and
of those 97.6% are lexicalized. Table 4 shows PoS
for compounds in NDS logs and entries.</p>
      <p>Parts
PoS
N
A
Prop
V
Adv</p>
      <p>Logs</p>
      <p>
        Entries
L
We evaluate error detection for syntactic
compound errors (i.e. words that are written apart
and should be a compound) in GramDivvun in two
ways. Firstly, we compare last year’s results in
Wiechetek (2019a) with a newer version of
GramDivvun, from now on referred to as the
Nodalida-corpus. Last year’s results are based on
version r183544
        <xref ref-type="bibr" rid="ref15 ref16">(Wiechetek et al., 2019a)</xref>
        9. The new
results are based on version r2851010 of
GramDivvun.
      </p>
      <p>However, as the focus in the last analysis was a
different one, i.e. we evaluated other error types as
well, we ran a second evaluation on a 2 363
wordcorpus11 specifically made to test compound
error detection, i.e. every sentence contains a
potential compound. These sentences are hand-selected
from SIKOR.</p>
      <p>The results of the evaluation are presented in
Table 5. We can see that precision has gone
significantly up, i.e. the average precision is 95.5%.</p>
      <p>9https://github.com/giellalt/lang-sme/
releases/tag/nodalida-2018 on 2019-09-26
10https://github.com/giellalt/lang-sme/
releases/tag/clicit on 2020-09-07</p>
      <p>11http://gtsvn.uit.no/freecorpus/orig/sme/
odda_mahppa/compounds.correct.txt
However, the recall has gone down to average 46%.
We are investigating the reasons for that. But in
general, a high precision is desirable in grammar
checking, even at the cost of a lower recall.</p>
      <p>The results of the evaluation of GramDivvun
compound grammar checking are shown in Table 5.</p>
      <sec id="sec-4-1">
        <title>Measure</title>
      </sec>
      <sec id="sec-4-2">
        <title>Precision</title>
      </sec>
      <sec id="sec-4-3">
        <title>Recall</title>
      </sec>
      <sec id="sec-4-4">
        <title>F1-Score TP FP FN</title>
        <p>False negatives are typically due to the lack
of lexicalization. Many of those are proper
noun combinations which are very productive,
e.g. Murmánska-aviisa ‘Murmansk newspaper’,
Várggát-festiválas ‘at the Várggát festival’,
kmgalba ‘km sign’ and Divttasvuotna-regiovnna
‘Divttasvuotna region’.</p>
        <p>Other reasons are certain (unlikely) analyses of
especially the first element, e.g. that generally
suggest a syntactic construction rather than a
compound as in ex. (6). Here the first element duorastat
‘Thursday’ has a finite verb reading as well.
dán duorastat veaiggi.
this.gen Thursday twilight.gen
‘this Thursday evening’
(6)
(7)</p>
        <p>The false positive is due to an error in the
recognition of the span of the target. In ex. (7), lulli sámi
guvlui is concatenated, but it should only be lulli
sámi.</p>
        <p>dohko lulli sámi guvlui.
thither South Sámi area.ill
‘thither towards the South Sámi area.’
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We have shown that the lexicalization of
compounds – in addition to their dynamic treatment
– is useful and necessary for two language
applications for North Sámi, an e-dictionary (NDS) and a
grammar checker (GramDivvun). The evaluation
of NDS shows that we get a good coverage: 98%
of the compounds logged do get a translation and
72% are lexicalized in the FST. The evaluation of
GramDivvun has shown that we manage to identify
compound errors with a precision of 98% and a
recall of 49% utilising a combination of information
from the lexicon and syntax.</p>
      <p>We conclude that there are perfectly good
reasons for lexicalizing compounds, i.e. providing
idiomatic translations for when it cannot be derived
from the parts, and to support compound
grammar checking. At the same time, lexicalization can
dissimulate word formation information in corpus
tools. This can be resolved and we have already
implemented a solution in Constraint Grammar to
make the information available in a future version
of the corpus tool. As dynamic compounding is
limited to few PoS at the moment, in the future
we want to investigate and model compounding
of other PoS (in the FST). Also experiments with
neural network approaches and a comparison of the
results to our rule-based grammar checker could be
an interesting future project.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Lene</given-names>
            <surname>Antonsen</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Sámegielaid modelleren - huksen ja heiveheapmi duohta giellamáilbmái. [Modeling Saami languages. Construction and adaptation to real-world linguistic issues]</article-title>
          .
          <source>Ph.D. thesis</source>
          , UiT The Arctic University of Norway, Tromsø.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Kenneth R.</given-names>
            <surname>Beesley</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lauri</given-names>
            <surname>Karttunen</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Finite State Morphology. CSLI Studies in Computational Linguistics</article-title>
          .
          <source>CSLI Publications</source>
          , Stanford.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Eckhard</given-names>
            <surname>Bick</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tino</given-names>
            <surname>Didriksen</surname>
          </string-name>
          .
          <year>2015</year>
          . CG-3
          <article-title>- beyond classical Constraint Grammar</article-title>
          . In Beáta Megyesi, editor,
          <source>Proceedings of the 20th Nordic Conference of Computational Linguistics (NoDaLiDa</source>
          <year>2015</year>
          ), pages
          <fpage>31</fpage>
          -
          <lpage>39</lpage>
          . Linköping University Electronic Press, Linköpings universitet.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Lars</given-names>
            <surname>Borin</surname>
          </string-name>
          , Markus Forsberg, and
          <string-name>
            <given-names>Johan</given-names>
            <surname>Roxendal</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Korp - the corpus infrastructure of språkbanken</article-title>
          . In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Jan Odĳk, and Stelios Piperidis, editors,
          <source>Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC</source>
          <year>2012</year>
          ).
          <article-title>European Language Resources Association (ELRA).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Daiga</given-names>
            <surname>Deksne</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Bidirectional lstm tagger for latvian grammatical error detection</article-title>
          . In Ekštein K. (eds) Text, Speech, and Dialogue.
          <source>TSD 2019. Lecture Notes in Computer Science</source>
          , vol
          <volume>11697</volume>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Fred</given-names>
            <surname>Karlsson</surname>
          </string-name>
          , Atro Voutilainen, Juha Heikkilä, and
          <string-name>
            <given-names>Arto</given-names>
            <surname>Anttila</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Constraint Grammar: A Language-Independent System for Parsing Unrestricted Text</article-title>
          . Mouton de Gruyter, Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Fred</given-names>
            <surname>Karlsson</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>Constraint Grammar as a Framework for Parsing Running Text</article-title>
          . In Hans Karlgren, editor,
          <source>Proceedings of the 13th Conference on Computational Linguistics (COLING</source>
          <year>1990</year>
          ), volume
          <volume>3</volume>
          , pages
          <fpage>168</fpage>
          -
          <lpage>173</lpage>
          , Helsinki, Finland. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Sjur N.</given-names>
            <surname>Moshagen</surname>
          </string-name>
          , Tommi A.
          <string-name>
            <surname>Pirinen</surname>
            , and
            <given-names>Trond</given-names>
          </string-name>
          <string-name>
            <surname>Trosterud</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Building an open-source development infrastructure for language technology projects</article-title>
          .
          <source>In NODALIDA.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Tommi A.</given-names>
            <surname>Pirinen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Krister</given-names>
            <surname>Lindén</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Stateof-the-art in weighted finite-state spell-checking</article-title>
          .
          <source>In Proceedings of the 15th International Conference on Computational Linguistics and Intelligent Text Processing -</source>
          Volume
          <volume>8404</volume>
          ,
          <year>CICLing 2014</year>
          , pages
          <fpage>519</fpage>
          -
          <lpage>532</lpage>
          , Berlin, Heidelberg. Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Riektačállinrávvagat. 2015.</given-names>
            <surname>Riektačállinrávvagat</surname>
          </string-name>
          . Sámedikki giellaossodat/Sámedikki oahpahusossodat, Guovdageaidnu.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Alla</given-names>
            <surname>Rozovskaya</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Roth</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Grammar error correction in morphologically rich languages: The case of russian</article-title>
          .
          <source>In Transactions of the Association for Computational Linguistics</source>
          , vol.
          <volume>7</volume>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Trond</given-names>
            <surname>Trosterud Ryan Johnson</surname>
          </string-name>
          , Lene Antonsen.
          <year>2013</year>
          .
          <article-title>Using finite state transducers for making efficient reading comprehension dictionaries</article-title>
          .
          <source>In Proceedings of the 19th Nordic Conference of Computational Linguistics (NoDaLiDa</source>
          <year>2013</year>
          ),
          <source>Proceedings Series</source>
          <volume>16</volume>
          :
          <fpage>59</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>SIKOR.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>SIKOR uit norgga árktalaš universitehta ja norgga sámedikki sámi teakstačoakkáldat</article-title>
          ,
          <source>veršuvdna 06.11</source>
          .
          <year>2018</year>
          . http://gtweb.uit.no/korp. Accessed:
          <fpage>2018</fpage>
          -11-06.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Gary F.</given-names>
            <surname>Simons</surname>
          </string-name>
          and Charles D. Fennig, editors.
          <source>2018. Ethnologue: Languages of the World. SIL International</source>
          , Dallas, Texas, twenty
          <article-title>-first edition</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Linda</given-names>
            <surname>Wiechetek</surname>
          </string-name>
          , Kevin Brubeck Unhammer, and Sjur Nørstebø Moshagen. 2019a.
          <article-title>Seeing more than whitespace - Tokenisation and disambiguation in a North Sámi grammar checker</article-title>
          .
          <source>In Proceedings of the third Workshop on the Use of Computational Methods in the Study of Endangered Languages</source>
          , pages
          <fpage>46</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Linda</given-names>
            <surname>Wiechetek</surname>
          </string-name>
          , Sjur Nørstebø Moshagen, Børre Gaup, and Thomas Omma. 2019b.
          <article-title>Many shades of grammar checking - launching a constraint grammar tool for north sámi</article-title>
          .
          <source>In Proceedings of the NoDaLiDa Linda Wiechetek</source>
          .
          <year>2012</year>
          .
          <article-title>Constraint Grammar based correction of grammatical errors for North Sámi</article-title>
          . In G. De Pauw,
          <string-name>
            <surname>G-M de Schryver</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          <string-name>
            <surname>Forcada</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Sarasola</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          <string-name>
            <surname>Tyers</surname>
          </string-name>
          , and P.W. Wagacha, editors,
          <source>Proceedings of the Workshop on Language Technology for Normalisation of Less-Resourced Languages (SALTMIL 8/AFLAT 2012)</source>
          , pages
          <fpage>35</fpage>
          -
          <lpage>40</lpage>
          , Istanbul, Turkey, may.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>