<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Morphological Features of the Irish Universal Dependency Treebank</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>ADAPT Centre, School of Computing, Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computing, Macquarie University</institution>
          ,
          <addr-line>Sydney</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2008</year>
      </pub-date>
      <fpage>111</fpage>
      <lpage>122</lpage>
      <abstract>
        <p>The Universal Dependencies Project1 (Nivre, [9]; Nivre et al., [10]) is an ongoing effort towards creating a set of harmonised dependency treebanks that are annotated and structured according to universal guidelines. This paper reports on the addition of morphological features to the Irish Universal Dependencies Treebank (IUDT). Our feature set subscribes to the feature inventory of the UD Project and has been mapped from Irish morpho-syntactic tags - the output of a Finite State Morphological Analyser for Irish (Uí Dhonnchadha and van Genabith [16]). Irish, a Celtic language, has some relatively unusual morphological features that require language-specific labels not covered by the universal feature set. In this paper, we summarise the Irish-specific features that we have added to this set by explaining the linguistic properties that they each describe. We also report on the first parsing experiments using the IUDT by assessing the effect that the inclusion of morphological features has on parsing accuracy.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The motivation behind the Universal Dependencies Project (Nivre, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]; Nivre et al.,
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]) is to create a set of harmonised dependency treebanks that will facilitate
improved multilingual parsing and better cross-linguistic analysis. Treebank
development for any language is both resource- and time-intensive. Building a large-scale,
fully-annotated treebank, usually requires a large team of linguists and/or
computational linguists, all of whom are collectively responsible for both the design and
annotation of the dataset. Low-resource languages, however, lack the financial
investment enjoyed by those better resourced and more widely researched languages,
and as such, resources such as treebanks often take much longer to produce.
      </p>
      <p>
        Irish is a relatively low-resourced language amongst the UD treebank
collection. The inclusion of low-resourced languages in large projects like this means
that the language can benefit from the experience and contributions of the wider
research community. The Irish Universal Dependency Treebank (IUDT) was
developed as a result of a mapping of the Irish Dependency Treebank (IDT)
annotation scheme (Lynn [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) to the UD annotation scheme (Lynn and Foster [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). Both
treebanks are currently relatively small in size (1020 trees) and lack morphological
features in their initial release. The natural next step in development is therefore
the inclusion of such features.
      </p>
      <p>In this paper, we discuss our extension of the IUDT to include morphological
features. We follow the UD guidelines2 for feature inclusion. While the UD project
defines a feature set that aims to cover linguistic universals shared across the
numerous3 languages that are part of the project, it is also possible for each language
group to define their own language-specific features where necessary. For
example, one of the Irish-specific UD features relates to initial mutation – a linguistic
phenomenon that is common across Celtic languages (see Section 4.1). However,
initial mutation does not feature widely in other languages, and is therefore not
deemed ‘universal’. Here, we list the morphological tags from the UD feature set
that are relevant to Irish and additional new Irish-specific features.</p>
      <p>
        Additionally, we consider that parsing models for morphologically rich
languages can benefit from the inclusion of morpho-syntactic features given sufficient
training data. This additional level of annotation in dependency treebanks, at the
morpho-syntactic level, contributes to a deeper understanding of the data that may
not be achieved through part-of-speech (POS)-tags or dependency labels alone.
This has been demonstrated in various studies, including shared tasks on the
parsing of morphologically rich languages (e.g. Seddah et al, [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]; [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). To this end,
we evaluate the inclusion of these features in the IUDT through empirical methods.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The morphology of Irish</title>
      <p>
        In computational terms, Irish is regarded as a morphologically rich language (Lynn
et al [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]), and as such, Irish parsing models should benefit from the inclusion of
morphological features in the training data. Stenson [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] describes the Irish
language as an “inflectional language, tending more toward isolating than
polysynthetic in general". She also notes that inflection is primarily realised through
suffixation, yet initial mutation (a characteristic of Celtic languages) is also common
and appears in the form of eclipsis (e.g. bord; ar an mbord ‘on the table’) or
lenition (e.g. dath an bhoird ‘the colour of the table’). Another prominent feature (and
also of Scottish Gaelic and Manx), which influences inflection, is the existence of
two sets of consonants, referred to as ‘broad’ and ‘slender’ consonants (Ó Siadhail
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]). For example, buail ‘hit’ + adh (verbal noun suffix) ! bualadh. The
fol
      </p>
      <sec id="sec-2-1">
        <title>2http://universaldependencies.org/u/feat/index.html 3The latest release of UD treebanks (v.4.) includes 64 treebanks, covering 47 languages.</title>
        <p>
          lowing is a short summary of the main inflectional processes in Irish. For a more
detailed description, see The Christian Brothers [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] or Lynn [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>Nouns In terms of nominal inflection, Modern Irish only uses three cases:
common, genitive and vocative. The common case covers nominative, accusative and
dative, yet it should be noted that the dative case is marked on personal pronouns.
Each noun falls into one of five declensions. Case, declension and gender are all
expressed through inflection.</p>
        <p>Verbs Verbs can be marked for both tense and aspect, and inflect for person and
number (e.g. ith ‘eat’; d’ithimis ‘we used to eat’).</p>
        <p>
          Adjectives The Christian Brothers ([
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] p.63) note eight main declensions of
adjectives. They can inflect for genitive singular masculine, genitive singular
feminine and nominative plural (e.g. bacach ‘lame’; bacaigh Gen.Sg.Masc).
Prepositions Simple prepositions can inflect for a pronominal object, indicating
person and number (e.g. le ‘with’; liom ‘with me’; lei ‘with her’).
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Irish Morphosyntactic Tagset</title>
      <p>
        The IDT and IUDT were built upon a gold-standard POS-tagged corpus of Modern
Irish, which was developed by Uí Dhonnchadha [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The POS tags in this
corpus are the output format of the Irish Morphological Analyser (Uí Dhonnchadha
and van Genabith [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]). They are based on a mapping from the Irish PAROLE
Morphosyntactic Tagset (ITÉ [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) to a Finite State Morphological Feature Tagset.
      </p>
      <p>The following is an example sentence labelled with FST tagset output: Tá sé
soiléir ‘It is clear’. Each token is followed by a string of POS tags and
morphosyntactic features, separated by ‘+’. The string sé sé+Pron+Pers+3P+Sg+Masc tells
us that sé (with the same lemma form) is a 3rd person singular personal pronoun,
of masculine gender.</p>
      <p>(1) Tá sé soiléir ‘It is clear’.</p>
      <p>Tá bí+Verb+PresInd
sé sé+Pron+Pers+3P+Sg+Masc
soiléir soiléir+Adj+Base</p>
      <p>Until now, the IUDT contained only the POS tags (coarse- and fine-grained)
of this feature set, i.e. +Pron+Pers, indicating personal pronoun. Our purpose of
now including additional morphosyntactic features in treebank data is to mark
additional lexical and grammatical properties of words that are not available through
POS tags alone. As shown in (1) above, such additional morphological features for
Irish data are readily available for inclusion in the data set, but need to be converted
to the correct representation and align with the UD guidelines.</p>
      <p>
        The IUDT is in the CoNLL format (Buchhloz [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]), where morphological
features are labelled in the FEATS column, and every feature has the form
FeatureName=Value. Every word can have any number of features, separated by the
vertical bar (i.e. a = xjb = yjc = z). The features for the pronoun (sé) in (1) above would
therefore be represented as Gender=Masc|Number=Sing|Person=3. According to
UD guidelines, multiple features (and where applicable, multiple values) should be
ordered in alphabetical order.4
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Mapping Irish Morphosyntactic Tagset to the</title>
    </sec>
    <sec id="sec-5">
      <title>UD feature set</title>
      <p>The Universal Dependencies feature set is a standardised list of morphological
features, which is based on the Interset system (Zeman [17]). It is divided into
lexical and inflectional features. Mapping to the UD feature set was straightforward
for the majority of the tags in the Irish Morphosyntactic Tagset. However, we have
introduced a number of non-universal features and values in order to fully represent
specific linguistic phenomena in the Irish language (see Section 4.1). Table 1
presents the full inventory. Rather than discuss the details of entire feature set here,
we only discuss in detail the newly introduced Irish language-specific features (see
bolded features). For information on all other standard (universal) features, we
point the reader to the Universal Dependencies website.5
4.1</p>
      <sec id="sec-5-1">
        <title>Description of Irish-specific features</title>
        <p>There are a number of morphosyntactic features in Irish for which the UD universal
feature set does not fully account. This is not surprising, as while all languages
share certain linguistic universals, what makes languages unique is often seen in
their differences across syntax or morphology. For a more fine-grained annotation,
we extend the UD feature set to cover these nuances in Irish. Here, we provide a
short description of each of the Irish-specific features.6
Dialect=Munster, Connaught, Ulster There are three main dialects of Irish and
there are a number of lexical items that differ across them. Some surface forms
differ, while sharing the same lemma. For example, the past tense form of the
lemma cuir ‘put’ is chuireas in the Munster dialect, but chuir in the Connaught
and Ulster dialects. Other lexical items differ in both surface and lemma form. For</p>
        <sec id="sec-5-1-1">
          <title>4http://universaldependencies.org/u/overview/morphology.html</title>
          <p>5http://universaldependencies.org/u/feat/all.html
6Not requiring much discussion, note that the Feature Value NomAcc will replace the incorrectly
used Com tag in future versions of the IUDT to indicate for the common case in Irish. Com is used in
other UD treebanks to indicate the Comitative case.
Feature Name</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>Case</title>
          <p>Degree
Dialect†
Definite
Form†
Gender
Mood
Negative
NounType†
NumType
Number</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>PartType†</title>
        </sec>
        <sec id="sec-5-1-4">
          <title>Person</title>
          <p>Poss
PrepForm†
PronType
Reflex
Tense
VerbForm
Voice</p>
          <p>Feature Value ILneflxeiccatiloonral?
Dat, Gen, NomAcc†, Voc Inflectional
Cmp, Pos, Sup Inflectional
Connaught, Munster, Ulster Inflectional
Def, Ind Inflectional
Ecl, Emph, hPref, Len, VF Inflectional
Fem, Masc Inflectional
Cnd, Imp, Ind, Int†, Sub Inflectional
Neg Inflectional
NotSlender, Slender, Strong, Weak Inflectional
Card, Ord, Pers Lexical
Plur, Sing Inflectional
IAndf,, NCummp,l,PCaot,mVpb,,CVoopc, Deg, Lexical
1, 2, 3 Inflectional
Yes Lexical
Cmpd Lexical
Art, Dem, Ind, Int, Prs, Rel Lexical
Yes Lexical
Fut, Past, Pres Inflectional
Cop†, Ger, Inf, Part Inflectional
Auto† Inflectional
Prevalence
example, the word achan ‘every’ is used in Ulster dialect, but gach ‘every’ is used
in the Munster and Connaught dialects.</p>
          <p>
            Form=Ecl, Emph, Len, hPref Initial mutation is a feature of Celtic languages.
It is triggered by a preceding word and affects the spelling of nouns, adjectives
and verbs. Eclipsis (Ecl) on verbs occurs following clitics such as interrogative
particles (an, nach); complementisers (go, nach); and relativisers (a, nach) (Stenson
[
            <xref ref-type="bibr" rid="ref14">14</xref>
            ], pp. 21-26). For example, tugann sé ‘he gives’; an dtugann sé. ‘does he give?’
An example of lenition (Len) on proper nouns is following the vocative particle
a. E.g. Máire ‘Mary’; A Mháire! In addition, some words can trigger the
hprefix (hPref) on the following word. For example, le hinstitiúidí ‘with institutes’.
Finally, (Emph) notes the emphatic marker in Irish. For example, liom ‘with me’;
liomsa ‘with me’.
Form=VF VF (Vowel Form) is an indicator of spelling changes that occur in
copular verbs when followed by a word that begins with a vowel or a lenited
consonant. For example, Is féidir liom ‘I can’; Conas ab fhéidir liom ‘How can I?’.
Mood=Int While regular verbs make use of interrogative particles (e.g. ar chuala
tú? ‘did you hear?’) to indicate a question construction, the copula has a number
of different forms to indicate the interrogative mood. For example Is maith leat
‘you like’; an maith leat? ‘do you like?’; ar mhaith leat? ‘would you like?’; nach
mhaith leat? ‘do you not like?’; nár mhaith leat? ‘would you not like?’
NounType=Weak, Strong Plural nouns are either referred to as Weak Plurals
or Strong Plurals. The form of a Strong Plural remains unchanged, regardless
of grammatical case. In other words it does not inflect. For example: Tá na
ríthe ag troid ‘the kings are fighting’ (Com.Pl); bás na ríthe ‘death of the kings’
(Gen.Pl). On the other hand, the plural form of a noun is weak if it meets
specific criteria related to the common plural form (The Christian Brothers [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] p.34).
The weak/strong nature of nouns in turn affects the declension of their modifying
adjectives. Therefore, this feature applies to both Nouns and Adjectives.
NounType=NotSlender, Slender Irish has broad (NotSlender) and slender
consonants. This feature is influenced by a consonant’s preceding vowel: a broad
consonant is preceded by a broad vowel and a slender consonant by a slender vowel.
In some cases, this consonantal feature can impact the spelling of subsequent
adjectives. For example, if a plural noun has a slender ending, the following adjective
is lenited (e.g. deonach ‘big’; eagrais dheonacha ‘voluntary organisations’). This
feature applies to Adjectives.
          </p>
          <p>PartType=Ad, Cmpl, Comp, Cop, Deg, Inf, Num, Pat, Vb, Voc Irish makes
use of a broad range of particles. It is important to differentiate these particles,
because in some cases they share the same form, yet have different functions. For
example, a can be a vocative particle, relative particle, infinitive particle or
quantifier particle, and trigger a variety of different spelling changes (e.g. eclipsis,
h-prefix or lenition).</p>
          <p>PrepForm=Cmp Compound prepositions contain a simple preposition and a
noun. Nouns that follow compound prepositions are inflected in the genitive case.
For example, ar fud ‘throughout’ and domhan ‘world’, when combined, becomes:
ar fud an domhain ‘throughout the world’.</p>
          <p>
            VerbForm=Cop Irish has two verbs ‘to be’ – the copular verb and the substantive
verb (Stenson [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ] p.92). The copula is used for (i) classification (‘he is a man’),
(ii) identification (‘Mary is the doctor’), (iii) to express ownership (iv) to mark
emphasis through clefting and (v) in making comparisons. The substantive is used
in all other cases. The copula does not inflect for mood, gender or number in the
way regular verbs do. As the UD POS tagset subsumes the copula as VERB, it
is important to distinguish between them through their morphological features, as
they follow a different argument structure. The substantive verb follows the general
VSO (Verb Subject Object) order, whereas the copula construction mostly follows
a Copula-Predicate-Subject order.
          </p>
          <p>
            Voice=Auto Irish does not have an equivalent to the English passive construction
(The Christian Brothers (1988, p.120) and Stenson [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ] p.145). Stenson identifies
autonomous verbs as one type of construction that is used instead. Autonomous
verbs have an ‘understood’ subject and therefore the noun which follows is usually
the direct object. For example, chonacthas iad ‘they were seen’ translates literally
as ‘someone saw them’. The dative case of the pronoun iad ‘them’ (vs siad ‘they’)
clearly marks it as the object of the verb.
5
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Experiments</title>
      <p>Treebank Features LAS (test) UAS (test)
IUDT (baseline) no 68.27 76.29
IUDT yes 70.63 77.57
IUDT + Opt yes 72.18 78.3
IUDT universal only 69.26 76.99
IDT no 69 78.33</p>
      <p>LAS (10-fold) UAS (10-fold)
IUDT 10-fold yes 72.19 79.12
LAS (dev) UAS (dev)
70.59 78.73
71.79 79.45
73.33 79.52
70.70 78.66
70.66 79.8</p>
      <p>
        In this section we discuss a number of parsing experiments carried out on the
IUDT. We use MaltParser (Nivre et al., [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]) for all our experiments.7 We first look
at how the inclusion of morphological features impacts parsing models trained on
the IUDT. We then follow on to optimise the inclusion of these features through
the use of MaltOptimizer. Following that, we demonstrate the impact of specifying
the Irish-specific features in our set. Then, we put our parsing results into context
by providing parsing results for the IDT. Finally, we provide parsing results based
on 10-fold cross-validation in order to provide a more realistic parsing accuracy
report.
      </p>
      <sec id="sec-6-1">
        <title>Evaluation of inclusion of morphological features</title>
        <p>
          In order to assess the benefit of including morphological features in the IUDT,
we report here on parsing experiments that are based on models trained on two
7MaltParser v1.7, available to download from http://www.maltparser.org/download.html
different versions of the treebank:
our baseline, which is the original IUDT (v1.0)8, (Lynn and Foster [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ])
– without morphological features
the newly updated IUDT (v1.4)9
– includes the morphosyntactic information (Section 2).
        </p>
        <p>In this experiment, we follow the development/test/training treebank split, as
per the UD convention, i.e. roughly 10%-10%-80%: dev (150 trees)/ test (150
trees)/ training (remaining = 720 trees). We can see from the results presented in
Table 2 that the inclusion of the morphological features in the IUDT leads to a clear
increase in parsing accuracy. As the only difference across the two experiments is
the presence or absence of morphological data (e.g. there is no change in annotation
scheme, standard treebank data or MaltParser settings), we can conclude that the
morphological information has a positive impact on parsing accuracy.
5.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Optimising feature inclusion</title>
        <p>
          With the promising results we get from including morphological information in
the training data, we are prompted to evaluate the use of MaltOptimizer
(Ballesteros [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]) as part of our efforts to fully optimise the inclusion of such rich
features. MaltOptimizer is a freely available tool that can be used in conjunction
with MaltParser to optimise parsing, based on an analysis of training data. The
analysis phase collects information such as the number of words/sentences,
percentage of projective/non-projective trees, existence of covered roots10, features in
LEMMA, FEATS, CPOS (coarse-grained part-of-speech) and POS (fine-grained
part-of-speech) columns. Based on this, the output provides the user with the most
effective combination of options (user-defined parameters) to use in the MaltParser
training and parsing phases.
        </p>
        <p>MaltOptimizer defined the optimal training options for the IUDT dataset as the
Nivre arc-eager parsing algorithm with normal root handling and a LIBLINEAR
learner classifier. As we show in Table 2, this configuration resulted in a LAS
increase of 1.55 and UAS increase of 0.73 on the test set. The improvement on the
results of the dev set were similar for LAS (1.54), but less so for UAS (0.07).
5.3</p>
      </sec>
      <sec id="sec-6-3">
        <title>Impact of Irish-specific morphological features</title>
        <p>In Section 4.1, we list and describe the Irish language-specific features that we have
introduced to our feature set. We acknowledge that we are motivated to specifying
these new features as we feel that this linguistic phenomena is integral to the syntax
of the Irish language and should be captured in the data. This intuition comes
8http://hdl.handle.net/11234/1-1464
9https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-1827
10Covered roots are cases where the root (head value = 0) is crossed by one or more arcs.
from a knowledge of the language, yet it is also important to empirically justify
this inclusion. For this reason, we experiment on a version of the IUDT treebank
that includes only the universal features, and excludes the Irish-specific features.
Our results are also presented in Table 2, where we can see that removing the
Irish specific features results in a drop in parsing accuracy, thus supporting the
motivation for their general inclusion in the latest version.</p>
        <p>On closer analysis of the parser output (development set), we observe exactly
how some of these features help to improve parsing. For example, the feature
Voice=Auto – which we explain in Section 4.1 indicates the autonomous verb –
appears to help the parser in correctly identifying the direct object (dobj) argument
of the verb. Without this feature, confusion arises when the parser assumes that the
noun following the autonomous verb is the subject, thus assigning the nsubj
label attachment. In some instances, correct identification of the verb–direct object
attachment resulted in a knock-on effect of improved parsing of the overall tree.
Likewise, we observe that the morphological feature PrepForm=Cmpd has helped
the parser to more accurately identify the correct dependency attachment and
labelling of case to compound prepositions, where previously parser confusion led
to the incorrect assignment of either root or nmod labels. Again, improvements in
this labelling had a knock-on effect in some cases, thus improving the overall parse
of the tree.
5.4</p>
      </sec>
      <sec id="sec-6-4">
        <title>Comparison with IDT parsing results</title>
        <p>To put the parsing results for Irish Universal Dependencies into context, we look
at comparing our IUDT parsing results with those of the original Irish Dependency
Treebank (IDT). It should be noted that the IDT does not yet contain morphological
features. The results are presented towards the bottom of Table 2.</p>
        <p>
          A number of factors can affect the quality of parsing, such as the design of
an annotation scheme, the number of dependency labels that are used (including
granularity of labels) and the settings used in a parsing framework. As reported
by Lynn and Foster [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], the IDT has 47 labels, compared to just 35 used in the
IUDT. In addition, due to the different annotation schemes, both the label names
and structural analyses differ across treebanks. It is therefore unsurprising that the
parsing results based on the IDT differ from those based on the IUDT. We remark
here that accuracy dropped across the board for the parsing model trained on the
IUDT. It is interesting to note that a scheme that is designed specifically for Irish
appears to be easier to parse than UD, which instead aims to be universal, and is
informed by a large collection of languages.
        </p>
        <p>We also note that adding morphological features to the IUDT increased parser
accuracy to a level which is on a par with the IDT parser (without morphological
features). This could suggest that the inclusion of morphological features
overcomes the potential loss of information introduced when mapping treebank
annotations from an Irish-language inspired scheme to a universal scheme. However,
it should be noted that strong conclusions cannot be drawn from this, as the
MaltParser settings chosen for these parsing experiments were optimised and tuned to
the IDT and may not necessarily be optimal for the IUDT.11
5.5</p>
      </sec>
      <sec id="sec-6-5">
        <title>Cross-validation</title>
        <p>It is worth noting that the baseline results provided here for the IDT are
actually lower than those previously reported by Lynn (2016) (LAS 71.4% and UAS
80.1%). The reason for this is the difference in the data splitting approach for the
dev/test/ train sets. As mentioned in Section 5.1, our IDT (and IUDT) results
presented here are based on the data split convention of the UD project. Prior reported
IDT experiments, however, have been carried out using k-fold cross-validation.</p>
        <p>While a conventional data-split approach across a large project is well
motivated, parsers trained and tested on small datasets may not always reflect the full
truth of attainable parsing accuracy. Instead, parsing model evaluation is often
carried out using k-fold cross-validation when working with smaller data sets. This
approach involves splitting the data into k sets of equal size, one set used as a test
set and the remaining as training data, and repeated iteratively until all partitions
have served as a test set once. The result is an average across k test set results.</p>
        <p>Thus, in order to acquire a more realistic set of results in our current
experiments, we perform 10-fold cross-validation on the IUDT with morphological
features to get a clearer idea of parsing quality. Our results show an increase across
both LAS (70.63% – 72.19%) and UAS (77.57% – 79.12%).
6</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>We have described the morphological features recently introduced to the Irish
Universal Dependency treebank, in the form of both standard UD features and features
introduced specifically for the Irish language. Parsing experiments with MaltParser
suggest that these morphological annotations are helpful, and results can be further
improved using MaltOptimizer. Ablation experiments demonstrate that the
Irishspecific morphological information has a useful role to play.
7</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>We would like to thank the anonymous reviewers for their valuable feedback and
Lamia Tounsi for her analysis advice. This work has been supported by the ADAPT
Centre for Digital Content Technology (www.adaptcentre.ie) at Dublin City
University, funded under the SFI Research Centres Programme (Grant 13/RC/2106)
and is co-funded under the European Regional Development Fund.</p>
      <p>11The MaltParser configurations on previously reported Irish parsing experiments included
stacklazy and liblinear learner algorithms. The feature set included word form, lemma and both coarse
and fine-grained POS tags.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <surname>Miguel</surname>
          </string-name>
          (
          <year>2012</year>
          )
          <article-title>MaltOptimizer: A System for MaltParser Optimization</article-title>
          ,
          <source>In Proceedings of the Eighth International Conference on Linguistic Resources and Evaluation (LREC)</source>
          , Istanbul,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Buchhloz</surname>
          </string-name>
          , Sabine and Marsi,
          <string-name>
            <surname>Erwin</surname>
          </string-name>
          (
          <year>2006</year>
          )
          <article-title>CoNLL-X shared task on Multilingual Dependency Parsing</article-title>
          ,
          <source>In Proceedings of the 10th Conference on Computational Natural Language Learning (CoNLL)</source>
          , New York City, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>The</given-names>
            <surname>Christian Brothers</surname>
          </string-name>
          (
          <year>1998</year>
          ) New Irish Grammar Dublin: CJ Fallon
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>ITÉ</surname>
          </string-name>
          (
          <year>2002</year>
          )
          <article-title>PAROLE Morphosyntactic Tagset for Irish. Institiúid Teangeolaíochta Éireann</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Lynn</surname>
          </string-name>
          , Teresa and Foster,
          <string-name>
            <surname>Jennifer</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <article-title>Universal Dependencies for Irish</article-title>
          .
          <source>In Proceedings of the Second Celtic Language Technology Workshop</source>
          , Paris, France,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Lynn</surname>
          </string-name>
          ,
          <string-name>
            <surname>Teresa</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <article-title>Irish Dependency Treebanking and Parsing</article-title>
          .
          <source>PhD Thesis</source>
          , Dublin City University and Macquarie University Sydney,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Lynn</surname>
          </string-name>
          , Teresa, Foster, Jennifer, Dras, Mark (
          <year>2013</year>
          )
          <article-title>Working with a small dataset - semi-supervised dependency parsing for Irish</article-title>
          .
          <source>In Proceedings of Statistical Parsing of Morphologically Rich Languages</source>
          , Seattle, USA,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Nivre</surname>
          </string-name>
          , Joakim, Hall, Johan, Nilsson, Jens (
          <year>2006</year>
          )
          <article-title>MaltParser: A DataDriven Parser-Generator for Dependency Parsing</article-title>
          .
          <source>In Proceedings of the Fifth International Conference on Language Resources and Evaluation</source>
          , Genoa, Italy,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Nivre</surname>
          </string-name>
          ,
          <string-name>
            <surname>Joakim</surname>
          </string-name>
          (
          <year>2015</year>
          )
          <article-title>Towards a Universal Grammar for Natural Language Processing</article-title>
          .
          <source>In Computational Linguistics and Intelligent Text Processing</source>
          , Springer International Publishing,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Nivre</surname>
            , Joakim, de Marneffe, Marie-Catherine, Ginter, Filip, Goldberg, Yoav, Hajic, Jan, Manning,
            <given-names>Christopher D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McDonald</surname>
          </string-name>
          , Ryan, Petrov, Petrov, Pyysalo, Sampo, Silveira, Natalia, Tsarfaty, Reut, Zeman, Daniel (
          <year>2016</year>
          ).
          <article-title>Universal Dependencies v1: A Multilingual Treebank Collection</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation</source>
          . Portoroz, Slovenia,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Ó</given-names>
            <surname>Siadhail</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mícheál</surname>
          </string-name>
          (
          <year>1989</year>
          ), Modern Irish:
          <article-title>Grammatical structure and dialectal variation</article-title>
          . Cambridge: Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Seddah</surname>
          </string-name>
          , Djamé, Kubler, Sandra, Reut, Tsarfaty (
          <year>2014</year>
          )
          <article-title>Introducing the SPMRL 2014 Shared Task on Parsing Morphologically-Rich Languages</article-title>
          .
          <source>In Proceedings of the first joint meeting of Statistical Parsing of Morphologically Rich Languages and Syntactic Analysis of Non-Canonical English</source>
          , Dublin, Ireland,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Seddah</surname>
          </string-name>
          , Djame and Tsarfaty, Reut and Kuebler, Sandra and Candito, Marie and Choi, Jinho and Farkas , Richard and Foster,
          <article-title>Jennifer and Goenaga, Iakes and Gojenola, Koldo and Goldberg, Yoav and Green, Spence and Habash, Nizar and Kuhlmann, Marco and Maier, Wolfgang and Nivre, Joakim and Przepiórkowski, Adam and Roth, Ryan and Seeker, Wolfgang and Versley, Yannick and Vincze, Veronika and Wolinski, Marcin and Wróblewska, Alina</article-title>
          and Villemonte de la Clérgerie,
          <source>Eric</source>
          (
          <year>2013</year>
          )
          <article-title>Overview of the SPMRL 2013 shared task: cross-framework evaluation of parsing morphologically rich languages</article-title>
          .
          <source>In Proceedings of the Fourth Workshop on Statistical Parsing of Morphologically Rich Languages</source>
          , Seattle, WA,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Stenson</surname>
          </string-name>
          ,
          <string-name>
            <surname>Nancy</surname>
          </string-name>
          (
          <year>1981</year>
          )
          <article-title>Studies in Irish Syntax</article-title>
          . Tübingen: Gunter Narr Verlag
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Uí</surname>
            <given-names>Dhonnchadha</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elaine</surname>
          </string-name>
          (
          <year>2009</year>
          )
          <article-title>Part-of-Speech Tagging and Partial Parsing for Irish using Finite-State Transducers and Constraint Grammar</article-title>
          ,
          <source>PhD Thesis</source>
          , Dublin City University,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Uí</surname>
            <given-names>Dhonnchadha</given-names>
          </string-name>
          , Elaine, van Genabith,
          <string-name>
            <surname>Josef</surname>
          </string-name>
          (
          <year>2006</year>
          )
          <article-title>A Part-of-speech tagger for Irish using Finite-State Morphology and Constraint Grammar Disambiguation</article-title>
          .
          <source>In Proceedings of the 5th International Conference on Language Resources and Evaluation (LREC</source>
          <year>2006</year>
          ), Genoa, Italy,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>