<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Dec</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Highway to Hell. Towards a Universal Dependencies Treebank for Dante Alighieri's Comedy</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Corbetta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Passarotti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Flavio Massimiliano Cecchini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Moretti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Although in earlier stages of linguistic research there were claims of similarity between Old Italian</institution>
          ,
          <addr-line>particu-</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Università Cattolica del Sacro Cuore</institution>
          ,
          <addr-line>largo A. Gemelli 1, 20123 Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Università degli studi di Bergamo</institution>
          ,
          <addr-line>via Salvecchio 19, 24129 Bergamo</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Università di Pavia</institution>
          ,
          <addr-line>corso Strada Nuova 65, 27100 Pavia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>02</volume>
      <issue>2023</issue>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>In this paper, we describe the creation in Universal Dependencies of a treebank for Dante's Comedy, the first syntactically annotated text for Old Italian following a dependency-based schema. We detail the phase of treebanking the first part of the Comedy, the Inferno, and we describe some annotation issues. Then, we perform an evaluation of automated dependency parsing with models trained on the currently available annotated portion of the text.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Dante Alighieri</kwd>
        <kwd>Old Italian</kwd>
        <kwd>treebank</kwd>
        <kwd>Universal Dependencies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>parison. Currently, the project boasts 245 treebanks for
141 languages,2 including historical languages such as
Over the past two decades, there has been a growing con- Ancient Greek, Latin, Old French, Akkadian and
Classivergence between the world of corpora for ancient lan- cal Chinese. With regard to the Italian language, there
guages and the scholarly community working in the area are 9 ud treebanks, covering a diverse range of genres,3
of technologies for Natural Language Processing (nlp). amounting to 879 657 tokens and 37 871 sentences.
Because of the absence of native speakers and newly writ- This paper details the process of developing a ud
treeten texts, dealing with ancient languages means lacking bank out of Dante’s Comedy, starting from the annotation
the possibility of introspective analysis or field inquiries. of the Inferno, the first out of the three parts ( cantiche)
The only empirical evidence historical linguists can en- of the work. The motivation for this is the current
abgage with is confined to old texts, many of which are sence of any dependency-based treebank for Old Italian.4
fortunately digitally available today. Enhancing these Besides providing the scholarly community of historical
data sources with meta-linguistic annotation provides linguistics with a valuable resource, we create gold data
scholars with enriched data to support their investiga- that can be used for the supervised training and testing
tions. Moreover, building annotated sets of textual data of stochastic nlp tools.
for an ancient language following de facto standards is This paper is organized as follows: in Section 2, we
a way to make these old texts compatible with several introduce Old Italian and the resources available for this
ready-made nlp tools, as well as to make them compara- language, with a specific focus on the DanteSearch
corble with annotated corpora for other (modern) languages. pus. In Section 3, we describe the creation of the treebank,</p>
      <p>
        Universal Dependencies1 (ud) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is an annotation starting from the Inferno. In Section 4, we describe
trainframework started in 2015 which aims to provide a uni- ing and evaluation of a number of models for parsing.
versal formalism for dependency-based syntactic anno- Section 5 concludes the paper by summarizing our
findtation, with the goal of facilitating cross-linguistic com- ings and sketching future work.
2ud version 2.12, May 2023 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
3Including “legal, news, wiki, nonfiction, government legal, social,
learner-essays and grammar-examples”. No literary texts have been
included thus far.
4Whereas, with regard to Dante Alighieri, his works in Latin are
already part of ud, see [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ].
larly Dante Alighieri’s vernacular, and Modern Italian,5 Strictly related to the the historical dictionary of Old
Italespecially when compared to the evolution of other Ro- ian built by ovi is the Tesoro della Lingua Italiana delle
mance languages like French, where diferences between Origini corpus (tlio) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which collects 3 173 texts for
old and modern varieties are more pronounced [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], nu- a total of 23 685 634 occurrences. Additionally, there are
merous studies have now recognized and emphasized the corpora that cover a wider temporal span, such as the
distinction between Old and Modern Italian [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], particu- midia corpus [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], a lemmatized and morphologically
anlarly from a syntactic perspective [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. notated collection of Italian texts from the 13th century
      </p>
      <p>
        The Grammatica dell’italiano antico (gia; ‘Grammar to the first half of the 20th century, and the codit corpus
of Old Italian’) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], defines Old Italian as the language [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], a diachronic corpus of Italian that covers the period
spoken in Florence during the 13th century and the early from the 13th century until 1947.
14th century. The authors of the gia justify their choice Although a preliminary efort has been made towards
of selecting Florentine texts (later expanded to texts from the creation of a digital corpus of Old Italian with respect
all the Tuscan region) on the basis of the abundant docu- to the quotations reported in the Grande dizionario della
mentation of vernacular scripta in Florence, driven also lingua italiana [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ],7 no dependency-based syntactic
anby the diligence and productivity of the Florentine scribes. notation of Old Italian texts is currently available.
However, it should be noted that there are numerous
written varieties that characterize Medieval Italy, albeit in a 2.2. DanteSearch
minority when it comes to documentation and written
evidence. Among the resources available for Old Italian,
Dante
      </p>
      <p>
        Regardless of whether Old Italian should be strictly Search (ds) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is an annotated corpus containing all of
limited to the Tuscan area or can also encompass non- Dante Alighieri’s works, including both the Latin and
Tuscan varieties, the significance and influence of Tuscan the vernacular texts. The resource has been developed by
on the evolution of the Italian language is undeniable. the University of Pisa and consists of a set of
(downloadTherefore, while choosing an Old Italian text for a ud able)8 xml files providing both textual data and linguistic
treebank, it seems obvious to select a Tuscan text, specif- annotation.
ically a Florentine one, namely the Comedy of Dante Concerning the Comedy, the text included in ds is
Alighieri. based on Petrocchi’s edition [17] and is recorded in two
      </p>
      <p>
        Dante Alighieri was born in Florence in 1265 and he separate xml files: one file provides the grammatical
is legitimately considered one of the greatest poets and layer of annotation (featuring tokens, lemmas, and tags
writers of the Middle Ages. His most important work is representing both parts of speech and morphosyntactic
the Comedy, which was written between 1308 and 1320, features), while the other contains a clause-based layer
and is crucial to Italian literature, due to its historical of syntactic annotation [18].
(and still continuing) success among readers, and rele- The clause-based annotation of syntax distinguishes
vance among scholars. The decision of Dante to write main and subordinate clauses, the latter being assigned a
the Comedy in the Florentine vernacular represents a label for their function, such as “declarative”, “temporal”,
pivotal moment in the history of Italian literature and and “relative” [19].
language, as it contributed to spreading and elevating
the vernacular to a literary language [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. 3. Treebanking Dante’s Comedy :
      </p>
      <p>
        Together with the undeniable significance of the text,
the availability of a digital resource, DanteSearch [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the Inferno
containing all of Dante’s works enhanced with a number
of fundamental layers of annotation, further supports
our decision to choose the Comedy as the text for the
ifrst ud treebank of Old Italian.
      </p>
      <sec id="sec-1-1">
        <title>Dante’s Comedy is composed of three parts, called can</title>
        <p>tiche, which are Inferno ‘Hell’, Purgatorio ‘Purgatory’ and
Paradiso ‘Heaven’. These cantiche are divided
respectively into 34, 33 and 33 subsections called canti. This
Section details the process of annotating the Inferno
according to ud’s formalism.</p>
        <sec id="sec-1-1-1">
          <title>2.1. Resources for Old Italian</title>
          <p>There is quite a substantial amount of texts and lexical
resources in digital format available for Old Italian. Among
them, the Opera del Vocabolario Italiano corpus6 (ovi)
contains Old Italian texts dating before the 15th century
and is one of the major corpora, containing 3 443 texts
of Old Italian for a total of 30 176 628 word occurrences.
5As exemplified by a statement by [5, p. 124], cf. [6, ch. vi].
6http://www.ovi.cnr.it/Il-Corpus-Testuale.html
7The work by Favaro consists of a conversion from an xml source
ifle to the CoNLL-U format adopted in ud, for tokenization,
lemmatization, and morphological annotation.
8https://dantesearch.dantenetwork.it.</p>
          <p>no. tokens
lemma(s)
tag(s)</p>
        </sec>
        <sec id="sec-1-1-2">
          <title>3.1. From DanteSearch to ud</title>
          <p>In ds, the Inferno consists of 33 416 tokens out of a total mentre che
of 99 390 (without punctuation marks). ADV SCONJ</p>
          <p>We perform a conversion from the grammatical xml Table 1
ifle of the Inferno provided by ds to the CoNLL-U format Example of locution mentre che
adopted by ud’s treebanks.9 The conversion focuses on
tokens (i. e. forms), lemmas, parts of speech (PoS), and
morphological features. However, in the CoNLL-U file ds ud
we do not report the syntactic annotation contained in no. tokens 1 2
wthiethxtmhle swyonrtda-cbtaicsefilde uodf sydnst,adcutiec taonaitlsysiinsc[o1m,§p2a.2ti]b.ility lemtamg(as)(s) FilipponArgenti FPiRliOpPpNo APrRgOePnNti</p>
          <p>The conversion of tags happens on a 1:1 basis (ds:ud) Table 2
whenever possible. Diferent criteria for the assignment Example of the proper noun Filippo Argenti
of PoS and morphological tags between the two
annotation styles are managed case by case. For instance, ds
alternately assigns the tag for “pronouns” (p) or “adjec- Further, we also want to adjust the lemmatization of
tives” (a) to possessives such as mio ‘my’, while in ud we articles. In ds, there are separate lemmas la/una and
always tag them as “determiners” (DET). il/uno for the definite/indefinite feminine and masculine</p>
          <p>With regard to tokenization and lemmatization, in a articles respectively, whereas, following the convention
few cases we modify the criteria followed by ds to fit the of most ud Italian treebanks, we lemmatize both under
ones of ud. Specifically, this applies to the tokenization the respective masculine forms.
and lemmatization of what are referred to as locuzioni
‘locutions’ in ds, i. e. sets of two or more words arranged 3.2. Syntactic annotation
in a fixed sequence [ 20], such as mentre che ‘while’ and
davanti a ‘in front of’. In ds, such multiword expressions We perform the syntactic annotation of the Inferno
manare analyzed as single tokens, while the ud annotation ually13 using ConlluEditor [22] and with the support of
schema requires that the words they are composed of be a few critical commentaries on the work, namely those
analyzed individually and considered as separate tokens. by Chiavacci Leonardi [23] and Inglese [24]. Following
As a consequence, for locutions we employ a distinct the ud guidelines, annotation is made at sentence level;
tokenization, lemmatization, and PoS tagging in contrast we base sentence splitting on full stops and question or
to ds, as shown in Table 1 with regard to the following exclamation marks followed by an uppercase letter,
acexample:10 cording to Petrocchi’s edition of the Comedy recorded in
ds [17].</p>
          <p>Inferno, v, vv. 95–96 A sentence corresponds to a syntactic tree, i. e. an
noi udiremo e parleremo a voi, / mentre acyclic, oriented, rooted graph [25], whose nodes
corche ’l vento, come fa, ci tace. respond to tokens in the text.14 Nodes are related to each
other through dependencies, i. e. hierarchical binary
relations, which are labeled with a syntactic function, such as
‘will please us, too, to hear and speak with
you, / now while the wind is silent, in this
place.’11</p>
          <p>Modifications of lemmatization and PoS tagging are
required also for multiword proper nouns, which are
lemmatized under a unique lemma in ds in contrast to
ud. Table 2 shows the example of the multiword proper
name Filippo Argenti:12
9CoNLL-U is a format with tab-separated values where lines
contain the annotation of tokens into 10 fields; see https://
universaldependencies.org/format.html.
10In this example, the ds tag clst stands for a subordinating
conjunction (cs) used in a locution (l) within a temporal clause (t),
while the ud PoS tags ADV and SCONJ stand respectively for
“adverb” and “subordinating conjunction”.
11The English translations of the examples from the Comedy are by</p>
          <p>Allen Mandelbaum, available at: https://digitaldante.columbia.edu/
dante/divine-comedy/.
12The ds tag n stands for “onomastics” and the ud tag PROPN stands</p>
          <p>for “proper noun”.
13The syntactic annotation is performed by a single annotator with
expertise in Italian studies. Annotating pre-parsed data has been
ruled out after evaluating the accuracy of the UDPipe model [21]
trained on the largest ud treebank of Italian (isdt) and tested on
the first three canti of Inferno: its LAS score is 63,52% (see Section
4).
14In ud, a distinction between “token” and “syntactic word” is made:
while “token” refers to an orthographic unit of segmentation,
“syntactic word” refers to the actual level of analysis in the syntactic
tree. These two levels often, but not always, coincide, e. g. the
token nel ‘in the’ would be analyzed into the syntactic words in
‘in’ and il ‘the’, each bearing its own annotation. Refer to https:
//universaldependencies.org/u/overview/tokenization.html. In this
paper, the term “token” will be used throughout as an equivalent
to ud’s “syntactic word”.
nsubj for “nominal subject”. Dependency-based
annotation schemes are predicate-centered, with the sentence’s
main predicate serving as the tree’s root. In ud’s
formalism, function words depend on the content words they
modify.15</p>
          <p>While annotating the Inferno according to the ud
formalism, we encounter several issues that require taking
specific decisions. In the following, we discuss the
annotation of ellipses and comparative clauses.</p>
          <p>The total number of sentences in the Inferno is 1 228,
for a total of 41 367 tokens.
3.2.1. Ellipsis
“Ellipsis” refers to the omission of words or phrases that
can be inferred from the context of a sentence or
utterance.16 While annotating the Inferno, we encounter
several cases of ellipses, including nominal ellipses, i. e. [27,
p. 526]:
diferent types of anaphoric phenomena
involving a gap within the internal
structure of the nominal phrase.
and predicate ellipsis [28, p. 504]:
a type of ellipsis that leaves the main
predicate of the clause unpronounced, most
often together with one or more of its
internal arguments or (low) adjuncts.</p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>In the matter of nominal ellipsis, we follow the solution</title>
        <p>of promotion, as outlined in the ud guidelines.17 We
present here an example of nominal ellipsis (Figure 1):
Inferno, ix, vv. 28–29:
Quell’è ’l più basso loco e ’l più oscuro / e
’l più lontan dal ciel che tutto gira
‘That is the deepest place and the darkest
place, / the farthest from the heaven that
girds all’
where oscuro ‘dark’ and lontan ‘far’ depend on the omitted
noun (NOUN) loco ‘place’, as shown by the repetition of
the definite article ( DET) ’l ‘the’, which modifies the noun.
In this case, we promote the adjectives (ADJ) oscuro and
lontan to heads of their respective coordinate clauses
using the dependency relation conj “conjunct”.</p>
        <p>Following to the ud guidelines, we handled predicate
ellipsis by using the dependency relation orphan
(orphan relation), like in the following example:
15This is not the case for all dependency-based schemes, like for
instance for the analytical layer of annotation of the Prague
Dependency Treebank for Czech (pdt), where e. g. conjunctions govern
conjuncts and adpositions are the heads of adpositional phrases.
Refer to https://ufal.mf.cuni.cz/pdt2.0/doc/manuals/en/a-layer/
where the predicate of the sentence, namely the
verbum dicendi, is omitted. This structure is extremely
common to introduce a reported speech. As shown in Figure
2, the omission of the predicate requires promoting the
subject of the sentence, elli “he”, to the root of the tree
(root) and annotating the underlying oblique relation
(obl) of the phrase a lui “to him” with an orphan relation
(orphan).</p>
        <p>Currently, the syntactic annotation in UD handles
these cases of ellipsis with the promotion mechanism,
which involves promoting an element to function as the
omitted element in the sentence and replacing it in its
dependency relation without explicitly signaling this
omission, and the use of the orphan dependency relation,
whose function is to indicate that the element subject to
html/index.html.
16See [26] for an introduction to the topic.
17Promotion involves selecting an element to take the place of the
omitted element in the syntactic tree, following a specific
hierarchy. Promotion is used without explicitly signaling the ellipsis.
See ud guidelines: https://universaldependencies.org/u/overview/
specific-syntax.html#ellipsis.
the orphan relation does not have an overt dependent
element in the syntactic structure.
3.2.2. Comparative clauses</p>
      </sec>
      <sec id="sec-1-3">
        <title>In the Inferno, we find a diverse usage of comparative</title>
        <p>clauses, ranging from sentences where the comparative
clause is longer than the main clause it depends on, to
others where comparatives consist of just a few tokens.
In light of the long-lasting discussion on the treatment
of comparative clauses in ud,18 we annotate such clauses
by labeling their head tokens with the dependency
relation advcl “adverbial clause modifier” specified for the
subtype cmp for comparative clauses.19</p>
        <p>A number of issues concerns cases of clauses where
our annotation, following the ud framework, parts from
the interpretation provided by ds, like in the following
example (Figure 3):</p>
      </sec>
      <sec id="sec-1-4">
        <title>Inferno, vi, v. 19:</title>
      </sec>
      <sec id="sec-1-5">
        <title>Urlar li fa la pioggia come cani</title>
        <p>‘That downpour makes the sinners howl
like dogs’</p>
      </sec>
      <sec id="sec-1-6">
        <title>In the annotation of ds, the portion come cani ‘like</title>
        <p>dogs’ is considered a phrase that is part of a declarative
clause. Instead, we consider come cani as a comparative
clause with an elliptical predicate, namely Urlar li fa la
pioggia come [fa urlare] i cani ‘That downpour makes the
sinners howl like [it makes] dogs [howl]’.</p>
        <p>We observe a few cases where a come-clause can be
considered either a comparative clause or a secondary
predication. In such cases, we rely on the interpretation
provided by commentaries, like in the following sentence
(Figure 4):</p>
        <p>In this sentence, the come-clause can be interpreted
either as a secondary predication (therefore, annotated
using the subtyped relation advcl:pred20), ‘He, being
a comprehensive person, answered to me’, or as a
comparative clause (with subtyped relation advcl:cmp), ‘He
answered to me like a comprehensive person’. In this
case, we follow the interpretation of Chiavacci Leonardi
[23] in considering the come-clause as a comparative
construction.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Evaluation</title>
      <sec id="sec-2-1">
        <title>We use the manually annotated Inferno to train models</title>
        <p>with UDPipe 121 [29] and to assess their performances
in view of employing them for parsing Purgatorio and
Paradiso, so as to facilitate their subsequent manual
annotation.22 In our evaluation framework, we employ a
cross-validation based on 10%/90% splits of the data: each
test set will then consist of approximately 4 137 out of
41 367 tokens and 123 out of 1 228 sentences, while train
sets of approximately 37 230 tokens and 1 105 sentences.</p>
        <p>The evaluation of the models’ accuracies is performed
by measuring Labeled (las) and Unlabeled Attachment
Score (uas) [30].</p>
        <p>The training and evaluation process is based on one
eleven- and one tenfold partition of the data, for a total
of 11+10 iterations: the first partition patterns upon the
original division into canti, with batches of 3 consecutive
20Cf. documentation at https://universaldependencies.org/la/dep/
advcl-pred.html (for Latin).
18Cf. the discussion group on comparatives in ud: https:// 21https://github.com/ufal/udpipe.</p>
        <p>universaldependencies.org/workgroups/comparatives.html. 22We acknowledge that doing tests within a single cantica may not
19Cf. documentation at https://universaldependencies.org/la/dep/ guarantee the same performances when compared to other
canadvcl-cmp.html (for Latin). tiche.</p>
        <p>Partition
random
consecutive
random
consecutive</p>
        <p>Scenario
+Morph
+Morph
-Morph
-Morph</p>
        <p>Avg. uas
81,95± 0,94%
81,79± 1,38%
75,32± 0,91%
74,90± 1,37%
canti23 assigned to the test set and the remaining 3124 across the text, as also standard deviation is very low.
forming the training set; the second partition is obtained On the other hand, las and uas metrics improve
signifiby fully random selection of sentences.25 Moreover, eval- cantly when the text is already enriched with linguistic
uation is carried out according to two scenarios: one annotation. This allows us to have positive expectations
(+Morph) in which lemmas, parts of speech and morpho- with regard to the parsing of Purgatorio and Paradiso,
logical features are given, and one (-Morph) in which cantiche for which lemmatization and morphosyntactic
every annotation level has to be tagged from scratch.26 taggings are inherited from the conversion from ds.</p>
        <p>The accuracy of each model is calculated using
eval.py,27 an evaluation script provided by the UD
project. As shown in Table 3, evaluations conducted 5. Conclusions and future
on the random partition result into slightly higher av- perspectives
erage accuracy scores than those based on triplets28 of
consecutive canti: in the +Morph scenario, a diference Building a ud treebank for Dante’s Comedy is the first
of 0,16% is observed for uas, whereas in the opposite step towards incorporating Old Italian among the
lan-Morph scenario the improvement is more marked, but guages of ud. This paper describes the development of
still minor, at 0,42% for uas and 0,26% for las. The only the first part of this treebank, which consists of the first
exception regards las in the +Morph scenario, though cantica of the Comedy, the Inferno.
the diference of 0,02% encountered there is negligible. We also present the results of an experiment of
su</p>
        <p>Consistently with our expectations, we also observe pervised automated dependency parsing using both as
that parsing performed with prior assignment of the training and test sets data from the Inferno. We run this
other annotation levels produces better results compared experiment to understand to what extent the process of
to the case where the parser has to handle all annota- syntactic annotation of the Comedy, which has been
pertion levels simultaneously. Specifically, in the +Morph formed so far fully manually, can benefit from the results
scenario the average of models trained on the random of the application of an nlp tool. Although the accuracy
partition exhibits an improvement of 6,63% for uas and rates reported in the paper are fairly good (≈ 77% las), in
9,10% for las , and similarly models trained on consec- the near future we will have to evaluate how and to what
utive canti show an improvement of 6,89% for uas and extent they will drop once a model trained and evaluated
9,38% for las. on the Inferno is applied to a diferent cantica. Should the</p>
        <p>We can conclude that, on the one hand, sampling the accuracy rates drop heavily, even such a negative result
dataset randomly or by selecting consecutive parts of the might prove helpful in pointing out syntactic diferences
text does not seem to significantly afect performances, between the three cantiche. Moreover, the use of other
and this could point to the fact that, at least in this cantica, parsers, based on diferent algorithms and resources (like
morphosyntactic phenomena are uniformly distributed embeddings), might lead to better and, most importantly,
diverging results and errors.
23We actually note that, since the number of canti, 34, is not divisible As for annotation issues, we will suggest to introduce
by 3, one canto would be left out, and is instead aggregated to the a specific subtype, e. g. ellp, in ud’s documentation, so
last batch, which then consists of 4 consecutive canti (31, 32, 33, as to properly identify cases of ellipses, as they are not
34). explicitly captured by the current annotation strategies
24Or 30; see fn. 23. mentioned in the paper, namely promotion and the use
25IPnlefearsneor_etfreeretboatnhke fGoirtHthuebdpaatageanhdttpdse:t/a/gilietdhusbta.ctoismti/cCsloanudtihaeCpoarbrtei/- of the relation orphan: the former does not signal the
tions. presence of ellipsis, while the latter obscures the real
26Corresponding respectively to –parse and –tag –parse options dependency relations which are replaced by it. While
for UDPipe; see https://ufal.mf.cuni.cz/udpipe/1/users-manual, adopting a subtype like ellp would make it possible to
27§h3tt.6p.s://github.com/UniversalDependencies/tools/blob/master/ collect cases of ellipses, their resolution is up to the
annoeval.py. tation of so-called enhanced dependencies, which are a
28Or a quadruplet; see fn. 23. kind of advanced annotation that augments dependency
labels to facilitate disambiguation.29</p>
        <p>We plan to engage additional annotators with
expertise in Old Italian to expedite the process of annotation
of Purgatorio and Paradiso. Additionally, we intend to
apply error detection processes (like, for instance, those
described in [31]) to retrieve possible mistakes or
inconsistencies in syntactic annotation.</p>
        <p>Another task we intend to address is the extension of
the ud documentation for Italian in order to make the
validator30 correctly deal with some peculiarities of Old
Italian, like for instance enclitic adpositions (e.g., meco
’with me’), which require the introduction of the feature
Clitic=Yes combined with the PoS tag ADP, currently
permitted only with the PoS tag PRON.</p>
        <p>Finally, we plan to include enhanced dependencies 31
in the ud treebank of Dante’s Comedy, once the basic
syntactic annotation of the entire work will be completed.
29https://universaldependencies.org/u/overview/enhanced-syntax.</p>
        <p>html.
30https://github.com/UniversalDependencies/tools/blob/master/</p>
        <p>validate.py
31https://universaldependencies.org/u/overview/enhanced-syntax.</p>
        <p>html
Association (elra), Marseille, France, 2022, pp. 94– Matematicko-fyzikální fakulta, Prague, Czech
Re100. URL: https://aclanthology.org/2022.lt4hala-1. public, 2007. URL: https://dspace.cuni.cz/handle/20.
13/. 500.11956/12614?locale-attribute=en.
[17] D. Alighieri, La Commedia secondo l’antica vul- [26] J. Merchant, Ellipsis: A survey of analytical
gata voll. i–iv, number 7 in Edizione nazionale delle approaches, in: J. van Craenenbroek, T.
TemOpere di Dante Alighieri a cura della Società Dan- merman (Eds.), The Oxford Handbook of
Elliptesca Italiana, Le Lettere, Florence, Italy, 1994. URL: sis, Oxford Handbooks, Oxford University Press,
https://www.lelettere.it/libro/9788871661483, edi- Oxford, uk, 2019. URL: https://academic.oup.com/
tor: Giorgio Petrocchi. edited-volume/41718/chapter/353990361.
[18] M. Tavoni, Allestimento, fruizione e prospettive [27] A. Saab, Nominal Ellipsis, in: J. van
Craedi DanteSearch, in: E. Cresti, M. Moneglia (Eds.), nenbroek, T. Temmerman (Eds.), The Oxford
Corpora e Studi Linguistici. Atti del liv Congresso Handbook of Ellipsis, Oxford Handbooks, Oxford
della Società di Linguistica Italiana (Online, University Press, Oxford, uk, 2019, pp. 526–561.
8-10 settembre 2021), number 6 in nuova serie, URL: https://academic.oup.com/edited-volume/
Oficinaventuno, Milan, Italy, 2022, pp. 255–273. 41718/chapter-abstract/353995808. doi:10.1093/
URL: https://www.societadilinguisticaitaliana. oxfordhb/9780198712398.013.26.
net/wp-content/uploads/2022/11/017_ [28] A. Lobke, W. Harwood, Predicate Ellipsis,
Tavoni_Atti_LIV_Congresso_SLI.pdf . in: J. van Craenenbroek, T. Temmerman (Eds.),
doi:10.17469/O2106SLI000017. The Oxford Handbook of Ellipsis, Oxford
Hand[19] S. Gigli, La codifica sintattica della Commedia di books, Oxford University Press, Oxford, uk,
Dante, in: M. D’Amico (Ed.), Sintassi dell’italiano 2019, pp. 504–525. URL: https://academic.oup.com/
antico e sintassi di Dante. Atti del seminario di studi edited-volume/41718/chapter-abstract/353995176.
(Pisa 15/16 ottobre 2011), Felici, Ghezzano (pi), Italy, [29] M. Straka, J. Straková, Tokenizing, POS Tagging,
2015, pp. 81–95. Lemmatizing and Parsing ud 2.0 with UDPipe, in:
[20] L. Serianni, A. Castelvecchi, Grammatica italiana, J. Hajič, D. Zeman (Eds.), Proceedings of the CoNLL
Universitaria, second ed., utet Università, Turin, 2017 Shared Task: Multilingual Parsing from Raw
Italy, 2006. Text to Universal Dependencies, Association for
[21] M. Straka, J. Hajič, J. Straková, UDPipe: Trainable Computational Linguistics (acl), Vancouver, bc,
Pipeline for Processing CoNLL-U Files Performing Canada, 2017, pp. 88–99. URL: https://aclanthology.
Tokenization, Morphological Analysis, POS Tag- org/K17-3009. doi:10.18653/v1/K17-3009.
ging and Parsing, in: N. Calzolari, K. Choukri, T. De- [30] S. Buchholz, E. Marsi, CoNLL-X Shared Task on
clerck, S. Goggi, M. Grobelnik, B. Maegaard, J. Mar- Multilingual Dependency Parsing, in: L. Màrquez,
iani, H. Mazo, A. Moreno, J. O. Odijk, S. Piperidis D. Klein (Eds.), Proceedings of the Tenth
Confer(Eds.), Proceedings of the Tenth International Con- ence on Computational Natural Language Learning
ference on Language Resources and Evaluation (CoNLL-X), Association for Computational
Linguis(lrec’16), European Language Resources Associa- tics (acl), New York City, nj, usa, 2006, pp. 149–164.
tion (elra), Portorož, Slovenia, 2016, pp. 4290–4297. URL: https://aclanthology.org/W06-2920.</p>
        <p>URL: https://aclanthology.org/L16-1680. [31] M. Dickinson, Error Detection and
Cor[22] J. Heinecke, ConlluEditor: a fully graphical edi- rection in Annotated Corpora, Ph.D.
thetor for Universal Dependencies treebank files, in: sis, The Ohio State University, 2005. URL:
A. Rademaker, F. Tyers (Eds.), Proceedings of the https://sifnos.sfs.uni-tuebingen.de/decca/
Third Workshop on Universal Dependencies (udw, publications/dickinson-dissertation.html.
SyntaxFest 2019), Association for Computational
Linguistics (acl), Paris, France, 2019, pp. 87–93.</p>
        <p>URL: https://aclanthology.org/W19-8010. doi:10.</p>
        <p>18653/v1/W19-8010.
[23] D. Alighieri, Inferno, number 613 in Oscar
classici, Arnoldo Mondadori, Milan, Italy, 2005. Editor:</p>
        <p>Anna Maria Chiavacci Leonardi.
[24] D. Alighieri, Commedia. Inferno, number 1 in</p>
        <p>Opere, Carocci, Rome, Italy, 2007. Editor: Guglielmo</p>
        <p>Inglese.
[25] J. Havelka, Mathematical Properties of
Dependency Trees and their Application to Natural
Language Syntax, Ph.D. thesis, Univerzita Karlova –</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>M.-C. de Marnefe</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Nivre</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Zeman</surname>
          </string-name>
          , Universal Dependencies,
          <source>Computational Linguistics</source>
          <volume>47</volume>
          (
          <year>2021</year>
          )
          <fpage>255</fpage>
          -
          <lpage>308</lpage>
          . URL: https://direct.mit.edu/coli/article/ 47/2/255/98516/Universal-Dependencies. doi:
          <volume>10</volume>
          .1162/coli_a_
          <fpage>00402</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Zeman</surname>
          </string-name>
          , et alii,
          <source>Universal dependencies 2.12</source>
          ,
          <year>2023</year>
          . URL: http://hdl.handle.net/11234/1-5150,
          <article-title>LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL)</article-title>
          ,
          <source>Faculty of Mathematics and Physics</source>
          , Charles University. Available at http://hdl.handle.net/11234/1-
          <fpage>5150</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Cecchini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          , G. Moretti, M. Passarotti,
          <article-title>UDante: First Steps Towards the Universal Dependencies Treebank of Dante's Latin Works</article-title>
          , in: J.
          <string-name>
            <surname>Monti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Dell'Orletta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Tamburini</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the Seventh Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2020</year>
          , Bologna, Italy, March 1-3
          <year>2021</year>
          ),
          <article-title>Associazione italiana di linguistica computazionale (ailc</article-title>
          ), Accademia University Press, Turin, Italy,
          <year>2020</year>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>105</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2769</volume>
          /paper_14.pdf .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Passarotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Cecchini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          , G. Moretti, UDante, Studi Danteschi lxxxvi (
          <year>2022</year>
          )
          <fpage>309</fpage>
          -
          <lpage>338</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G. I.</given-names>
            <surname>Ascoli</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          '
          <article-title>Italia dialettale, Archivio glottologico italiano viii (1882-</article-title>
          <year>1885</year>
          )
          <fpage>98</fpage>
          -
          <lpage>128</lpage>
          . Available at https://archive.org/details/ archivioglottolo08fireuoft/page/n5/mode/2up.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Tomasin</surname>
          </string-name>
          , Il caos e l'ordine, Piccola Biblioteca Einaudi, Giulio Einaudi, Turin, Italy,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dardano</surname>
          </string-name>
          (Ed.),
          <article-title>Sintassi dell'italiano antico</article-title>
          , Lingue e Letterature Carocci, Carocci, Rome, Italy,
          <year>2013</year>
          . URL: https://www.carocci.it/prodotto/ sintassi-dellitaliano-antico.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dardano</surname>
          </string-name>
          , G. Frenguelli (Eds.), SintAnt, Aracne, Rome, Italy,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tesi</surname>
          </string-name>
          ,
          <article-title>Parametri sintattici per la definizione di "Italiano antico"</article-title>
          , in: M.
          <string-name>
            <surname>Dardano</surname>
          </string-name>
          , G. Frenguelli (Eds.),
          <source>SintAnt. La sintassi dell'italiano antico, Aracne</source>
          , Rome, Italy,
          <year>2004</year>
          , pp.
          <fpage>425</fpage>
          -
          <lpage>444</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salvi</surname>
          </string-name>
          , L. Renzi (Eds.),
          <article-title>Grammatica dell'italiano antico</article-title>
          , il Mulino, Bologna, Italy,
          <year>2010</year>
          . URL: https: //www.mulino.it/isbn/9788815134585.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Manni</surname>
          </string-name>
          ,
          <article-title>La lingua di Dante, Le vie della civiltà</article-title>
          , il Mulino, Bologna, Italy,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tavoni</surname>
          </string-name>
          ,
          <article-title>DanteSearch: il corpus delle opere volgari e latine di Dante lemmatizzate con marcatura grammaticale e sintattica</article-title>
          , in: A.
          <string-name>
            <surname>Cerbo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Mondola</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Žabjek</surname>
          </string-name>
          , C. D. Fiore (Eds.),
          <source>Lectura Dantis</source>
          <year>2002</year>
          -2009.
          <article-title>Omaggio a Vincenzo Placella per i suoi settanta anni</article-title>
          , volume
          <volume>2</volume>
          (
          <fpage>2004</fpage>
          -2005), Il Torcoliere - Oficine
          <string-name>
            <surname>Grafico-Editoriali di</surname>
            <given-names>Ateneo</given-names>
          </string-name>
          , Naples, Italy,
          <year>2011</year>
          , pp.
          <fpage>583</fpage>
          -
          <lpage>608</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P. G.</given-names>
            <surname>Beltrami</surname>
          </string-name>
          ,
          <article-title>Il Tesoro della Lingua Italiana delle Origini (tlio)</article-title>
          , in: N.
          <string-name>
            <surname>Maraschio</surname>
            ,
            <given-names>T. Poggi</given-names>
          </string-name>
          <string-name>
            <surname>Salani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bongi</surname>
          </string-name>
          , M. Palmerini (Eds.),
          <article-title>Italia linguistica anno mille. Italia linguistica anno duemila</article-title>
          .
          <source>Atti del xxxiv Congresso Internazionale di Studi della Società di Linguistica Italiana, Firenze</source>
          <volume>19</volume>
          -21 ottobre
          <year>2000</year>
          ,
          <article-title>number 45 in Società di linguistica italiana</article-title>
          , Bulzoni, Rome, Italy,
          <year>2003</year>
          , pp.
          <fpage>695</fpage>
          -
          <lpage>698</lpage>
          . URL: https://www.torrossa.com/it/catalog/preview/ 2280318. doi:
          <volume>10</volume>
          .1400/28371.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>P. D'Achille</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          Grossmann (Eds.),
          <article-title>Per la storia della formazione delle parole in italiano</article-title>
          , Quaderni della Rassegna, eighth ed.,
          <string-name>
            <surname>Franco</surname>
            <given-names>Cesati</given-names>
          </string-name>
          , Florence, Italy,
          <year>2017</year>
          . URL: https://www.francocesatieditore.com/catalogo/ per-la
          <article-title>-storia-della-formazione-delle-parole-in-italiano/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Micheli</surname>
          </string-name>
          ,
          <string-name>
            <surname>codit.</surname>
          </string-name>
          <article-title>A new resource for the study of Italian from a diachronic perspective: Design and applications in the morphological field</article-title>
          ,
          <source>Corpus</source>
          <volume>23</volume>
          (
          <year>2022</year>
          ). URL: https://journals.openedition. org/corpus/7306. doi:
          <volume>10</volume>
          .4000/corpus.7306.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Favaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Guadagnini</surname>
          </string-name>
          , E. Sassolini,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bifi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          ,
          <article-title>Towards the Creation of a Diachronic Corpus for Italian: A Case Study on the gdli Quotations</article-title>
          , in: R. Sprugnoli, M. Passarotti (Eds.),
          <source>Proceedings of the Second Workshop on Language Technologies for Historical and Ancient Languages (lt4hala)</source>
          ,
          <source>European Language Resources</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>