<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Annotation of Liber Abbaci, a Domain-Specific Latin Resource</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francesco Grotto</string-name>
          <email>francesco.grotto1@sns.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rachele Sprugnoli</string-name>
          <email>rachele.sprugnoli@unicatt.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Margherita Fantoli</string-name>
          <email>margherita.fantoli@kuleuven.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Simi</string-name>
          <email>maria.simi@unipi.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Flavio Massimiliano Cecchini</string-name>
          <email>flavio.cecchini@unicatt.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Passarotti</string-name>
          <email>marco.passarotti@unicatt.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>. Scuola Normale Superiore</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. KU Leuven</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>. Università Cattolica del Sacro Cuore</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>. Università degli Studi di Pisa</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2003</year>
      </pub-date>
      <abstract>
        <p>The Liber Abbaci (13th century) is a milestone in the history of mathematics and accounting. Due to the late stage of Latin, its features and its very specialized content, it also represents a unique resource for scholars working on Latin corpora. In this paper we present the annotation and linking work carried out in the frame of the project Fibonacci 1202-2021. A gold-standard lemmatization and part-ofspeech tagging allow us to elaborate some ifrst observations on the linguistic and historical features of the text, and to link the text to the Lila Knowledge Base, that has as its goal to make distributed linguistic resources for Latin interoperable by following the principles of the Linked Data paradigm. Starting from this specific case, we discuss the importance of annotating and linking scientific and technical texts, in order to (a) compare and search them together with other (non-technical) Latin texts (b) train, apply and evaluate NLP resources on a non-standard variety of Latin. The paper also describes the fruitful interaction and coordination between NLP experts and traditional Latin scholars on a project requiring a large range of expertise.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Latin texts have a wide diachronic and diatopic
extension that corresponds to a similarly large
diversity of the textual genres they represent. Besides</p>
      <p>Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
literary ones, a huge amount of Latin texts of
several different genres can be found spread all over
Europe and beyond. An important textual genre
is represented by scientific treaties, which in many
cases are interesting not only for their contents, but
also because of the technical terminology they
feature.</p>
      <p>
        This is precisely the case for the Liber Abbaci
‘the book of the abacus’ by Leonardo of Pisa (also
known as Fibonacci). Written in the very first
years of the 1200s, it is a book on arithmetic
promoting a style of calculation based on Arabic
numerals without aid of an abacus. Fibonacci
12022021 is a project financed by the Tuscany
Region and involving the University of Pisa and the
Galilei Museum in Florence, following the
publication of a critical edition of the Liber Abbaci
by Enrico Giusti
        <xref ref-type="bibr" rid="ref12">(Fibonacci, 2020)</xref>
        . The goal of
the project is to produce an enhanced digital
edition of this work by leveraging advanced
publishing tools and investigating the use of
computational linguistics techniques in order to uncover
the wealth of linguistic, scientific and historical
information contained in the book.
      </p>
      <p>Besides its scientific interest, the Liber Abbaci
features a very peculiar lexicon, not often
represented in the currently available (linguistically
annotated) corpora for Latin. In order to fill this gap,
in the context of the project Fibonacci 1202-2021
we have started performing the linguistic
annotation of the Liber Abbaci, beginning from
part-ofspeech (PoS) tagging and lemmatization of a
specific chapter of the book, chosen for its linguistic
and historical interest. The dataset is freely
available online1.</p>
      <p>This paper describes the process of annotation
of the Liber Abbaci and two applications of its
1http://dialogo.di.unipi.it/
LiberAbbaci
results, namely (a) the evaluation of a number
of trained models for PoS tagging and
lemmatization for Latin in out-of-domain fashion and
(b) the interlinking of the annotated chapter with
other linguistic resources for Latin through the
Lila Knowledge Base (KB)2.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The research area dealing with the creation of
linguistic resources and Natural Language
Processing (NLP) tools for ancient languages has seen a
remarkable growth during the last decade
        <xref ref-type="bibr" rid="ref21 ref22 ref25 ref5 ref6">(Sprugnoli and Passarotti, 2020)</xref>
        . This has primarily
concerned Latin and Ancient Greek as essential media
to access and understand the so-called Classical
heritage. In particular, several annotated corpora
of Latin texts are currently available in digital
format: they follow different guidelines and tagsets
and feature different layers of linguistic
annotation. This section wants to provide a (far from
exhaustive) overview of such resources to show how
the dataset presented in this paper stands with
respect to the state of the art.
      </p>
      <p>
        The LASLA corpus contains 2,500,000
semimanually annotated tokens. It covers a large
portion of the extant Classical Latin literature. It was
started in 1961 by the LASLA research center at the
Université de Liège3 and is still being expanded4.
The corpus is considered to be a gold standard,
since the annotation of every token has been
manually verified by a philologist. The linguistic
information consists of lemmatization,
morphological tagging, and an additional syntactic layer for
verbs
        <xref ref-type="bibr" rid="ref25">(Verkerk et al., 2020)</xref>
        . Texts cover various
literary genres (theater, poetry, prose) and have a
chronological extension ranging from the
comedies of Plautus to the texts of Suetonius and Pliny
the Younger. Recent additions reach later stages
of Latin literature 5, but include neither Medieval
nor Neo-Latin works. Natural sciences and
technical works are weakly represented in the
corpus, the treatise De Agri Cultura ‘on agriculture’
by Cato and the recently added Naturales
Quaestiones ‘investigations about nature’ by Seneca
being the only examples.
      </p>
      <p>2https://lila-erc.eu
3http://web.philo.ulg.ac.be/lasla/
presentation-du-laboratoire/</p>
      <p>4See http://web.philo.ulg.ac.be/lasla/
textes-latins-traites/.</p>
      <p>5Of which some are already available: see
http://web.philo.ulg.ac.be/lasla/
textes-latins-en-cours-de-traitement/.</p>
      <p>
        The corpus of Latin Lemmatized Texts released
by
        <xref ref-type="bibr" rid="ref7 ref8">Thibault Clérice (Clérice, 2021</xref>
        a) is formed by
21,222,911 tokens (17,804,769 without
punctuation marks) and includes a large set of
Classical and Late Latin texts available in a a number
of open access corpora6. Clérice’s corpus
covers a very ample chronological span (up until the
9th century) as well as different genres: from
Classical literature (Horace, Ovid, etc.), to Christian
religious texts and legal texts. The linguistic
annotation consists of lemmatization and full
morphological description of the tokens , produced
automatically by applying the Pie Latin LASLA+
model 0.0.6
        <xref ref-type="bibr" rid="ref18">(Manjavacas et al., 2019)</xref>
        , fine-tuned
on ca. 1,500,000 tokens taken from the LASLA
corpus (Clérice, 2021b), with very good results
concerning lemmatization and PoS tagging7.
However, results appear to be less good on unknown
tokens8. This difference underlines the difficulty
of using automatic annotation tools on texts with
a very specialized language, surely not found in
LASLA, as is the case for Fibonacci’s Liber
Abbaci.
      </p>
      <p>
        As for syntactically annotated corpora, vfie
treebanks are currently available for Latin. They
are the Index Thomisticus Treebank (IT-TB)
        <xref ref-type="bibr" rid="ref19">(Passarotti, 2019)</xref>
        , the PROIEL treebank
        <xref ref-type="bibr" rid="ref1 ref11 ref15">(Haug and
Jøhndal, 2008; Eckhoff et al., 2018)</xref>
        , the Latin
Dependency Treebank by the Perseus Digital
Library (part of the Ancient Greek and Latin
Treebank)
        <xref ref-type="bibr" rid="ref3">(Bamman and Crane, 2007)</xref>
        , the Late Latin
Charter Treebank (LLCT)
        <xref ref-type="bibr" rid="ref22 ref5 ref6">(Cecchini et al., 2020b)</xref>
        and the UDante treebank
        <xref ref-type="bibr" rid="ref22 ref5 ref6">(Cecchini et al., 2020a)</xref>
        .
The treebanks include texts of different genres
(literary, historical, philosophical and documentary)
and periods (from Classical to Medieval), but
technical works are not represented.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Dataset Creation and Analysis</title>
      <p>The Liber Abbaci is made up of more than 270,000
tokens and is divided into 15 chapters of varying
length. The choice of starting our manual
annotation from chapter VIII de reperiendis pretiis
mercium per maiorem guisam ‘on finding out the price
of goods through the “greater means”’ is due to the
6For the full list, see https://github.com/
lascivaroma/latin-lemmatized-texts/tree/
0.1.2.</p>
      <p>7For lemmatization, accuracy: 0.9734 . For PoS tagging,
accuracy: 0.9651 .</p>
      <p>8For lemmatization, accuracy: 0.8716 . For PoS tagging,
accuracy: 0.9232 .
peculiarity of its content. Here, Fibonacci treats
many simple business negotiations using
proportions and referring to many examples taken from
the entire Mediterranean world. The examples
concern weight and monetary systems as well as
the main products bought and sold in the 13th
century. This means that the text is rich of
terminology specific of the mathematical domain but also
of trade and commerce. Chapter VIII is made up of
29,858 tokens (including punctuation marks), thus
covering about 10% of the total length of the Liber
Abbaci.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Annotation</title>
        <p>
          The manual annotation of chapter VIII is carried
out by a master’s degree student in Classical
languages, with excellent knowledge of Latin but
without any previous expertise in either
linguistic annotation or computational linguistics. The
overall effort of the work amounts to a total of
227 hours, including: training sessions, study
of the guidelines and of terminology related to
measures, coins and trade in the Middle Ages
(Marcinkowski, 2003; Martinori, 1915), the actual
annotation, the reconciliation after evaluation of
inter-annotator agreement (IAA, see Section 3.2),
periodic checks with supervisors, the linking of
the annotated text to the LiLa KB (see Section
5). We make use of a large number of
dictionaries as references: the Oxford Latin Dictionary
(OLD)
          <xref ref-type="bibr" rid="ref20">(Souter, 1968)</xref>
          , the Lexicon Totius
Latinitatis
          <xref ref-type="bibr" rid="ref13">(Forcellini, 1965)</xref>
          , the Dictionnaire illustré
latin-français (herafter: Gaffiot)
          <xref ref-type="bibr" rid="ref14">(Gaffiot, 2016)</xref>
          and the Thesaurus Linguae Latinae9 for Classical
Latin, but also the Dictionary of Medieval Latin
from British Sources
          <xref ref-type="bibr" rid="ref16">(Latham and Howlett, 1975)</xref>
          and the Glossarium mediae et infimae latinitatis
          <xref ref-type="bibr" rid="ref10">(du Cange et al., from 1883 to 1887)</xref>
          for Medieval
Latin. Tokenization and sentence splitting are
performed manually on a text editor, then
lemmatization and PoS tagging are carried out on a shared
spreadsheet following the Universal Dependencies
(UD) formalism
          <xref ref-type="bibr" rid="ref9">(de Marneffe et al., 2021)</xref>
          , in
particular both the universal and the language-specific
guidelines relative to the latest release of the UD
treebanks (v 2.9)10.
        </p>
        <p>The implementation of the UD guidelines to
the linguistic peculiarities of the text does not
9https://thesaurus.badw.de/
das-projekt.html</p>
        <p>10https://universaldependencies.org/
guidelines.html
always happen straightforwardly. Chapter VIII
of the Liber Abbaci, as well as the work in its
entirety, presents several typical features of
Medieval Latin, both graphically (e. g. the
monophthongization ae → e and the spelling nichil instead
of the Classical nihil ‘nothing’), morphologically
(e. g the presence of analytical verb forms such as
the “perfect”, i. e. present perfective, subjunctive
habeat . . . honeratum, instead of the Classical
onerauerit, from onero ‘to load’) and syntactically
(e. g. the nearly exclusive use of quod ‘that’ to
introduce declarative clauses, instead of accusative
and infinitive 11). It is also worth noting the very
limited use of enclitic particles (in the whole
chapter VIII, Fibonacci uses the enclitic conjunction
que ‘and’ only 3 times, appended to the auxiliary
verb form erunt ‘they will be’) and the presence
of syntactic calques of vernacular constructions
(e. g. secundum quod uadis multiplicando
‘according to what you are multiplying’, where uado is
preferred to the more Classical eo ‘to go’ and
further assumes an auxiliary function, and the use
of the gerundive form multiplicando is an
innovation).</p>
        <p>But the main peculiarities of the text concern the
lexicon. Chapter VIII presents indeed a rich set of
toponyms, units of measurement, names of coins
and Arabisms often not even reported by Medieval
Latin dictionaries. This is the case, for example,
of some names of places, such as Bugea, today’s
Big˘a¯ya/Bgayet in Algeria (a city where Fibonacci
spent a period of his childhood, learning the art
of calculation), and Septis, today’s Ceuta/Sabta on
the Strait of Gibraltar; or, among the numismatic
terms, of bolsonalia, a word designating a certain
amount of broken silver or mixture coins which
were sold to goldsmiths because they were
adulterated or out of date.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Inter-Annotator Agreement</title>
        <p>
          The IAA is calculated on 30 sentences (1,010
tokens), with the participation of a second scholar
with a background in Classical languages. We
register an almost perfect agreement with a Cohen’s
κ
          <xref ref-type="bibr" rid="ref1 ref15">(Artstein and Poesio, 2008)</xref>
          of 0.97 for
lemmatization and 0.94 for PoS tagging.
        </p>
        <p>
          The comparison between the two annotations
highlights two main issues. The first concerns the
choice of the UPOS (Universal Part Of Speech) tag
          <xref ref-type="bibr" rid="ref9">(de Marneffe et al., 2021, §2.2.2)</xref>
          for terms such as
11See for example
          <xref ref-type="bibr" rid="ref24">(Traina and Bertotti, 2015, C. XVI)</xref>
          .
nam ‘certainly’ and enim ‘namely’, because
different corpora and dictionaries adopt different
conventions: e. g. nam is labeled as adverb in the
Lila KB and Df in the Latin PROIEL treebank,
both possibly equivalent to UPOS ADV12; as S13,
standing for conjonction de coordination (UPOS:
CCONJ) in the LASLA corpus, and more
generically conjonction (servant à confirmer/causale)
(UPOS either CCONJ or SCONJ) in the Gaffiot;
ifnally particle (not necessarily corresponding to
UPOS PART) in the OLD, and similarly
particule in one sense in the Gaffiot. The treatment of
the etymologically related and functionally
similar enim is mostly identical for all sources, only
with the Gaffiot reporting a sense as adverbe
instead of particule, followed by the LASLA
corpus in using both labels S and M (generic for
adverbe), the latter though very marginally. These
terms have been discussed and finally assigned the
UPOS PART, used in the latest Latin UD treebanks
to label discoursive particles like these. Such
difficulties derive on one hand from the “volatile” and
diachronically variable nature of similar elements,
but on the other hand, and relatedly, to traditional
grammars overlooking them and more generally
skipping over pragmatic phenomena, in favour of
“more Classical” parts of speech (hence the
frequent inclusion of nam, enim, etc. in the catchall
category of “adverbs”).
        </p>
        <p>
          The second issue is the UPOS to be used for
unus ‘one’. Fibonacci often uses unus to indicate a
generic entity, as is clearly visible when paralleled
by alter ‘other’. In this case, unus is tagged as DET
(determiner), like alter14. In a number of other
contexts, however, unus specifies the quantity of
a certain object. In such cases it is considered a
NUM (numeral)15. The difficulty here originates
from a well known and complex linguistic change
that will eventually produce a clear indefinite
article from the numeral in Romance languages, but
for which, being so gradual, we cannot pinpoint
12Cf.
          <xref ref-type="bibr" rid="ref11">(Eckhoff et al., 2018, §5)</xref>
          13With only very few exceptions when it is seen as part of
a compound expression with tmesis, thus not receiving an
autonomous PoS; cf. Pl. Am. 2.1, 49-50: Quo id, malum, pacto
potest nam (mecum argumentis puta) fieri, nunc uti tu et hic
sis et domi?, interpreted as an instance of quonam ‘whither
pray?’, itself receiving K meaning pronom interrogatif.
        </p>
        <p>14For instance, in the clause ita est pretium unius ad
pretium alterius (VIII, 8) ‘so the price of the one
[merchandise] is to the price of the other’.</p>
        <p>
          15For instance, in the clause . . . que multiplica per summam
denariorum unius libre (VIII, 20) ‘which you have to
multiply by the amount of denarii of which one pound consists’.
an exact historical moment; cf.
          <xref ref-type="bibr" rid="ref17">(Ledgeway, 2012,
§4.2.1)</xref>
          .
Table 1 reports accuracy scores computed on our
gold standard processed with UDPipe using the
UD v2.6 models for Latin
          <xref ref-type="bibr" rid="ref23">(Straka and Straková,
2017)</xref>
          . The scores clearly show that current
models are not good enough to process the Latin of
Fibonacci. The best accuracy for lemmatization
is achieved by the model trained on the LLCT
treebank, which contains a set of Early Medieval
charters written in Tuscany. However, this scores are
lower than state-of-the-art ones: the best
participating system at the EvaLatin 2020 evaluation
campaign achieves an accuracy of 96, 19% for
lemmatization and 96, 74% for PoS tagging on
the corresponding test set
          <xref ref-type="bibr" rid="ref21 ref22 ref5">(Sprugnoli et al., 2020)</xref>
          ,
i. e. about 33 and 15 points more than the results
obtained on Fibonacci.
        </p>
        <sec id="sec-3-2-1">
          <title>EvaLatin2020 IT-TB</title>
          <p>LLCT
Perseus
PROIEL</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Lemma</title>
          <p>63.60
65.58
68.81
67.54
60.25</p>
          <p>UPOS
81.90
77.14
82.79
78.37
51.64</p>
          <p>Taking into consideration lemmatization, the
percentage of out-of-vocabulary lemmas, that is,
lemmas present in the text by Fibonacci but not
in the training texts of the models, is very high
(&gt; 50% of lemma types). The majority of errors
are registered for numbers and common nouns.
The first problem is due to the fact that some
models do not recognize Arabic numbers, because
they have not seen them in their training data,
while others lemmatize them with a special
“metalemma” of the kind of num. arab., eschewing
lexical forms. As for common nouns, most errors
related to lemmatization concern the lexical classes
discussed in Section 3. For example, the tokens
libris and libre are often lemmatized as liber ‘free’
(ADJ) instead of libra ‘pound’ (NOUN).</p>
          <p>Table 2 shows the F1 score per UPOS tag. We
observe that an F1 above 70% is achieved by any
model only on 5 tags: ADP, NOUN, NUM, SCONJ
and VERB. No model recognizes the SYM tag
(used for mathematical operators such as
parenSYM
AUX
ADJ
PRON
PART
ADV
CCONJ
SCONJ
VERB
NOUN
DET
PROPN
NUM
ADP</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Global</title>
        <p>theses), because it is not present in their
respective training data. The same is true for the tag
PART in IT-TB (up until UD v2.8)16, Perseus and
PROIEL, and for the tag DET in Perseus. In old
versions of the IT-TB, DET is limited to the
protoarticle ly (8 occurrences), while in Perseus the
tag PROPN appears only for the lemma Aefulanus
(1 occurrence). The IT-TB-based model, too,
registers a near-zero F1 score for PROPN: in the
corresponding training data, this tag is used for a
restricted (116 types of lemmas) set of terms mostly
specific to the domains of philosophy and
religion (e. g. Aristoteles, Maria), not present in our
dataset. Low performances are registered also for
the AUX tag, the annotation of which is not
consistent in training data: in Perseus, this tag is not used
at all, while in EvaLatin 2020 it marks only the
auxiliaries in periphrastic passive (including
deponent) constructions, while in the other treebanks it
is applied also to verbal copulas, as per UD
guidelines. Further, the Liber Abbaci sees the rise (1
occurrence) of habeo ‘to have’ as a possible
auxiliary (cf. Section 3.1), unheard of in Classical Latin
and only attested (albeit marginally) in LLCT.</p>
        <p>16Annotation discrepancies with respect to other Latin UD
treebanks for INTJ, NUM, PART, PRON and DET have been
resolved in IT-TB in its last version (2.9), released in
November 2021; however, the model adopted in this paper and
currently available in UDPipe is based on an older version of the
data.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Linking and Querying in LiLa</title>
      <p>
        The LiLa KB makes linguistic resources for Latin
interoperable by linking tokens in corpora and
entries in dictionaries/lexica to a collection of
canonical forms for Latin called Lemma Bank
        <xref ref-type="bibr" rid="ref21 ref22 ref6">(Passarotti et al., 2020)</xref>
        . In order to connect the
lemmas of chapter VIII to LiLa’s KB, a string
match is first performed between the lemmas in
the texts and those in the KB, also taking into
account their parts of speech. Using this strategy,
88.8% of the lemmas are directly connected to
a single entry in the KB. The remaining
unconnected lemmas fall into two possible categories:
ambiguous lemmas, that is, with possible
connection to more than one entry in the KB; and
lemmas absent from the KB. More specifically,
we find 44 ambiguous lemmas (corresponding to
631 tokens): for example, colligo can be
connected to two entries: either a first-conjugation
verb colligare17 ‘to bind’, or a third-conjugation
verb collige˘re18 ‘to gather’. These cases are
manually disambiguated, checking each context of use.
The remaining, not directly connected lemmas are
not present in the KB and need to be manually
added: these are mainly words denoting weight
and monetary units (e. g. karatus ‘carat’), or
different written representations of lemmas already
in LiLa (e. g. torscellus is a graphic variant of
tor17https://lila-erc.eu/data/id/lemma/94854
18https://lila-erc.eu/data/id/lemma/94855
cellus19, a unit of length). Thanks to the linking,
each lemma of our dataset becomes part of an
interoperable ecosystem made of resources of
different kinds. We can thus query different interlinked
resources using SPARQL and the LiLa endpoints20.
For example, we can find the lemmas appearing
only in chapter VIII21 and not in the other texts that
are currently linked to the KB: the Summa
Contra Gentiles by Thomas Aquinas (from the Index
Thomisticus), those found in UDante (a corpus of
5 works mostly by Dante Alighieri, or attributed to
him, manually annotated following the UD
formalism), and the Querolus siue Aulularia (an
anonymous comedy dating back to the 5th c. AD).
      </p>
      <sec id="sec-4-1">
        <title>Lemma</title>
        <p>rotulus
soldus
virgula
byzantius
cantare</p>
      </sec>
      <sec id="sec-4-2">
        <title>Gloss</title>
        <p>unit of weight
monetary unit
bar of a fraction
monetary unit
unit of weight</p>
        <p>Table 3 shows the 5 most frequent distinctive,
i. e. exclusively found in the Liber Abbaci,
lemmas retrieved using a SPARQL query22. They are
all related to mathematics, coins and units of
measurement, confirming the specificity of the domain
of our dataset. In particular, rotulus and canta¯re
are two units of weight, both deriving from
Arabic, respectively from ra.tl (in turn, a
metathetical adaptation of Greek λίτρα litra ‘pound’) and
qin.ta¯r, which designates a weight of 100
rotuli23. The term soldus, instead, indicates a
unit of measurement used for monetary quantities.
Among the many currencies mentioned in chapter
VIII, Fibonacci often cites the byzantius, a golden
19https://lila-erc.eu/data/id/lemma/133810
20https://lila-erc.eu/sparql/
21https://lila-erc.eu/data/corpora/
CorpusFibonacci/id/corpus/Liber Abbaci
22https://github.com/CIRCSE/
SPARQL-queries/blob/main/
distinctivelemmas-Fibonacci.rq</p>
        <p>
          23It should be noted that Fibonacci alternates a
thirddeclension canta¯re (gen. sing. canta¯ris) with a
seconddeclension cantarium (gen. sing. cantarii). During
lemmatization of the text, the various attested singular
forms have been linked to their respective lemmas; the
nom./acc. plur. cantaria, which theoretically could derive
both from canta¯re and cantarium, has been linked to the
lemma canta¯re for simple reasons of probability, as it is the
most frequently used by Fibonacci among these two forms.
coin minted in Constantinople24. Finally, virgula
(diminutive of virga, properly a ‘rod’, used by
Fibonacci in the same sense of virgula) primarily
denotes the bar between the numerator and
denominator of a fraction, but it can also designate the
fraction itself
          <xref ref-type="bibr" rid="ref4">(Bocchi, 2004)</xref>
          .
6
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>This paper describes the annotation of one
chapter of the Liber Abbaci by Fibonacci, and reports
on the linguistic peculiarities of this text and the
ensuing challenges.</p>
      <p>The results of existing UDPipe models in
lemmatization and tagging show low accuracy and
F1 scores when compared to the state of the art
for these tasks in the recent EvaLatin 2020
evaluation campaign. This, on the one hand, can be
attributed to the characteristics of the genre of
Fibonacci’s texts, which are representative of
scientific Medieval Latin texts, and on the other hand
can be explained with the different choices in
annotation style of Latin treebanks released under
the UD project. Substantial improvements can be
expected with models trained on new releases of
Latin treebanks which have already undertaken the
effort of resolving annotation discrepancies and of
making the annotation style across treebanks more
homogeneous. Further improvements will
however require new annotated chapters and
experiments in domain adaptation, which are scheduled
as future work.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work is a contribution to the Fibonacci
12022021 project, financed by the Tuscany Region.
Part of the work has been funded by the
European Research Council (ERC) under the
European Union’s Horizon 2020 research and
innovation programme – Grant Agreement No. 769994.
The authors want to thank: prof. Andrea
Bocchi, dott. Alessandro Gelsumini, prof. Pier Daniele
Napolitani and prof. Enrica Salvatori for their
linguistic and historical advice.
Edoardo Martinori. 1915. La Moneta: vocabolario
generale. Instituto italiano di numismatica.</p>
      <p>Marco Passarotti, Francesco Mambrini, Greta Franzini,
Flavio Massimiliano Cecchini, Eleonora Litta,
Giovanni Moretti, Paolo Ruffolo, and Rachele
Sprugnoli. 2020. Interlinking through lemmas. the
lexical collection of the lila knowledge base of
linguistic resources for latin. Studi e Saggi Linguistici,
58(1):177–212.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Ron</given-names>
            <surname>Artstein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Poesio</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Inter-coder Agreement for Computational Linguistics</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>34</volume>
          (
          <issue>4</issue>
          ):
          <fpage>555</fpage>
          -
          <lpage>596</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>24Also mentioned is the byzantius saracenatus, equivalent to the hyperperus, that is, a byzantius with inscriptions in Kufic characters</article-title>
          (Martinori,
          <year>1915</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>David</given-names>
            <surname>Bamman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Gregory</given-names>
            <surname>Crane</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>The Latin Dependency Treebank in a Cultural Heritage Digital Library</article-title>
          .
          <source>In Proceedings of the Workshop on Language Technology for Cultural Heritage Data (LaTeCH</source>
          <year>2007</year>
          ), pages
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          , Prague, Czech Republic, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Bocchi</surname>
          </string-name>
          .
          <year>2004</year>
          . In Michelangelo Zaccarello and Lorenzo Tomasin, editors,
          <article-title>Storia della lingua e iflologia. Per Alfredo Stussi nel suo sessantacinquesimo compleanno</article-title>
          , chapter Sì nel Livero de l'abbecho, pages
          <fpage>121</fpage>
          -
          <lpage>158</lpage>
          . SISMEL - Edizioni
          <string-name>
            <surname>del Galluzzo</surname>
          </string-name>
          , Florence, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Flavio M. Cecchini</surname>
            , Rachele Sprugnoli, Giovanni Moretti, and
            <given-names>Marco</given-names>
          </string-name>
          <string-name>
            <surname>Passarotti</surname>
          </string-name>
          .
          <year>2020a</year>
          .
          <article-title>UDante: First Steps Towards the Universal Dependencies Treebank of Dante's Latin Works</article-title>
          .
          <source>In Seventh Italian Conference on Computational Linguistics</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          , Bologna. CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Flavio</given-names>
            <surname>Massimiliano</surname>
          </string-name>
          <string-name>
            <surname>Cecchini</surname>
          </string-name>
          , Timo Korkiakangas, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Passarotti</surname>
          </string-name>
          .
          <year>2020b</year>
          .
          <article-title>A New Latin Treebank for Universal Dependencies: Charters between Ancient Latin and Romance Languages</article-title>
          .
          <source>In Proceedings of the 12th Language Resources and Evaluation Conference</source>
          , pages
          <fpage>933</fpage>
          -
          <lpage>942</lpage>
          , Marseille, France, May. European Language Resources Association.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Thibault</given-names>
            <surname>Clérice</surname>
          </string-name>
          .
          <source>2021a. lascivaroma/latinlemmatized-texts: 0.1</source>
          .
          <fpage>2</fpage>
          <string-name>
            <surname>- HN</surname>
            <given-names>PSL</given-names>
          </string-name>
          , May. DOI:
          <volume>10</volume>
          .5281/zenodo.4661034; project online at https://github.com/ lascivaroma/latin-lemmatized-texts.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Thibault</given-names>
            <surname>Clérice</surname>
          </string-name>
          . 2021b.
          <article-title>Latin Lasla Model, Apr</article-title>
          . DOI:
          <volume>10</volume>
          .5281/zenodo.4661034.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
            , Christopher D. Manning, Joakim Nivre, and
            <given-names>Daniel</given-names>
          </string-name>
          <string-name>
            <surname>Zeman</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <string-name>
            <given-names>Universal</given-names>
            <surname>Dependencies</surname>
          </string-name>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>47</volume>
          (
          <issue>2</issue>
          ):
          <fpage>255</fpage>
          -
          <lpage>308</lpage>
          ,
          <fpage>07</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <article-title>Charles du Fresne sieur du Cange</article-title>
          , bénédictins de la congrégation de Saint-Maur, d. Pierre Carpentier, Johann Christoph Adelung,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Louis</surname>
          </string-name>
          <string-name>
            <surname>Henschel</surname>
          </string-name>
          , Lorenz Diefenbach, and Léopold Favre.
          <article-title>from 1883 to 1887</article-title>
          .
          <article-title>Glossarium mediae et infimae latinitatis</article-title>
          . Favre, Niort, France.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Hanne</given-names>
            <surname>Martine</surname>
          </string-name>
          <string-name>
            <surname>Eckhoff</surname>
          </string-name>
          , Kristin Bech, Gerlof Bouma, Kristine Eide, Dag Haug, Odd Einar Haugen, and
          <string-name>
            <given-names>Marius</given-names>
            <surname>Jøhndal</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The PROIEL treebank family: a standard for early attestations of IndoEuropean languages</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>52</volume>
          (
          <issue>1</issue>
          ):
          <fpage>29</fpage>
          -
          <lpage>65</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>Leonardus Bigollus Pisanus vulgo Fibonacci</source>
          .
          <year>2020</year>
          . Liber Abbaci, volume
          <volume>79</volume>
          of Biblioteca di «Nuncius». Leo S. Olschki, Florence, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Egidio</given-names>
            <surname>Forcellini</surname>
          </string-name>
          .
          <year>1965</year>
          .
          <article-title>Lexicon totius latinitatis</article-title>
          .
          <source>Arnaldo Forni</source>
          , Bologna, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Félix</given-names>
            <surname>Gaffiot</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Dictionnaire Latin-Français. Accessible at gaffiot</article-title>
          .fr.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Dag</given-names>
            <surname>Trygve Truslew Haug</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marius</given-names>
            <surname>Jøhndal</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Creating a Parallel Treebank of the Old IndoEuropean Bible Translations</article-title>
          .
          <source>In Proceedings of the Second Workshop on Language Technology for Cultural Heritage Data (LaTeCH</source>
          <year>2008</year>
          ), pages
          <fpage>27</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Ronald</given-names>
            <surname>Edward Latham and David R Howlett</surname>
          </string-name>
          .
          <year>1975</year>
          .
          <article-title>Dictionary of Medieval Latin from British Sources: Fascicule V: IJKL</article-title>
          . OUP Oxford.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Adam</given-names>
            <surname>Ledgeway</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>From Latin to Romance, volume 1 of Oxford studies in historical and diachronic linguistics</article-title>
          . Oxford University Press, Oxford, UK.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Enrique</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          , Ákos Kádár, and
          <string-name>
            <given-names>Mike</given-names>
            <surname>Kestemont</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Improving lemmatization of nonstandard languages with joint learning</article-title>
          .
          <source>In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pages
          <fpage>1493</fpage>
          -
          <lpage>1503</lpage>
          , Minneapolis, Minnesota, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Passarotti</surname>
          </string-name>
          ,
          <year>2019</year>
          . volume
          <volume>10</volume>
          of Age of Access?
          <article-title>Grundfragen der Informationsgesellschaft, chapter The Project of the Index Thomisticus Treebank</article-title>
          , pages
          <fpage>299</fpage>
          -
          <lpage>320</lpage>
          . De Gruyter Saur, Berlin, Germany; Boston, MA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Souter</surname>
          </string-name>
          .
          <year>1968</year>
          . Oxford Latin dictionary: OLD. Clarendon Press.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Rachele</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Passarotti</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <source>Proceedings of LT4HALA 2020-1st Workshop on Language Technologies for Historical and Ancient Languages. In Proceedings of LT4HALA 2020-1st Workshop on Language Technologies for Historical and Ancient Languages.</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Rachele</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          , Marco Passarotti, Flavio Massimiliano Cecchini, and
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Pellegrini</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of the EvaLatin 2020 evaluation campaign</article-title>
          .
          <source>In Proceedings of LT4HALA 2020 - 1st Workshop on Language Technologies for Historical and Ancient Languages</source>
          , pages
          <fpage>105</fpage>
          -
          <lpage>110</lpage>
          , Marseille, France, May.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Milan</given-names>
            <surname>Straka</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jana</given-names>
            <surname>Straková</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Tokenizing, POS tagging, lemmatizing and parsing UD 2.0 with UDPipe</article-title>
          .
          <source>In Proceedings of the CoNLL</source>
          <year>2017</year>
          <article-title>Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies</article-title>
          , pages
          <fpage>88</fpage>
          -
          <lpage>99</lpage>
          , Vancouver, Canada, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Alfonso</given-names>
            <surname>Traina</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tullio</given-names>
            <surname>Bertotti</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Sintassi normativa della lingua latina</article-title>
          .
          <source>Pàtron</source>
          , Bologna, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Philippe</given-names>
            <surname>Verkerk</surname>
          </string-name>
          , Yves Ouvrard, Margherita Fantoli, and
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Longrée</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <string-name>
            <surname>L.A.S.L.A.</surname>
          </string-name>
          and
          <article-title>Collatinus: a convergence in lexica</article-title>
          .
          <source>SSL</source>
          ,
          <volume>1</volume>
          (LVIII):
          <fpage>95</fpage>
          -
          <lpage>120</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>