<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IRCDL</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Automatic Annotation of Legal References (Allegationes) in the Liber Extra's Ordinary Gloss</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Esuli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincenzo Roberto Imperia</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Puccetti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Istituto di Scienza e Tecnologie dell'Informazione “A. Faedo”</institution>
          ,
          <addr-line>via G. Moruzzi 1 - 56124 Pisa PI</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Università degli Studi di Palermo - Dipartimento di Giurisprudenza</institution>
          ,
          <addr-line>via Maqueda, 172 - 90134, Palermo PA</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>21</volume>
      <fpage>20</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>The study of normative corpora of the past is a key activity in the fields of Religious Studies and Legal History. The development of intelligent software tools that support this activity is of paramount importance to support the digital transformation of the community. We present an interdisciplinary activity that lead to an accurate automatic annotation of legal references in the Liber Extra's Ordinary Gloss. An index of legal references as been derived from the annotations enabling the creation of novel navigation and data analysis tools. The contribution of this work is twofold: the actual index is already by itself valuable resource for the discipline, and we detail the process that lead to its production, showing that an efective result can be delivered by a small team with limited resources. Both the index and the code are made publicly available.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Legal references</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>Conditional Random Fields</kwd>
        <kwd>Dataset</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>1.1. Glossa and Allegationes, from Medieval Books to Computer Science</title>
        <p>Before examining the specific case study presented here and evaluating the concrete applications in
the field of legal history resulting from the development of automatic annotation techniques for legal
allegations, it is necessary to clarify the meaning of terms such as gloss, ordinary gloss and allegations,
and in particular to consider the nature of allegations in the intellectual context of medieval jurists.</p>
        <p>
          The term glossa (gloss) refers to “a brief annotation composed and written to explain a text and
addressing either its terminology and its exterior trappings or its animating spirit and its underlying
principles” [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. From the late 11th century and for several centuries thereafter, the gloss served as the
principal paratextual tool through which law masters in law schools, later Universities, explained and
commented on Roman and Canon law compilations, enabling students and practitioners to engage
with technically complex texts. The term glossa ordinaria (ordinary gloss) refers to apparatus of glosses
compiled by some eminent jurists, renowned for their exceptional quality and thoroughness. Their
widespread acceptance ensured their consistent reproduction alongside the normative reference text,
establishing them as the ordinary apparatus [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The phenomenon of normative texts accompanied by
ordinary gloss is evident both in the manuscript production of the late Middle Ages, the earliest printed
works of the late 15th century, continuing into the significant editorial initiatives of the 16 th century.
        </p>
        <p>
          Manuscripts and printed texts retained the
standard page layout essentially unchanged, with the
normative text and the ordinary gloss arranged
according to a model that has been defined, in its
essential features, as the “agora template” [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], as
shown in Figure 1. This spatial configuration, with
the normative text at the center framed by the
glosses, mirrored the dialogical nature of the
legal interpretation method: the authoritative text
was accompanied by commentary that explained,
questioned or contradicted it. It can be afirmed
that it is the very content of the apparatus that
qualifies them as "hypertexts" [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], since they demand
the active participation of the reader. The margins
of the normative text contain various components,
each serving a specific function, but collectively
acting as tools to help the user understand the main
text, making it more accessible and consultable.
Particular attention will be paid here to one of these
components, the allegationes, defined by Hermann
Kantorowicz [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] as references to auctoritates that
must justify every assertion or, at least, deal with
those that at first glance appear to contradict it.
        </p>
        <p>
          These references are, in the vast majority of cases,
citations of legal norms, presented in a highly abbre- Figure 1: A page from the Liber Extra’s Ordinary
viated form, using abbreviations and numbers that Gloss [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The box of text in the upper center of the
correspond to the conventional criteria for identi- page is the normative text from Liber Extra. The
fying the legal source containing the norm. These surrounding text are the glosses that also contain
references, initially sporadic in the earliest layers the legal references.
of glosses, gradually increased in number until they
became quantitatively predominant, alongside the refinement of interpretive techniques developed
by jurists. Many recent work [
          <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13">10, 11, 12, 13</xref>
          ] reafirmed and clarified that allegations are not merely
citations: the use of a reference constitutes an appeal to an authority that not only lends greater solidity
to the legal argument, but forms its very foundation. The use of legal references by medieval jurists
had thus a complex significance. It demonstrated that interpretative operations were carried out on
legitimate grounds, based, on the one hand, on applicable and valid legal norms and, on the other, on
the consistent use of legal categories shared within a common cultural and intellectual context [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>
          These conceptual premises are closely related and provide the backdrop for considerations concerning
more practical aspects. The first pertains to the citation style of allegationes, which presupposed a
normative text that was now stable, fixed, and unalterable. This allowed the reader to pinpoint the
referenced passage with certainty [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The second concerns the nature and purpose of these references.
The primary aim was to create a network in which the various sedes materiae were organised and linked,
thus enabling the reader-user to navigate vast normative compilations [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          Moreover, the use of legal references could take on an even more incisive argumentative function
if, in addition to a principle or rule derived from the normative text, allegations pro and contra were
included. This technique quickly evolved into a distinct literary genre, with many works constructed
according to these specific criteria [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          The formal and substantive nature of the allegationes in the works of the jurists underscores the need
to exploit the potentials of computer science, and specifically machine learning, to develop accurate
tools capable of preserving their peculiarities in future digital editions of medieval legal texts [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Annotation of legal references</title>
      <p>
        The complex and stratified genesis of the ordinary apparatus to normative compilations, closely linked
to the concept of the authoritative text itself in the Middle Ages, prevented the preparation of critical
editions of them in the modern sense. Consultation of them still requires recourse to manuscripts or
early modern printed editions [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This is also the case for the Corpus iuris canonici. For this reason,
the Glossa Ordinaria (Ordinary Gloss) to the Decretales Gregorii IX, best known as the Liber Extra, was
chosen for the design and development of an automatic annotation system for the legal allegations.
Promulgated by the Pope with the bull Rex Pacificus , the legal collection is subdivided into 5 books, 185
titles, 1971 chapters, with a total of 9872 lemmas. The gloss of each lemma contains a variable number
of legal references. From the text of the 1582 Editio Romana [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which is the common reference edition,
a complete digital transposition was made by Edward A. Reno III, as part of the project “The Digital
Decretals” [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The files containing the digital text can be found on the project’s website, together with
specific references to the sections of the apparatus not included in the transcription work, as well as
the interventions made to standardise elements of spelling, abbreviation, punctuation and numbering
in the text.
      </p>
      <p>Figure 2 illustrates the process that lead to an accurate automatic annotation of the whole Liber
Extra’s Ordinary Gloss based on the annotation of a small subset by a human expert, followed by the
creation of an index of all the annotation in which every legal reference, pointing to a specific norm, is
linked to a lemma, chapter, and title, of the Liber Extra. The next Sections will detail the process.</p>
      <sec id="sec-2-1">
        <title>2.1. The annotation schema</title>
        <p>An annotation is a span of text marked in the text to identify some relevant information of a specific
type. The annotation schema we defined identifies the four types of entities:
• Annotation of type glossed lemma, “Lemma glossato”, indicates the specific terms that are glossed.</p>
        <p>Legal references are included in the gloss text.
• Annotations of type legal reference, “Allegazione normativa” annotate the references to legal
norms or regulations.
• The last two annotations types are title, “Titolo”, and chapter, “Capitolo”, which together with the
lemma precisely identify the position of the legal reference in the Liber Extra.</p>
        <p>The annotations of these entities thus enables building a link between the Liber Extra and the legal
norms and regulations that are crucial for the interpretative framework of the book. The annotation
data also include the exact character position of the beginning and the end of the annotated text in the
digital text.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Expert annotation</title>
        <p>
          An expert of the domain performed the annotation of the legal references on a subset of the book. The
expert annotated 12 titles out of 185, making a total of 4578 annotations. The titles have been randomly
sampled across the whole book to better cover the diferent content of each section of the book, i.e.,
following Reno’s notation [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]: 1.02, 1.11, 1.33, 2.02, 2.23, 3.02, 3.26, 4.17, 4.19, 5.01, 5.03, 5.23. The expert
focused only on annotating the legal references, as the annotations of the other three types have been
made successively in a completely automatic way, as described in Section 2.3.
        </p>
        <p>
          The tool used by the expert to perform the manual annotation is INCEpTION [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. INCEpTION is a
popular annotation platform designed for collaborative and eficient text annotation. It supports a wide
range of tasks, including named entity recognition, relation extraction, and classification. INCEpTION
helps the manual annotation with a dedicated GUI design, and also by providing automated suggestions
for spans of text that are evaluated as potential new annotations, which the expert can validate with a
single mouse click. The automated suggestions are based on the definition of a recommender, i.e., a
machine-learning algorithm that trains a model by continuously observing the annotations made by
the expert. The automatic annotation model trained by INCEpTION is a valuable help for the annotator,
yet it is not of suficient accuracy to be used to perform a complete automatic annotation of the rest of
the book. We thus used the 12 annotated titles as a training set for a proper batch training process of
more accurate models, as detailed in Section 2.3.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Training a model for the automatic annotation of legal references</title>
        <p>
          The code to replicate the automatic annotation is published with an open source license [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. We focus
here only on the legal references, as the annotation of the other types of entities has been solved using
diferent methods, as detailed in Section 2.3.1.
        </p>
        <p>
          The automatic annotation process is modeled as a word labeling problem. We adopted the BILOU
labeling schema [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] in which each word can be assigned to one of a set of labels, either if is part of a
legal reference or not (O label). For words that are part of a legal reference, diferent labels are used if
the word is at beginning of the annotation (B label), it is the last (L label), it inside the sequence of words
that form the legal reference (I label), or the legal reference is composed of a unique word (U label).
        </p>
        <p>
          We tested two approaches to train the automatic annotation model: using a “traditional” statistical
machine learning algorithm, i.e., Conditional Random Fields [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] (CRFs), and fine-tuning a
transformerbased Large Language Model (LLM) [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. A preliminary test of few-shot prompting [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] of a generative
LLM reported a very low accuracy and was discarded.
        </p>
        <p>CRFs are probabilistic graphical models that are used to model sequential or structured data by
defining conditional probabilities of a set of target variables given a set of observed variables, allowing
for the incorporation of context and interdependencies among the variables in the target space. In
natural language processing (NLP), CRFs are often used for tasks like named entity recognition or
partof-speech tagging, where the output labels (e.g., tags) are interdependent. We tested two configurations
for the CRFs graph: a simple configuration in which the observed variables are all the word bigrams in
a context window of three words before and after the one to be annotated, and a rich configuration that
considered bigrams and trigrams in a context windows of size seven.</p>
        <p>
          For the fine tuning of LLMs we tested two models: the original BERT [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], as a baseline, and Latin
BERT [23], a state-of-the-art model for Latin. The fine tuning of LLMs was made using the Transfomers
python package [24], training for 10 epochs, using a learning rate of 2 · 10− 5, a weight decay value of
0.01, and a batch size of 32. Sequences longer than 512 token were split in multiple sequences.
        </p>
        <p>The comparison of the methods is based on a 10-fold cross validation on the expert-annotated data.
The annotated data is split into ten parts. The splits are kept the same across all the tested methods.
For every fold one tenth of the annotated data is considered to be the test data and the remaining nine
tenth are the training data. The process repeats for all the folds, collecting for each tested method an
automatic annotation of all data. The accuracy of the automatic annotation is determined by comparing
it with the annotations from the expert, using a token-and-blank evaluation model [25].</p>
        <p>Results in Table 1 show that CRFs in the rich configuration performed with close to perfect accuracy.
The more complex graph of the rich configuration required three times the training time of the basic
configuration, yet the training time is still short and the improvement is worth the additional cost.</p>
        <p>The comparison between BERT and Latin BERT shows the impact of the main training language
to a task. The tokenizer of BERT obviously struggles with Latin words. For example the five-words
expression “Quae est radix omnium malorum" is tokenized by BERT in 19 tokens, whereas Latin BERT
produces exactly five tokens. This allowed Latin BERT to better identify the relevant elements of the
language that are related to the expression of legal references, whereas BERT struggled with very short
sub-word tokens, which are evidently less statistically related to the concept of legal reference.</p>
        <p>
          The comparison of CRFs with LLMs highlights that in this annotation task CRFs have many advantage
points. The obvious one is that CRFs get the best accuracy. Not less important is that CRFs require less
computational resources. The training time of CRFs is based on using a desktop with a single Intel i9
CPU, while the LMMs time are based on a dedicated server with 4 A-40 NVIDIA GPUs, costing roughly
ten times the desktop computer. Similarly the sizes of the final models shows a clear advantage for CRFs.
The diferences in training time are not very relevant, considering the diferences in hardware, and the
fact that even the longest training time is relatively quick. CRFs makes possible to train an annotation
model on personal hardware, enabling the adoption of machine learning to small research groups with
limited computational resources. More experiments on the fine tuning of LLMs may have reduced
the gap with CRFs, yet the high accuracy obtained by CRFs satisfied our goals, thus not justifying the
additional computational costs.
2.3.1. Automatic annotation of other entities
The annotation of chapters and titles has been made using a regular expression, exploiting the specific
format used in The Digital Decretals [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. For the annotation of the glossed lemmas, we had only a
ordered list of the glossed lemmas, leaving us to identify their position in the text. We solved this
with a search algorithm that considered the constraints imposed by the list. For example, the word
“Omnipotens” occurs 31 times in the whole text, but the only instance as a glossed lemma occurs after
the glossed lemma “Incomprehensibilis” and before “Inefabilis”. Solving all the constraints for all the
glossed lemmas allowed us to find the exact position for all of them.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Building the index of legal references</title>
        <p>The automatic annotation of the whole text identified 41784 legal references [ 26], each linked to a
glossed lemma, a chapters, a title, and a book part (among the five parts of the Ordinary Gloss relative
to the five books of the Liber Extra). The automatic annotation thus had a multiplier efect of 9˜.1 with
respect to the number of annotation from the expert.</p>
        <p>We split each legal references in two parts, the “title” or “section” that contains the referenced norm,
and the string that identifies the norm. For example, the legal reference “22. q. 4, incommutabilis” points
to Gratian’s Decretum, Causa 22, Quaestio 4, and its chapter (or canon) “incommutabilis” (according to
today’s citation standard: C.22 q.4 c.9). This allowed us to identify a total of 1795 referenced titles or
sections, from various collections, i.e., the Corpus Iuris Canonici and the Corpus Iuris Civilis.</p>
        <p>The index enables the researchers to perform the activities presented in Section 1. In addition to the
use as a browsing resource, the index itself can be the subject of analysis. For example, Table 2 lists the
ten titles or sections with most references in the Ordinary Gloss, considering the whole Liber Extra and
each of its books separately. This is a simple example of a data analysis enabled by the index.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Qualitative analysis</title>
      <p>A qualitative evaluation of the results of the automatic annotation process is particularly relevant when
considering the degree of correctness and accuracy in relation to the amount of data processed.</p>
      <p>Regarding the correctness of the annotation, the system’s ability to distinguish legal references from
the rest of the text, which is usually written in a discursive and argumentative style, according to the
explanatory function inherent in the gloss, is impressive. Concerning the accuracy of the annotation, the
system has demonstrated, in most of the cases, its ability to delimit the exact scope of legal references.
There are a few exceptions when the text of the allegation is unusually long or contains overly specific
references. There are no particular issues in identifying allegations that refer to the Liber Extra or to
the three partes of Gratian’s Decretum, as each is clearly defined in its citation form. The same holds
true for the allegations to the Code or the Institutes of the Corpus Iuris Civilis, which are usually clearly
preceded by their respective symbols (C. or Inst.).</p>
      <p>The automatic annotation has a recurring imprecision when identifying references to the Digest,
preceded by the abbreviation “f.”, which is erroneously omitted in many cases. In this case the automatic
annotation model shown a preference for shorter annotations, considering that in most cases both
versions of the annotations, with or without “f.”, are potentially correct yet referring to diferent sources.
The specific nature of this error made it easy to solve with a simple automatic post processing of the
annotations, checking for any eventual “f.” preceding them and adding it to the annotation, as we
found no cases in which the “f.” expression preceding an annotation was to be excluded. The published
version of the index [26] is thus not afected by this issue.</p>
      <p>Quantitatively insignificant are the cases in which the system detects text passages as legal allegations
when they are not.Particularly interesting are the cases where the system, due to the syntactic structure
of the text, has detected actual allegations that are "non-legal", such as references to Gospel passages.
In conclusion, the large amount of correct annotations and the limited efort required to correct the
erroneous ones, which will be the next step of the project’s activities, confirms the validity and potential
of the approach to annotation we presented.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>In the field of Legal History, the possibility of using an automatic annotation system for legal allegations
is proving to be a valuable tool. The development of coordinated and interconnected databases of
normative texts and their commentary apparatus, as well as the creation of digital editions of fundamental
works of medieval jurists, are just a few examples of potential applications. Moreover, the creation
of specific reference ontologies could enable the training of increasingly complex retrieval systems,
capable of extracting and sorting a vast amount of data that would otherwise be unmanageable, making
larger projects practically infeasible.</p>
      <p>In the case of study presented in this work, we have been able to annotate a large amount of text
exploiting the annotation a human expert on a tenth of the corpus and then a very efective machine
learning setup to have a complete, accurate annotation of the whole text. The outcome of this research
activity is twofold: we produced a valuable resource, the index, that will contribute to the GNORM
software and support studies on the Ordinary gloss and the Liber extra; and we have proved an efective
low-resource pipeline that can be replicated for similar activities in the field of Religious Studies, Legal
History, and many related disciplines in the humanities.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work was supported by project "Italian Strengthening of ESFRI RI RESILIENCE" (ITSERR) funded
by the European Union under the NextGenerationEU funding scheme (CUP:B53C22001770006).
Conference of the North American Chapter of the Association for Computational Linguistics:
Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational
Linguistics, Minneapolis, Minnesota, 2019, pp. 4171–4186. URL: https://aclanthology.org/N19-1423.
doi:10.18653/v1/N19-1423.
[23] D. Bamman, P. J. Burns, Latin bert: A contextual language model for classical philology, arXiv
preprint arXiv:2009.10053 (2020).
[24] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M.
Funtowicz, et al., Transformers: State-of-the-art natural language processing, in: Proceedings of the
2020 conference on empirical methods in natural language processing: system demonstrations,
2020, pp. 38–45.
[25] A. Esuli, F. Sebastiani, Evaluating information extraction, in: International Conference of the</p>
      <p>Cross-Language Evaluation Forum for European Languages, Springer, 2010, pp. 100–111.
[26] A. Esuli, V. R. Imperia, G. Puccetti, Automatic Annotation of the Legal References in the Liber
Extra’s Ordinary Gloss (1.0) [Data set], 2024. doi:10.5281/zenodo.14381709.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] ITSERR (Italian Strengthening of the ESFRI RI RESILIENCE</article-title>
          ),
          <year>2024</year>
          . URL: https://itserr.it.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <article-title>Corpus iuris canonici</article-title>
          , in: J.
          <string-name>
            <surname>Otaduy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Viana</surname>
            ,
            <given-names>J. S.</given-names>
          </string-name>
          Rueda (Eds.), Diccionario general de derecho canónico,
          <source>Thomson Reuters Aranzadi</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>757</fpage>
          -
          <lpage>765</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] Bernard of Parma, Glossa Ordinaria to Decretals Gregory IX, in: Decretales D. Gregorii Papae IX. suae integritati una cum glossis restitutae</article-title>
          .
          <source>Cum privilegio Gregorii XIII. Pont. Max</source>
          . et aliorum Principum, Romae, In Aedibus Populi Romani,
          <volume>1582</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bellomo</surname>
          </string-name>
          ,
          <source>The common legal past of Europe</source>
          ,
          <fpage>1000</fpage>
          -
          <lpage>1800</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Dolezalek</surname>
          </string-name>
          ,
          <article-title>Glosses and the juridical genre “apparatus glossarum” in the middle ages</article-title>
          ,
          <source>Rivista Internazionale di Diritto Comune</source>
          <volume>32</volume>
          (
          <year>2021</year>
          )
          <fpage>9</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Decretales</surname>
            <given-names>D.</given-names>
          </string-name>
          <article-title>Gregorii Papae IX. suae integritati una cum glossis restitutae</article-title>
          .
          <source>Cum privilegio Gregorii XIII. Pont. Max</source>
          . et aliorum Principum, Romae, In Aedibus Populi Romani,
          <volume>1582</volume>
          , available at UCLA Library Digital Collections„
          <year>2024</year>
          . URL: https://digital.library.ucla.edu/catalog/ark:/21198/ zz0014rx7w?cv=
          <fpage>35</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Hespanha</surname>
          </string-name>
          ,
          <article-title>Form and content in early modern legal books</article-title>
          ,
          <source>Rechtsgeschichte-Legal History</source>
          <volume>12</volume>
          (
          <year>2008</year>
          )
          <fpage>12</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Speciale</surname>
          </string-name>
          ,
          <article-title>Apparatus: ipertesto vivo e aperto, Ius Commune</article-title>
          .
          <source>Zeitschrift für Europäische Rechtsgeschichte</source>
          <volume>28</volume>
          (
          <year>2001</year>
          )
          <fpage>47</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H. U.</given-names>
            <surname>Kantorowicz</surname>
          </string-name>
          ,
          <article-title>Die allegationen im späteren mittelalter</article-title>
          ,
          <source>Archiv für Urkundenforschung</source>
          (
          <year>1935</year>
          )
          <fpage>15</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Quaglioni</surname>
          </string-name>
          ,
          <article-title>Licet allegare poetas: formanti letterari del diritto fra medioevo ed età moderna, Poesia e diritto nel Due e Trecento italiano</article-title>
          .
          <source>-(Memoria del tempo; 65)</source>
          (
          <year>2019</year>
          )
          <fpage>209</fpage>
          -
          <lpage>219</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Menzinger</surname>
          </string-name>
          ,
          <article-title>Reflections on the connection between author and text in medieval juridical production</article-title>
          ,
          <source>Historia et ius 11</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Menzinger</surname>
          </string-name>
          ,
          <article-title>The past, the others, himself: The open dialogue of a medieval legal author with his text</article-title>
          ,
          <source>in: Sicut dicit: Editing Ancient and Medieval Commentaries on Authoritative Texts</source>
          , Thurnout,
          <year>2019</year>
          , pp.
          <fpage>273</fpage>
          -
          <lpage>299</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Menzinger</surname>
          </string-name>
          ,
          <article-title>Interazione tra testo e 'citazione' nella dottrina giuridica civilistica: secoli XII e XIII, in: Juristische Glossierungstechniken als Mittel rechtswissenschaftlicher Rationalisierungen</article-title>
          , Erich Schmidt Verlag,
          <year>2022</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Weimar</surname>
          </string-name>
          , Argumenta brocardica,
          <source>Studia Gratiana</source>
          <volume>14</volume>
          (
          <year>1967</year>
          )
          <fpage>89</fpage>
          -
          <lpage>123</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Reno</surname>
          </string-name>
          , The Digital Decretals, https://www.digitaldecretals.com/,
          <year>2024</year>
          . [Online; accessed 1- November-2024].
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>J.-C. Klie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bugert</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Boullosa</surname>
            , R. E. de Castilho,
            <given-names>I. Gurevych</given-names>
          </string-name>
          ,
          <article-title>The inception platform: Machineassisted and knowledge-oriented interactive annotation</article-title>
          ,
          <source>in: Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations, Association for Computational Linguistics</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>9</lpage>
          . URL: http://tubiblio.ulb.tu-darmstadt.de/106270/, event Title:
          <source>The 27th International Conference on Computational Linguistics (COLING</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Esuli</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Puccetti,</surname>
          </string-name>
          <article-title>A Machine Learning pipeline to automatically annotate legal references (allegationes) in the Liber Extra's Ordinary Gloss</article-title>
          ,
          <year>2024</year>
          . URL: https://github.com/aesuli/CIC_ annotation. doi:
          <volume>10</volume>
          .5281/zenodo.14381817.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ratinov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <article-title>Design challenges and misconceptions in named entity recognition</article-title>
          ,
          <source>in: Proceedings of the thirteenth conference on computational natural language learning (CoNLL2009)</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          , et al.,
          <article-title>An introduction to conditional random fields, Foundations and Trends® in Machine Learning 4 (</article-title>
          <year>2012</year>
          )
          <fpage>267</fpage>
          -
          <lpage>373</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , et al.,
          <article-title>Language models are unsupervised multitask learners</article-title>
          ,
          <source>OpenAI blog 1</source>
          (
          <year>2019</year>
          )
          <article-title>9</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , in: J.
          <string-name>
            <surname>Burstein</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Doran</surname>
          </string-name>
          , T. Solorio (Eds.),
          <source>Proceedings of the 2019</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>