<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using and evaluating TRACER for an Index fontium computatus of the Summa contra Gentiles of Thomas Aquinas</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Greta Franzini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>greta.franzini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>marco.passarottig@unicatt.it</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Moritz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Bu¨ chler</string-name>
          <email>mbuechlerg@etrap.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Georg-August-Universita ̈t G o ̈ttingen</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This article describes a computational text reuse study on Latin texts designed to evaluate the performance of TRACER, a language-agnostic text reuse detection engine. As a case study, we use the Index Thomisticus as a gold standard to measure the performance of the tool in identifying text reuse between Thomas Aquinas' Summa contra Gentiles and his sources.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Thomas Aquinas (1225-1274) was a prolific
medieval author from Italy: his 118 works, known
as the Corpus Thomisticum, amount to 8,767,883
words
        <xref ref-type="bibr" rid="ref27">(Portalupi, 1994, p. 583)</xref>
        and discuss a
variety of topics, ranging from metaphysical to
legal, political and moral theory
        <xref ref-type="bibr" rid="ref18">(Kretzmann and
Stump, 1993)</xref>
        . The web of references to biblical,
ecclesiastical and classical literature that stretches
the whole Corpus Thomisticum speaks to
daunting erudition. In the late 1940s, Humanities
Computing pioneer Father Roberto Busa (1913-2011)
spearheaded a scholarly effort, known as the
Index Thomisticus, to manually annotate reuse, both
explicit (i.e., explicitly introduced by Aquinas as
a quote) and implicit (i.e., reference to works
without quotation), in the texts of Thomas Aquinas
        <xref ref-type="bibr" rid="ref4">(Busa, 1980)</xref>
        . Four decades later, Portalupi noted:
Ancora piu` difficile sara` [. . .] il
tentativo di confrontare automaticamente
tutto Tommaso con tutti i testi di uno
o piu` autori, per rintracciare in modo
globale la presenza implicita di una
fonte. Per fare questo occorrerebbe che
si verificassero due condizioni: in primo
luogo, gli autori di cui si studiano le
presenze implicite in Tommaso
dovrebbero essere informatizzati e interrogabili
nella totalita` delle loro opere; in secondo
luogo, bisognerebbe disporre di un
software molto potente e raffinato.
        <xref ref-type="bibr" rid="ref27">(Portalupi, 1994, p. 583)</xref>
        1
Today, a once visionary task is conceivable, giving
way to studies such as the present, which poses
the following research question: to which extent
can historical text reuse detection (HTRD)
software detect explicit and implicit text reuse in the
writings of Thomas Aquinas ? To this end, we test
the performance of TRACER, a text reuse
detection framework, for the creation of an Index
fontium computatus (a computed index of text reuse).
The Summa contra Gentiles (ScG) was chosen as a
case study because the critical edition used for the
Index Thomisticus, the 1961 Marietti Editio
Leonina
        <xref ref-type="bibr" rid="ref12">(Gauthier et al., 1882)</xref>
        , is still in use today
and because an ongoing treebanking effort of the
text will, in future, provide us with the linguistic
data needed to further refine the experiments
described here
        <xref ref-type="bibr" rid="ref22">(Passarotti, 2011)</xref>
        .
      </p>
      <p>1. Our English translation reads: ‘It will be even harder to
automatically compare all of Thomas against all of the texts
of one or multiple authors to check for the presence of
implicit sources. Such a task would only be possible under two
conditions: firstly, the texts of the authors quoted by Thomas
would have to be digitised and searchable in their entirety;
secondly, one would need very powerful and sophisticated
software’.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <sec id="sec-2-1">
        <title>The significance of text reuse</title>
        <p>
          Text reuse (TR) can be summarily described as
the written repetition or borrowing of text and can
take different forms. Bu¨chler et al. (2014)
separate syntactic TR, such as (near-)verbatim
quotations or idiomatic expressions, from semantic TR,
which can manifest itself as a paraphrase, an
allusion or other loose reproduction. The study of
quotation is key to any philological examination
of a text, as it is not only indicative of the
intellectual and cultural endowment of an author, but
may shed light on the sources used, the relation
between works and literary influence. Crucially,
quotations may also preserve text that is now lost,
thus facilitating efforts of textual reconstruction. 2
Owing to the magnitude of the task, the
publication of a work’s complete index of references,
conventionally known as Apparatus fontium or
Index scriptorum, is rare
          <xref ref-type="bibr" rid="ref27">(Portalupi, 1994, p. 582)</xref>
          .
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Text reuse in Thomas Aquinas</title>
        <p>
          Like many of his Christian predecessors,
Aquinas’ body of work teems with references to secular
and Christian literature alike. In the ScG
(12591265) Aquinas cites 170 works both explicitly and
implicitly
          <xref ref-type="bibr" rid="ref12">(Gauthier et al., 1882, Vols. IV-XV)</xref>
          .
Explicit quotations provide information about the
source text and the author and/or work, and can
either be direct or indirect
          <xref ref-type="bibr" rid="ref12">(Gauthier et al., 1882,
vol. XVI, pp. XVI-XXII)</xref>
          . Implicit reuses, in the
ScG and in general, are more elusive, as they are
almost never syntactically nor lexically-faithful to
the original text, thus making them hard for both
machines and humans to spot
          <xref ref-type="bibr" rid="ref27">(Portalupi, 1994, p.
582)</xref>
          . 3 Durantel notes that Aquinas’ tendency in
TR is to borrow only what is necessary to fit the
flow of his narrative without significant semantic
or syntactic deviation from the original
          <xref ref-type="bibr" rid="ref8">(Durantel, 1919, p. 63)</xref>
          . And yet, Pelster’s observation
on Aquinas’ paraphrastic reuse of Aristotle might
suggest greater deviation
          <xref ref-type="bibr" rid="ref26">(Pelster, 1935, p. 331)</xref>
          . 4
2. One notable example is the fragmentary survival of
Alexandrian scholarship at the hands of Roman philologists
(who wrote commentaries known as scholia) and
grammarians
          <xref ref-type="bibr" rid="ref31">(Turner, 2014, p. 16)</xref>
          .
        </p>
        <p>
          3. For problems with implicit quotations, see
          <xref ref-type="bibr" rid="ref14">(Haverfield,
1916, p. 197)</xref>
          and
          <xref ref-type="bibr" rid="ref10">(Fowler, 1997, p. 15)</xref>
          . For automatic
allusion detection, see (Bamman and Crane, 2008).
        </p>
        <p>4. “Da Thomas die Schriften des Aristoteles [. . .]
gewo¨hnlich nur dem Gedanken nach, nicht wo¨rtlich anfu¨hrt.”
In English: ‘Since Thomas usually quotes paraphrastically,
not literally.’</p>
        <p>
          Roberto Busa’s effort in the late 1940s
resulted in the creation of the Index Thomisticus, a
manually-lemmatised version of Thomas
Aquinas’ opera omnia
          <xref ref-type="bibr" rid="ref16">(Jones, 2016)</xref>
          . Among the
annotations, the Index Thomisticus tags tokens
forming explicit quotations as QL if literal (ad
litteram) and QS if a paraphrase (ad sensum), and
tokens forming implicit quotations as QR to indicate
a reference or citation alluding to another text. An
example quotation in the ScG containing a mixed
annotation is:
[. . .] ratio(QL) vero (QL)
significata(QL) per(QL) nomen(QL)
est(QL) definitio(QL)
secundum(QR) philosophum(QR) in(QR)
IV(QR) Metaph.(QR) 5
        </p>
        <p>The (QL) portion of this example contains the
literal quote, while the second (QR) portion
provides the reference.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Historical text reuse detection</title>
        <p>
          HTRD is a Natural Language Processing (NLP)
task aimed at identifying syntactic and semantic
TR in historical sources. The computational
analysis of historical languages is particularly
challenging as tools at our disposal are often trained
on a synchronic rather than diachronic state of
a language 6 and on controlled textual corpora.
Eger et al. (2015) and Passarotti (2010) tested
the performance of seven different taggers,
including TreeTagger
          <xref ref-type="bibr" rid="ref28">(Schmid, 1994)</xref>
          , for different
training sets and tag-sets of medieval (church)
Latin texts showing accuracies tightly below 96%
and 96:75% for PoS-tagging, and around 90% and
89:90% for morphological analysis, respectively.
These results have yet to be generalised to other
variants of Latin and can be improved upon with
the provision of additional training corpora,
treebanked and semantically-tagged, the creation of
corpora containing intertexts, or with the
expansion of lexical resources, such as the Latin
WordNet
          <xref ref-type="bibr" rid="ref17 ref19">(Minozzi, 2017, p. 130)</xref>
          .
        </p>
        <p>
          The extent to which the limitations of these
resources and taggers (e.g., correct resolution of
homographs) affect HTRD tools, including
Tesserae
          <xref ref-type="bibr" rid="ref5">(Coffee et al., 2013)</xref>
          , Passim
          <xref ref-type="bibr" rid="ref29">(Smith et al.,
2015)</xref>
          7 and TRACER
          <xref ref-type="bibr" rid="ref3">(Bu¨chler, 2013)</xref>
          is not yet
5. Book 1, chap. 12, n. 4. Our English translation reads:
‘[. . .] according to the philosopher in Metaph. IV, the
meaning of a name is its definition’.
        </p>
        <p>
          6. See Janda and Joseph (2005) for the dichotomy.
7. https://github.com/dasmiq/passim
fully understood. Reasons for this are the
field’s lack of progress caused by “inconsistent
standards and the scattering of insights across
publications”
          <xref ref-type="bibr" rid="ref6">(Coffee, 2018)</xref>
          , the general failure of
HTRD studies to publish negative results, and the
quasi-absence of gold standards for testing. To our
knowledge, the only projects to have published
computed results from intertextual studies on
historical sources are the Proteus Project (English
and Latin)
          <xref ref-type="bibr" rid="ref32">(Yalniz et al., 2011)</xref>
          , the Chinese Text
Project (early Chinese)
          <xref ref-type="bibr" rid="ref30">(Sturgeon, 2017)</xref>
          ,
Commonplace Cultures (English and Latin) (Gladstone
and Cooney, forthcoming), SHEBANQ (Hebrew)
          <xref ref-type="bibr" rid="ref20">(Naaijer and Roorda, 2016)</xref>
          , Samtla (Search and
Mining Tools for Language Archives)
(languageindependent)
          <xref ref-type="bibr" rid="ref13">(Harris et al., 2018)</xref>
          , and Tesserae
(Latin), but of these only the latter discloses tool
configurations.
3
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <sec id="sec-3-1">
        <title>Gold Standard</title>
        <p>To facilitate the classification of
automaticallydetected reuse, all QL-, QS- and QR-annotated
tokens were extracted from the Index Thomisticus.
Of the total 24,416 sentences constituting the ScG,
the 7,396 (30:29%) containing any combination
of QL, QS and QR were stored in a tabular file,
which we define as the Index Thomisticus Gold
Standard of TR (hereafter IT-GS). The number of
sentences containing only QL tokens (1,139)
compared to that of sentences containing only QS
tokens (2,270) corroborates expert assertions about
Aquinas’ paraphrastic style of TR.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Text acquisition and preparation</title>
        <p>
          For the sake of processing efficiency, out of the
ScG’s 170 source works we began with a set of
five readily available texts. These are Philosophiae
Consolationis and De Trinitate of Boethius, De
Deo Socratis of Apuleius, Cicero’s De Divinatione
and the Moerbeke Latin translation of Aristotle’s
Metaphysica. The texts were acquired from
different sources and cleaned of all paratextual
information. The clean texts were then
segmentised by sentence, PoS-tagged and lemmatised with
the TreeTagger Brandolini parameter file (with an
average accuracy of 93:72%), whose tag-set
provides the degree of granularity needed in this
experiment. 8 Finally, a script was used to format
sen8. The Brandolini tag-set was manually mapped against
that of Morpheus
          <xref ref-type="bibr" rid="ref7">(Crane, 1991)</xref>
          , which TRACER uses as a
tences to TRACER requirements.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Text reuse detection with TRACER</title>
        <p>The HTRD on this corpus was performed
(server-side) with TRACER, a language-agnostic
framework comprising hundreds of information
retrieval (IR) algorithms designed to work with
historical and modern languages alike. 9 TRACER
is a Java command-line tool driven by an XML
configuration file, which users can modify to fit
their detection needs. TRACER follows a
sixstep architecture, 10 which demystifies the
detection process by storing the computed output of
each step on the disk so that users can more easily
follow and locate errors in the processing chain,
if any. TRACER is resilient to OCR-noise and
capable of detecting both (near-)verbatim quotations
and looser forms of TR. The detection of
paraphrase requires the use of linguistic resources to
help TRACER match a word against its synsets
and an inflected form against its base-form. For
synonym detection, we extracted synonymous
relations from the Latin WordNet. TR identified with
TRACER was manually compared against the
ITGS to separate the True (TP) from the False
Positives (FP), and to identify False Negatives (FN).
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <sec id="sec-4-1">
        <title>Philosophiae Consolationis</title>
        <p>To detect both verbatim quotations and
paraphrase, TRACER was optimised for recall over
precision and configured to work with single
words as features, to ignore the top 20% most
frequent words, 11 to link text pairs with a
minimum overlap of 5 features, 12 to expand the query
to synonyms, and to return only those aligned text
pairs presenting an overall sentence similarity of
at least 50%. 13 Of the eight reuses indicated in
the Editio Leonina, we were unable to precisely
locate one as it alludes to four paragraphs of
text; 14 of the remaining seven, as shown in Figure
1, TRACER identified three (42%). Upon close
inspection, two FNs were affected by the 20%
threshold of feature removal, for example:
Boethius 1.4.105 Unde haud iniuria tuorum
quidam familiarium quaesivit: “Si quidem deus”,
inquit, “est, unde mala ? 15
Aquinas 3.71.10 , introducit quendam
philosophum quaerentem: si deus est, unde malum ? 16</p>
        <p>Here, the tokens si, est and unde were ignored as
they fell within the pool of the 20% most frequent
words removed.</p>
        <p>One reuse was successfully identified on the
basis of feature overlap but did not amount to a 50%
sentence similarity; and the fourth reuse could
not be identified because of a missing
synonymous relation in the Latin WordNet (i.e.,
gaudiumbeatitudo) 17 and its insufficient feature overlap.
The resulting F1-score is 4; 6 10 3.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>De Trinitate</title>
        <p>
          Given the results of the previous analysis, for
this second investigation the feature removal and
the sentence similarity values were lowered to
10% and 40% respectively, thus optimising for
even higher recall (10,349 total sentences aligned).
Of the four known reuses, TRACER identified
three. The 40% similarity threshold was essential
to the identification of one reuse (where the score
is 0.4375); the FN, which was indeed found on the
basis of an eight-word overlap but did not meet
the minimum sentence similarity threshold,
revealed another missing synonymous relation in the
WordNet (i.e., disciplinatus-eruditus) 18 and a
failed alignment of the variants temptare (Boethius)
and tentare (Aquinas) owing to inconsistent
TreeTagger lemmatisation (tempto and tento,
respeclength
          <xref ref-type="bibr" rid="ref1">(Broder, 1997)</xref>
          .
        </p>
        <p>14. This reuse would have doubtless been overlooked by
TRACER too owing to the absence of features to compare.</p>
        <p>15. Our English translation reads: ‘It is not wrong that a
certain acquaintance of yours has questioned: ‘If in fact God
exists,’ he asks, ‘where is evil from ?”</p>
        <p>16. Our English translation reads: ‘(Boethius) introduces
a certain philosopher who asks: ‘If God exists, where is evil
from ?’.’</p>
        <p>17. Incidentally, this relation is also not mapped in
BabelNet (bn:00042905n) nor in ConceptNet (http://
conceptnet.io/c/la/gaudium) (as of 8 June 2018).
18. Also not present in neither BabelNet nor ConceptNet.
tively). The F1-score for this analysis was 5; 6
10 4.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>De Deo Socratis</title>
        <p>This work of Apuleius is quoted twice in the
ScG. Of the two reuses, TRACER was able to
detect one in full and only parts of the second. The
second reuse spans three sentences and is mostly
paraphrastic, with only three words annotated in
the Index Thomisticus as QL (sunt animo
passiva). 19 To capture the fullest range of reuse
diversity, TRACER’s feature removal was set to 10%,
the overlap to 3 and the overall similarity to 20%.
However, as sunt (form of the verb sum ‘to be’)
is the most frequent word across the texts,
TRACER’s inbuilt feature removal prevented the
detection of the short QL portion of the reuse; the
QR+QS portions, on the other hand, were
successfully detected. We counted both results as TPs,
resulting in an F1-score of 2; 6 10 5.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>De Divinatione</title>
        <p>The only recorded reuse that Aquinas makes of
Cicero’s text is implicit and alludes to a block of
text, making it difficult to manually pinpoint with
precision. To detect as loose a similarity as
possible, the TRACER search was cast with the same
configuration used in the previous analysis. No
reuse, however, was found.
4.5</p>
      </sec>
      <sec id="sec-4-5">
        <title>Metaphysica</title>
        <p>The Editio Leonina lists 97 reuses of
Aristotle’s Metaphysica. As previously mentioned,
Pelster describes Aquinas’ reuse of the Latin
translation of the Metaphysica as more paraphrastic
than literal. Our manual examination of the texts
and the results of TRACER confirmed this
observation, in that we could not manually locate
seven reuses (due to their strong allusiveness) and
a fault-tolerant TRACER configuration (removal
of the top 10% most frequent words, overlap of 3
features and an overall sentence similarity of 40%)
yielded 19 TPs only (6 out of 15 QL 20 and 13 out
of 75 QR+QS). The F1-score resulting from this
analysis is 3; 8 10 4.</p>
        <p>
          19. [daemones] [. . .] sunt animo passiva or ‘demons are
emotional in mind’
          <xref ref-type="bibr" rid="ref17">(Jones, 2017, pp. 372-373)</xref>
          .
        </p>
        <p>20. The QL quotations in the ScG seem to refer to a
different Latin translation than that available to us, which would
explain why some instances of QL went undetected.</p>
        <p>Our results show that the FNs emerging from
the computational analyses were largely caused
by Aquinas’ paraphrastic and allusive TR style,
which at times challenged our own ability to spot
similarities, even with the help of the critical
edition. The allusions that we could identify generally
retain the semantics of the alluded-to texts, thus
confirming Durantel’s insights. While a number of
these negative results were also directly tied to
lacunae in the Latin WordNet and to inconsistent
lemmatisation, the flexibility and methodological
transparency of TRACER allowed us to locate
error sources and accordingly tune configurations to
work around these issues (e.g., by increasing the
feature overlap and/or lowering the sentence
similarity scoring thresholds). Notwithstanding,
TRACER’s panlingual feature removal parameter
affected the retrieval of shorter instances of reuse,
particularly those containing forms of the highly
frequent verb sum.</p>
        <p>The manual evaluation of TRACER results
against the IT-GS for the creation of an Index
fontium computatus was time-consuming, not least
because of a number of reference inaccuracies in
the critical edition itself (in one case, the reference
is off by ten lines). Nevertheless, the creation of
the index is proving essential to the assessment of
TRACER’s fitness for purpose on Latin texts.</p>
        <p>
          As far as the usability of the tool is concerned,
TRACER’s detection power is offset by its
cumbersome setup, which is unfriendly to those who
are not familiar with the command line, NLP
basics and/or Java (stack traces). This issue is being
addressed with the development of a user manual
          <xref ref-type="bibr" rid="ref11">(Franzini et al., 2018)</xref>
          .
6
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>
        This article describes a computational text reuse
study on Latin texts designed to evaluate the
performance of TRACER, a language-agnostic IR
text reuse detection engine. The results obtained
were manually evaluated against a gold standard
and are contributing to the creation of an Index
fontium computatus to both assess TRACER’s
efficacy and to provide a test-bed against which
analogous IR systems can be measured and thus
compared to TRACER. Our study shows that despite
the known limitations of existing linguistic
resources for Latin, the diverse spectrum of
paraphrastic reuse encountered and its own
languageagnosticism, TRACER is equipped to detect a
wide range of explicit text reuse in the ScG, be
that short or long, verbatim or paraphrastic, and
implicit reuse only if coupled with explicit. To
increase the detection accuracy, we are
implementing a black/white list to give users the power
to control words or multi-word expressions to be
ignored or retained in the detection; furthermore,
we plan on re-running these analyses with the
disambiguated linguistic annotation currently being
added to the text of the ScG
        <xref ref-type="bibr" rid="ref25">(Passarotti, 2015)</xref>
        to
measure its impact on this particular IR task.
      </p>
      <p>The data used and generated in the current
study is available from: https://github.
com/CIRCSE/text-reuse-aquinas.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The authors would like to thank Eleonora Litta
for proofreading this article and the anonymous
reviewers for their valuable comments. This
research was funded by the German Federal
Ministry of Education and Research (No. 01UG1409).
David Bamman and Gregory Crane. 2008. The Logic
and Discovery of Textual Allusion. In Proceedings
of the ACL Workshop LaTeCH - Language
Technology for Cultural Heritage Data. ACL. http:
//hdl.handle.net/10427/42685.</p>
      <p>Clovis Gladstone and Charles Cooney. forthcoming.</p>
      <p>Opening New Paths for Scholarship: Algorithms to
Track Text Reuse in ECCO. Digitizing
Enlightenment.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Andrei Z.</given-names>
            <surname>Broder</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>On the resemblance and containment of documents</article-title>
          .
          <source>In Proceedings of the Compression and Complexity of Sequences</source>
          <year>1997</year>
          , SEQUENCES '
          <volume>97</volume>
          , pages
          <fpage>21</fpage>
          -
          <lpage>29</lpage>
          , Washington, DC, USA. IEEE Computer Society. http://dl.acm.org/citation. cfm?id=
          <volume>829502</volume>
          .
          <fpage>830043</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Bu</surname>
          </string-name>
          ¨chler, Philip R. Burns, Martin Mu¨ller, Emily Franzini, and
          <string-name>
            <given-names>Greta</given-names>
            <surname>Franzini</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Towards a Historical Text Re-use Detection</article-title>
          . In Chris Biemann and Alexander Mehler, editors,
          <source>Text Mining</source>
          , pages
          <fpage>221</fpage>
          -
          <lpage>238</lpage>
          . Springer International Publishing, Cham. http://link.springer.com/10. 1007/978-3-
          <fpage>319</fpage>
          -12655-5_
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Bu</surname>
          </string-name>
          ¨chler.
          <year>2013</year>
          .
          <article-title>Informationstechnische Aspekte des Historical Text Re-use</article-title>
          .
          <source>PhD Thesis</source>
          . http://www.qucosa.de/fileadmin/ data/qucosa/documents/10851/ Dissertation.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Busa</surname>
          </string-name>
          .
          <year>1980</year>
          .
          <article-title>The annals of humanities computing: The Index Thomisticus</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>14</volume>
          (
          <issue>2</issue>
          ):
          <fpage>83</fpage>
          -
          <lpage>90</lpage>
          , October. http:// www.jstor.org/stable/30207304.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Neil</given-names>
            <surname>Coffee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jean-Pierre</surname>
            <given-names>Koenig</given-names>
          </string-name>
          , Shakthi Poornima, Christopher W. Forstall, Roelant Ossewaarde,
          <string-name>
            <given-names>and Sarah L.</given-names>
            <surname>Jacobson</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The Tesserae Project: intertextual analysis of Latin poetry</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          ,
          <volume>28</volume>
          :
          <fpage>221</fpage>
          -
          <lpage>228</lpage>
          . https://doi. org/10.1093/llc/fqs033.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Neil</given-names>
            <surname>Coffee</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>An Agenda for the Study of Intertextuality</article-title>
          .
          <source>Transactions of the American Philological Association</source>
          ,
          <volume>148</volume>
          :
          <fpage>205</fpage>
          -
          <lpage>223</lpage>
          . https://muse. jhu.edu/article/693654.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Gregory</given-names>
            <surname>Crane</surname>
          </string-name>
          .
          <year>1991</year>
          .
          <article-title>Generating and Parsing Classical Greek</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          , page
          <volume>243</volume>
          -
          <fpage>245</fpage>
          . https://doi.org/10.1093/ llc/6.4.243.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Jean</given-names>
            <surname>Durantel</surname>
          </string-name>
          .
          <year>1919</year>
          . Saint Thomas et le Pseudo-Denis.
          <source>Librairie Fe´lix Alcan</source>
          , Paris. http://archive.org/details/ cuasaintthomaset00dura.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Steffen</given-names>
            <surname>Eger</surname>
          </string-name>
          ,
          <article-title>Tim vor der Bru¨ck, and</article-title>
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Mehler</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Lexicon-assisted tagging and lemmatization in Latin: A comparison of six taggers and two lemmatization methods</article-title>
          .
          <source>In In Proceedings of the 9th SIGHUM Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities</source>
          , pages
          <fpage>105</fpage>
          -
          <lpage>113</lpage>
          . http://www.aclweb. org/anthology/W15-3716.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Don</given-names>
            <surname>Fowler</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>On the Shoulders of Giants: Intertextuality and Classical Studies. Materiali e discussioni per l'analisi dei testi classici</article-title>
          ,
          <volume>39</volume>
          :
          <fpage>13</fpage>
          -
          <lpage>34</lpage>
          . http://www.jstor.org/ stable/40236104.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Greta</given-names>
            <surname>Franzini</surname>
          </string-name>
          , Emily Franzini, Kirill Bulert, Marco Bu¨chler, and
          <string-name>
            <given-names>Maria</given-names>
            <surname>Moritz</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>TRACER: A User Manual</article-title>
          . https://tracer.gitbook. io/-manual/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Gauthier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Bataillon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          , T. de Vio Cajetan,
          <article-title>Commissio Leonina, and</article-title>
          <string-name>
            <given-names>Dominicans. 1882. Sancti</given-names>
            <surname>Thomae Aquinatis Doctoris Angelici Opera Omnia iussu edita Leonis XIII P.M. Ex Typographia Polyglotta S.C. de Propaganda</surname>
          </string-name>
          <string-name>
            <surname>Fide</surname>
          </string-name>
          , Rome.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Martyn</given-names>
            <surname>Harris</surname>
          </string-name>
          , Mark Levene, Dell Zhang, and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Levene</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Finding Parallel Passages in Cultural Heritage Archives</article-title>
          .
          <source>Journal on Computing and Cultural Heritage</source>
          ,
          <volume>11</volume>
          (
          <issue>3</issue>
          ):
          <volume>15</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          :
          <fpage>24</fpage>
          . http:// doi.acm.
          <source>org/10</source>
          .1145/3195727.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Francis John Haverfield</surname>
          </string-name>
          .
          <year>1916</year>
          .
          <article-title>Tacitus during the Late Roman Period and the Middle Ages</article-title>
          .
          <source>The Journal of Roman Studies</source>
          ,
          <volume>6</volume>
          :
          <fpage>196</fpage>
          -
          <lpage>201</lpage>
          . https://doi. org/10.2307/296272.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Richard D.</given-names>
            <surname>Janda</surname>
          </string-name>
          and
          <string-name>
            <given-names>Brian D.</given-names>
            <surname>Joseph</surname>
          </string-name>
          .
          <year>2005</year>
          . On Language, Change, and
          <string-name>
            <surname>Language</surname>
            Change - Or, Of History, Linguistics, and
            <given-names>Historical</given-names>
          </string-name>
          <string-name>
            <surname>Linguistics</surname>
          </string-name>
          . In Brian D. Joseph and Richard D. Janda, editors,
          <source>The Handbook of Historical Linguistics</source>
          , pages
          <fpage>3</fpage>
          -
          <lpage>181</lpage>
          . Wiley-Blackwell, Oxford.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Steven E. Jones. 2016. Roberto</given-names>
            <surname>Busa</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. J.,</surname>
          </string-name>
          <article-title>and the Emergence of Humanities Computing: The Priest and the Punched Cards</article-title>
          . Routledge, March.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Christopher P.</given-names>
            <surname>Jones</surname>
          </string-name>
          , editor.
          <source>2017. Apuleius. Apologia</source>
          . Florida. De Deo Socratis, volume
          <volume>534</volume>
          of Loeb Classical Library. Harvard University Press, Loeb Classical Library.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Norman</given-names>
            <surname>Kretzmann</surname>
          </string-name>
          and Eleonore Stump, editors.
          <year>1993</year>
          . The Cambridge Companion to Aquinas. Cambridge University Press, Cambridge; New York, May.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Minozzi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Latin WordNet, una rete di conoscenza semantica per il latino e alcune ipotesi di utilizzo nel campo dell'Information Retrieval</article-title>
          . In Paolo Mastandrea, editor,
          <source>Strumenti digitali e collaborativi per le Scienze dell'Antichita`, number 14 in Antichistica</source>
          , pages
          <fpage>123</fpage>
          -
          <lpage>134</lpage>
          . http://doi. org/10.14277/
          <fpage>6969</fpage>
          -182-9/ANT-14-10.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Martijn</given-names>
            <surname>Naaijer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dirk</given-names>
            <surname>Roorda</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Parallel Texts in the Hebrew Bible, New Methods and Visualizations</article-title>
          . CoRR, abs/1603.01541. http://arxiv. org/abs/1603.01541.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Passarotti</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Leaving behind the lessresourced status. The case of Latin through the experience of the Index Thomisticus Treebank</article-title>
          . In 7th SaLTMiL Workshop on Creation and
          <article-title>use of basic lexical resources for less-resourced languages LREC 2010, Valletta</article-title>
          , Malta, 23 May
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Passarotti</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Language Resources. The State of the Art of Latin and the Index Thomisticus Treebank Project</article-title>
          . In Marie-Sol Ortola, editor, Corpus anciens et Bases de donne´es, ALIENTO.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <article-title>E´changes sapientiels en Me´diterrane´e</article-title>
          , volume
          <volume>2</volume>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          pages
          <fpage>301</fpage>
          -
          <lpage>320</lpage>
          . Presses universitaires de Nancy, Nancy.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Passarotti</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>What you can do with linguistically annotated data. From the Index Thomisticus to the Index Thomisticus Treebank</article-title>
          . In Vijgen Roszak Piotr, editor, Reading Sacred Scripture with Thomas Aquinas. Hermeneutical Tools,
          <source>Theological Questions and New Perspectives</source>
          , pages
          <fpage>3</fpage>
          -
          <lpage>44</lpage>
          . Brepols.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Pelster</surname>
          </string-name>
          .
          <year>1935</year>
          .
          <article-title>Die Uebersetzungen der aristotelischen Metaphysik in den Werken des hl</article-title>
          .
          <source>Thomas von Aquin: Ein Beitrag. Gregorianum</source>
          ,
          <volume>16</volume>
          (
          <issue>3</issue>
          ):
          <fpage>325</fpage>
          -
          <lpage>348</lpage>
          . http://www.jstor. org/stable/23567607.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Enzo</given-names>
            <surname>Portalupi</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>L'uso dell'“Index Thomisticus” nello studio delle fonti di Tommaso d'Aquino: Considerazioni generali e questioni di metodo</article-title>
          .
          <source>Rivista di Filosofia Neo-Scolastica</source>
          ,
          <volume>86</volume>
          (
          <issue>3</issue>
          ):
          <fpage>573</fpage>
          -
          <lpage>585</lpage>
          . http://www.jstor.org/ stable/43062344.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Probabilistic Part-ofSpeech Tagging Using Decision Trees</article-title>
          .
          <source>In Proceedings of International Conference on New Methods in Language Processing</source>
          , Manchester, UK. http://www.cis.uni-muenchen. de/˜schmid/tools/TreeTagger/data/ tree-tagger1.
          <fpage>pdf</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>David A.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ryan</given-names>
            <surname>Cordell</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Abby</given-names>
            <surname>Mullen</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Computational Methods for Uncovering Reprinted Texts in Antebellum Newspapers</article-title>
          . American Literary History,
          <volume>27</volume>
          (
          <issue>3</issue>
          ):
          <fpage>E1</fpage>
          -
          <lpage>E15</lpage>
          . http://dx. doi.org/10.1093/alh/ajv029.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Donald</given-names>
            <surname>Sturgeon</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Unsupervised identification of text reuse in early Chinese literature. Digital Scholarship in the Humanities</article-title>
          . https://academic.oup.com/dsh/ advance-article/doi/10.1093/llc/ fqx024/4583485.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <given-names>James</given-names>
            <surname>Turner</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Philology: The Forgotten Origins of the Modern Humanities</article-title>
          . Princeton University Press, Princeton and Oxford.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <given-names>Ismet</given-names>
            <surname>Zeki</surname>
          </string-name>
          <string-name>
            <surname>Yalniz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ethem F.</given-names>
            <surname>Can</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Manmatha</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Partial Duplicate Detection for Large Book Collections</article-title>
          .
          <source>In Proceedings of the 20th ACM International Conference on Information and Knowledge Management</source>
          ,
          <source>CIKM '11</source>
          , pages
          <fpage>469</fpage>
          -
          <lpage>474</lpage>
          . http://doi.acm.
          <source>org/10</source>
          .1145/ 2063576.2063647.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>