<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Computational Humanities Research Conference, November</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Obtaining More Expressive Corpus Distributions for Standardized Ancient Languages</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oliver Hellwig</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sven Sellmer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Nehrdich</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Comparative Language Science, University of Zürich</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Language and Information, Heinrich Heine Universität</institution>
          ,
          <addr-line>Düsseldorf</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Oriental Studies, Adam Mickiewicz University</institution>
          ,
          <addr-line>Poznań</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Khyentse Center for Tibetan Buddhist Textual Scholarship, Universität Hamburg</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>1</volume>
      <issue>4</issue>
      <fpage>7</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>This paper introduces a latent variable model for ancient languages that aims at quantifying the influence that early authoritative works exert on their literary successors in terms of lexis. The model jointly estimates the amount of word reuse, based on uni- and bigrams of words, and the date of composition of each text. We apply the model to a corpus of pre-Renaissance Latin texts composed between the 3rd c. BCE and the 14th c. CE. Our evaluation focusses on the structures of word reuse detected by the model, its temporal predictions and the quality of the inferred diachronic distributions of words, which last aspect is assessed using a newly designed task from the field of computational etymology.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Text reuse</kwd>
        <kwd>citations</kwd>
        <kwd>standardized languages</kwd>
        <kwd>historical corpora</kwd>
        <kwd>Bayesian mixture model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>standardized languages reflects the everyday language use of authors who often spoke ‘vulgar’
varieties of the standardized language or, later, vernacular languages stemming from these
varieties.</p>
      <p>
        Two factors are especially relevant here. First, many authors writing in standardized
languages show the tendency to reuse and paraphrase authoritative works which were considered
as a kind of gold standard (see e.g. Lee [
        <xref ref-type="bibr" rid="ref23">22</xref>
        ]; Roberts [34] for Latin). The influence of earlier
works can therefore bias and distort the distributions of words found in later ones. More
generally, such languages typically are conservative in that they preserve words that are no longer
current outside literary circles. An instructive example for such a trend is the Latin word equus
‘horse’ [see 9, p. 291]. While this word is the standard expression for ‘horse’ in Classical Latin
and does not have any archaic ring to it, the Romance languages, which originated from Latin
dialects spoken in the late Antiquity (“Vulgar Latin”, see Herman [
        <xref ref-type="bibr" rid="ref16">15</xref>
        ]), derive their words for
‘horse’ from Latin caballus (e.g. Fr. cheval, It. cavallo), which suggests that occurrences of
equus in post-Classical Latin texts no longer reflect the spoken language.
      </p>
      <p>
        Second, the temporal structure of ancient corpora may be (partly) unclear, making it even
more difficult to reliably construct diachronic lexical trajectories. The accumulated efects that
standardization, text reuse, semantic conservatism and temporal uncertainties exert on corpus
distributions are difficult to determine from the raw corpus data alone, making it necessary
to balance the corpus evidence with detailed qualitative – and time-consuming – studies of
individual words. Such issues are not restricted to Latin texts of the late Antiquity and
the Middle Ages [see e.g. 23], but are also found, for example, in Buddhist Chinese [see e.g.
28], in the Indic corpora composed in Sanskrit and Pāli, or in Classical Chinese [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ]. As these
languages are the ancestors of important modern language families, Classical Studies as well as
Linguistics can benefit from corpus distributions that distinguish between the actual language
use and the influence of authoritative works. 2
      </p>
      <p>
        This paper discusses a Bayesian mixture model for lemmatized texts that disentangles the
influence exerted by authoritative, frequently cited and paraphrased texts on the word usage
encountered in their literary successors. It aims at generating a clearer picture of the actual
practice in standardized languages, at quantifying the amount of word reuse and at unveiling
intellectual lineages in such corpora. For modelling word reuse, this paper builds on previous
research that quantifies the influence of cited authors in the context of scientific publications
[
        <xref ref-type="bibr" rid="ref28 ref8">8, 27</xref>
        ]. Unlike such bibliometric studies, citations in ancient corpora are mostly not (clearly)
marked as such and must therefore be inferred from the data in the approach presented in
this paper. The detection of literary influences can be further enhanced by inspecting lexical
n-grams. While many previous approaches represent the textual data as bags of words, one
may argue that text reuse and stylistic influences rather get manifest in collocations taken over
from earlier literary works. While the presence of the unigrams aurum ‘gold’ and pretiosus
‘precious’ only gives a weak indication of literary ancestry, a bigram formed of these two words
(in pretiosior auro ‘more precious than gold’) is a much clearer indication that the late Roman
author Maximianus has been influenced by the Augustan poet Ovid. Our model therefore
complements the bag of words representation with lexical bigrams [see 46] and makes the
decision for uni- or bigrams part of the inference process.
      </p>
      <p>
        Another important aspect is the time of composition. Most (Bayesian) mixture models with
2The expression ‘actual language use’ has to be taken in a technical sense that changes according to the
author: For authors speaking some form of Latin, it refers to the language they use in everyday situations; for
users of other languages, it denotes, somewhat artificially, the Latin they write, but from which the efects of
word reuse have been removed, so to speak.
a temporal component assume that the time of composition is an observed variable (e.g. Blei
and Laferty [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Wang, Blei, and Heckerman [
        <xref ref-type="bibr" rid="ref44">44</xref>
        ]). Such an assumption does not hold for
many ancient texts as their dates are either unknown or still under scrutiny. While there
exist some Latin texts whose dates of composition are strongly disputed (see e.g. Laurioux [
        <xref ref-type="bibr" rid="ref22">21</xref>
        ]
on the cookbook of Apicius), this problem is more urgent for ancient Indian corpora, where
dates proposed for early texts are often just educated guesses (see e.g. Olivelle [
        <xref ref-type="bibr" rid="ref32">30</xref>
        ], 7-13 on
the Sanskrit philosophical texts called Upaniṣads). We address this issue by modelling the
time of composition of each text as a latent variable that conditions the observed features
and incorporates the current state of scholarly research with the help of a temporal prior (see
Hellwig [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for a related approach for Vedic Sanskrit).
      </p>
      <p>We use Latin texts composed between the 3rd c. BCE and the 14th c. CE as a test
case. As Sec. 5 will show, many aspects of the evaluation rely on qualitative arguments, as
gold standards for these tasks are currently not available. Using the Latin corpus ofers the
advantage that the evaluation can build on a long history of literary and linguistic research,
so that our results can be compared against an extensive record of previous scholarship. The
initial application to the well-researched Latin tradition makes it easier to transfer the methods
developed here to more disputed textual traditions of South Asia.</p>
      <p>After an overview of related work in Computational Linguistics (Sec. 2), Sections 3 and
4 describe the data and the model. Section 5 assesses various choices in the model design
using posterior predictive checks (Sec. 5.1) and presents an evaluation of three prominent
aspects of our model: word reuse (Sec. 5.2), predicted times (Sec. 5.3) and the inferred corpus
distributions (Sec. 5.4), the latter being tested on a new task in computational etymology.
– Data and scripts are available at https://github.com/OliverHellwig/sanskrit/tree/master/
papers/chr2021.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related research</title>
      <p>
        Our model of word reuse builds on previous work on detecting citation activities in scientific
literature. Such activities have repeatedly been formalized using (ad-)mixture models, starting
with Cohn and Hofmann [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] whose generative model conditions citations on the presence
of hidden topics. Erosheva, Fienberg, and Laferty [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] extend Latent Dirichlet Allocation
by conditioning the generation of links on the same document-specific topic distributions as
the generation of words. The citation-influence model of Dietz, Bickel, and Schefer [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], also
assuming citations to be fully observed, splits the process of generating words in two branches:
a word in document d is either drawn from the topic distribution of a cited text (which is in turn
sampled from a document-specific multinomial distribution over citable documents) or from
a word distribution specific to d (“innovation”). Nallapati et al. [
        <xref ref-type="bibr" rid="ref28">27</xref>
        ] present two models that
treat citations as latent variables sampled on the basis of document-specific topic distributions.
Although not directly concerned with citations, the author-topic model of Rosen-Zvi et al. [35]
ofers an alternative view of what we want to achieve in this paper as some texts in ancient
standardized languages can indeed be considered the work of a collective of – not necessarily
contemporaneous – authors (see e.g. Colledge [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] on the composition of the Legenda aurea by
an anonymous group of authors).
      </p>
      <p>
        Previous research has proposed various admixture models that contain a temporal
component modelled either in discrete bins (e.g. Blei and Laferty [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or Frermann and Lapata
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] with Gaussian priors on logistic topic-word mixtures) or as continuous observed variables
(e.g. Wang and McCallum [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ]; Wang, Blei, and Heckerman [
        <xref ref-type="bibr" rid="ref44">44</xref>
        ]). More complex models as
e.g. proposed by Kawamae [
        <xref ref-type="bibr" rid="ref19">18</xref>
        ] split the generation of words in time- and document-specific
branches.
      </p>
      <p>
        Using bigrams in admixture models was first proposed by Wallach [
        <xref ref-type="bibr" rid="ref43">43</xref>
        ] (also see Nokel and
Loukachevitch [
        <xref ref-type="bibr" rid="ref30">29</xref>
        ] Nokel and Loukachevitch [
        <xref ref-type="bibr" rid="ref30">29</xref>
        ] for a survey). While Wallach models all data
points as bigrams, the collocation model of Griffiths, Steyvers, and Tenenbaum [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] makes
the decision for uni- vs. bigrams part of the model structure. Wang, McCallum, and Wei [
        <xref ref-type="bibr" rid="ref46">46</xref>
        ]
further make the decision for uni- vs. bigrams dependent from the hidden topic.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <p>The experiments described in this paper are based on the works of 166 Latin authors who were
active between the 3rd c. BCE and the 14th c. CE, the French philosopher Nicole Oresme
(1320-1382) being the latest one included. From among the available Latin corpora (for an
overview see McGillivray [24, ch. 2]), we chose the Latin library corpus of the CLTK library3
due to its wide coverage. An author is included if at least 50k of text are contained in the
CLTK library or if the author is considered important for (text-)historical reasons (e.g. the Res
gestae of Augustus). The raw source data are unbalanced (authors such as Cicero or Thomas
Aquinas are strongly over-represented), and individual works are often split into multiple files.
We therefore merge all works of one author into a single text, although, arguably, the preference
for citing and reusing text can vary inside the oeuvre of an author.</p>
      <p>
        Latin is a strongly inflectional language. In addition, the orthography of some source texts
has not been standardized, and especially the late Christian authors are responsible for some
variation so that working with raw textual data would result in very sparse feature matrices.
All texts are therefore lemmatized using Collatinus [
        <xref ref-type="bibr" rid="ref33">31</xref>
        ] (which manages to resolve many of
the non-standard spellings in the process) and these lemmatized versions constitute the data
used for all following steps of the processing pipeline. After removing 104 stop words such as
ad ‘to(wards)’, et ‘and’ or meus ‘my’ as well as lemmata that occur less than 30 times, the
corpus consists of 5,320,406 word tokens with 10,309 distinct lemmata (also see the summary
in Tab. 1). Public sources such as the Encyclopedia Britannica and Wikipedia are used for
gathering information about the lifetime of each author (ld, ud: birth and death years of author
d). If not specified otherwise, the date md of a text d denotes the mean of this time span, i.e.
md = 12 (ld + ud).
      </p>
      <p>3thelatinlibrary.com, http://cltk.org/</p>
    </sec>
    <sec id="sec-4">
      <title>4. Model</title>
      <p>The model discussed in this paper needs to deal with three types of uncertainty: (1) unknown
structures of word reuse; (2) fuzzy or unknown dates of composition; (3) the question whether
uni- or bigrams of words should be used as the observed features. This leads to the following
generative story (see eq. 2 for the complete specification): First, for the ith word in text d, the
source text cdi is drawn from a text-specific multinomial distribution ξd. Note that ξd includes
the text d itself. Such self-loops mean that the respective data point is peculiar to the actual
author of text d.4 While many citation models proposed so far can build on a given citation
structure (as e.g. defined by web links or scholarly citations in articles), this information is not
available for our data. The value of the prior αij (text i cites from text j) therefore needs to
be adapted during inference depending on the inferred latent times. After each iteration of the
Gibbs sampler (this means after running it once over all data points), the mean time slots µ
of all texts are calculated based on the current state of the latent temporal assignments, and
the value of αij is updated using a sigmoid function:</p>
      <p>
αij = </p>
      <p>1
 1+exp(−(μi−μj))
10 if i = j
0 if µ j − µ i &gt; 3
else
(1)
The high value for αii encourages the model to explain the words observed in a text by the
preferences of its author. Note that the zeros for the case µ j −µ i &gt; 3 are structural zeros so that
text j is not considered a possible source of i if αij = 0.5 In addition, we multiply each element
of α with a citation mask m ∈ {0, 1}D×D that is derived from running a Levenshtein-based
citation detector over the unlemmatized texts. The value mij is set to 1 if at least one sequence
of five or more words is shared by texts i, j; else to zero. Zero values in m are again interpreted
as structural zeros. The use of this mask is based on the idea that literal citations, as detected
by the Levenshtein algorithm, indicate the acquaintance of an author with a previous work
and thus increase the probability that individual words from this previous work are used as
well.</p>
      <p>Second, a time slot tdi is drawn from a text-specific multinomial temporal distribution ωcdi .
The prior βcdi of ωcdi incorporates the current state of scholarly knowledge about the time of
composition of text cdi, and possible time slots obtain a flat uniform prior in the range lcdi , ucdi
while slots outside [lcdi , ucdi ] are set to structural zeros.</p>
      <p>
        Third, the model draws a Bernoulli-distributed variable bdi that decides if the word xdi
and its successor xdi+1 typically form a bigram. Contrary to the model proposed by Wang,
McCallum, and Wei [
        <xref ref-type="bibr" rid="ref46">46</xref>
        ], this decision does not depend on the sampled time tdi and thus saves
(T − 1) · V 2 trainable parameters. Based on the sampled value of bdi, either the unigram xdi or
the bigram xdixd i+1 is drawn from time-specific multinomial distributions ϕ tUdi resp. ϕ tBdi xdi .
      </p>
      <p>
        With Θ denoting all trainable parameters and π all priors, the joint distribution is given by
4This choice is represented by the Beta distributed variable λ in Dietz, Bickel, and Schefer [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
5The diference of three time slots is motivated by the following idea: As will be shown in Sec. 5.1, 150 is
a good choice for the number of time slots. As the whole corpus covers a temporal range of about 1,700 years,
three time slots correspond to slightly more than 30 years, a span that may describe the active period of one
author.
      </p>
      <p>Variable
text → citation
citation → time
time → unigrams
time → bigrams
2 words → uni-/bigr.
+ (1 − bdi)(Cat(xdi|ϕ tUdi )))]]
(2)</p>
      <p>
        The blocked Rao-Blackwellized Gibbs Sampler [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is obtained by using Dirichlet-multinomial
integrals:
p(cdi = e, tdi = k, bdi = l, xdi = u, xd i+1 = v|c−di, t−di, x−di, Θ, π)
∝ (Ad−edi + αde)
      </p>
      <p>Be−kdi + βek  (L1xd(i−xddii)+1 + δ1) CkBu(v−di)+γvB
∑lT Be−ldi + βel  (L0xd(i−xndi)+1 + δ0) ∑Vw CkUw−di+γwU
∑VwCCkUukB−u(dw−i+diγ)u+UγwB
bdi = 1
bdi = 0</p>
      <p>A small, but important diference to models that operate with a known citation structure is
the selection of possible sources. In this paper, a text c is only considered as a possible source
for an observed uni- xdi or bigram xdixd i+1 if it also contains xdi or xdixd i+1. This condition
prevents the model from assigning too much weight to early authors such as Cicero and Vergil.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <p>This section reports qualitative and quantitative evaluations for the three relevant elements
of our model: the detected structure of word reuse (Sec. 5.2), the temporal predictions (Sec.
5.3) and the diachronic trajectories of words that can be inferred from it (Sec. 5.4).</p>
      <sec id="sec-5-1">
        <title>5.1. Architecture and Parameter Settings</title>
        <p>
          We use posterior predictive checks (PPC; Mimno, Blei, and Engelhardt [
          <xref ref-type="bibr" rid="ref27">26</xref>
          ]) to compare various
model architectures and parameter settings. Given a trained model, we draw textwise samples
of the observed words using Eq. 2 and compare these samples with the true distributions in
(a) No.
        </p>
        <p>0.5912</p>
        <p>of slots, η2 = (b) Prior γ, η2 = 0.4899 (c) Prior δ, η2 = 0.01540 (d) Cit. mask, η2 = 0.0004
each text using the Hellinger Distance. The values that result from 30 replications per text
are grouped by texts and z-standardized, and ANOVAs are performed in order to test for
significant diferences between settings. Figure 1 shows smoothed density estimates of these
z-scores for four central design choices: the number of temporal slots (Fig. 1a), the parameters
γ (time → feature; Fig. 1b) and δ (uni- or bigram; Fig. 1c) and the use of the precomputed
citation mask (Fig. 1d). While ANOVA points to (highly) significant diferences in all four
settings, the values of Cohen’s η2 which quantify the efect size and are displayed below each
subfigure indicate that only the number of slots and the prior γ have a relevant influence on
the outcome of the model, while the influence of δ and the citation mask must be considered
as very small. Based on this evaluation, we choose 150 time slots, γ = 0.5, δ = 0.01 for all
following experiments, and we apply the citation mask.</p>
        <p>Running another PPC for establishing the optimal number of iterations of the Gibbs sampler,
we found no significant diferences between models trained with 100, 300, 500 or 1,000 iterations
(p-value of the ANOVA: 0.163). This somehow unexpected result is certainly due to the fact
that our model already has rather strong priors induced by the structural zeros in the citation
mask and the temporal prior β so that only few iterations are required to obtain a good
representation of the data. We therefore run the sampler for 100 iterations and record the
sampled values once after the last iteration.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Word Reuse</title>
        <p>
          As mentioned in the introduction, understanding the intellectual lineages of historical corpora
is one important aim of this paper. Therefore, the evaluation starts with inspecting the inferred
structure of word reuse. We calculate, for each text d, the proportion of words labeled as reused,
i.e. for which cdi ̸= d according to the model output. These proportions can be expected to be
correlated with the true date of d, as later texts have more opportunities to reuse words than
earlier ones. In order to deal with this efect, we perform a partial correlation by fitting a linear
regression that predicts the proportions of words labeled as reused (y) based on the number of
possible source texts (x). The residuals of this regression, which capture how much the model
output deviates from the linear estimate, are plotted against the true date of each text (see
Fig. 2a). Here, the dashed horizontal line at y = 0 corresponds to a residual of 0 and thus to
a perfect prediction of the model output through the linear regression. The blue curved line is
a smoothed density estimate of the actual residuals. This smoothed estimate shows that the
proportions of word reuse conform to the values estimated by the linear model until the end of
(a) Residuals of a linear regression that predicts the
inferred number of reused words given the
number of available source authors. Individual
authors are labeled if their residuals fall in the 5%
resp. 95% quantiles.
(b) Schematical representation of word reuse,
grouped by literary periods. The source periods
are found at the bottom. The line width indicates
the strength of the activity.
the Late Antiquity (5th c. CE). We observe increasing word reuse in the 8th or 9th c. CE, a
period commonly known as the Carolingian Renaissance, which saw a revival of classical Latin
literature that accompanied the formation of the Carolingian state [see e.g. 39]. In the 10th c.
CE and later, the proportions of word reuse tend to fall below their expected values. Seen from
the perspective of literary history, the intensive word reuse in the Carolingian Renaissance is
connected with authors such as Hrabanus Maurus, Angilbert or Alcuin, who in his De rhetorica
freely mixes extracts from Cicero’s De inventione and other authoritative sources with his own
comments [see e.g. 19]. In the 10th c., a new form of Latin is constituted, which, though still
accepting the classical language as its gold standard, is strongly influenced by the idiom of
Christian theological authors (“Ecclesiastical Latin”, see e.g. Dinkova-Bruun [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]). This form
of medieval Latin can therefore be expected to share less lexical features with Classical Latin
than earlier forms of the language.
        </p>
        <p>
          Figure 2b presents another view of the literary influences. In this plot all texts are aggregated
by the five literary periods defined by Adamik [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], plus an extra period “Medieval Latin”
starting at 900 CE.6 The widths of the lines between target, i.e. “citing” (top), and source, i.e.
“cited” (bottom), periods indicate the relative amount of word reuse inferred by the model.
The plot shows that works from the classical era quite constantly remained important sources
of word reuse throughout all periods considered in this paper, although even their influence
begins to wane in the Transitional Period (600–900 CE) and the Middle Ages. Such a result
makes sense as the works of some classical authors did not survive the breaks in the political
and religious history and were only rediscovered in the Italian Renaissance or even later (see
e.g. Tutrone [
          <xref ref-type="bibr" rid="ref40">40</xref>
          ] on the limited reception of the important Roman philosopher Lucretius in
the (early) Middle Ages). The strong connections between Late Latin on one hand and the
transitional and medieval periods on the other are due to the numerous important Christian
texts composed in Late Antiquity, most notably the Latin translation of the Bible (Vulgata)
6We label the period called Vulgar Latin by Adamik as Late Latin in this paper in order to distinguish it
from the sub-standard variety discussed by Herman [
          <xref ref-type="bibr" rid="ref16">15</xref>
          ].
and the work of Augustine. In addition, Fig. 2b shows a decline in word reuse between the
Transitional Period, which comprises the Carolingian Renaissance just discussed, and Medieval
Latin – most authors from the Transitional Period were obviously not too much regarded in
later times.
        </p>
        <p>In order to understand which authors are mainly responsible for the distribution observed
in Fig. 2b, we collect, for each literary period, those three authors with the highest amount
of words marked as reused, applying a minimal threshold of 1,000. The resulting list contains
the following authors:
Old Cato (the Elder) is the only representative of old Roman literature, a result which is
in accordance with his extraordinary importance for the development of a genuinely
Latin literature. His compendia on agriculture and warfare as well as the collection of
his orations (compiled by himself) exerted a considerable influence on later authors [ 2,
pp. 340–41].</p>
        <p>Classical Ovid and Cicero can be seen as the top representatives of Latin poetry and prose,
while Livy stands for the genre of classical historiography.</p>
        <p>Late This period shows an interesting interference between the famous Christian author
Augustine and the Vulgata, a new translation of the Bible composed by Jerome. Diferent
from what may be expected, Augustine is more frequently marked as cited than the
Vulgata (161,501 vs. 74,445 times). A closer inspection of words and bigrams labeled
as cited reveals that the model has problems in assigning individual Biblical citations to
the Vulgata or the Vetus Latina, the older Latin version of the Bible preferably cited by
Augustine [see e.g. 16, pp. 36–39]. – The third representative of this period is Gregory
of Tours, best known for his historical writings.</p>
        <p>
          Transitional Here, only Beda has made it in the list – a result fully in accordance with his
popularity in the Middle Ages [
          <xref ref-type="bibr" rid="ref47">47</xref>
          ].
        </p>
        <p>Medieval While Thomas Aquinas is a central representative of medieval Latin and its focus
on theological discussions, Albert of Aix and William of Tyre represent the genre of
medieval historical writings with a special focus on the Crusades.</p>
        <p>To sum up this section, it appears that the model was able to recover structures of word
reuse that conform to scholarly expectations.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Timestamping</title>
        <p>We model the partly unclear times of composition as latent variables. In this section we assess
the quality of the resulting temporal predictions. We simulate a research setting in which
only approximate temporal information is available, by setting the temporal ranges of all D
texts d to the ranges of the literary periods containing them according to Adamik [1, p. 9].
These artificially obfuscated ranges are used as temporal priors βd (see eq. 2). All texts
are trained jointly, and we evaluate how well the model can recover the exact dates and the
correct temporal order of the texts. Notably, this experiment is not merely another academic
exercise, but bears practical implications when studying ancient Indian text corpora for which
only approximate temporal information is available [see 14]. – Table 3 reports two evaluation
measures:
1
• The period-wise mean absolute error (MAE) calculated as |{d∈P }| ∑d∈P ||md−µ d||1 where
µ d is the mean of the word-wise temporal assignments for text d, and P is the literary
period.
• Ranking accuracy: The texts are grouped by their literary periods, and all texts
belonging to one period are ordered by their true dates md. The ranking accuracy gives the
proportion of text pairs for which the predicted temporal order is the same as the true
one.</p>
        <p>
          The results in Tab. 3 show that dating texts composed in standardized languages is challenging.
Although the literary periods only extend over 200-300 years each, the MAEs vary between
40 and 140 years and thus cover substantial parts of each period. It may, however, be noted
that Kumar, Lease, and Baldridge [
          <xref ref-type="bibr" rid="ref21">20</xref>
          ] report slightly higher MAEs of 85-155 years for English
stories published between 1798 and 2008, which suggests that the results achieved by our model
are actually in an acceptable range. The values of the ranking accuracy are coupled with the
uncertainties in the temporal predictions and fall below the random baseline of 50% for three
of the five periods. Notably, both evaluation measures seem to get worse for post-Classical
texts when Latin gradually ceased to be used as a spoken language, and an ANOVA of the
MAEs as well as a Fisher-Yates test of the raw counts for the ranking accuracies both show
(highly) significant diferences between all periods (p-values: 0.00147 [MAE]; 0.0005 [ranking
acc.]).
        </p>
        <p>
          In order to assess if the temporal predictions improve when more reliable temporal
information is available, we perform a cross-validation experiment. A subset of fifteen authors 7 is
chosen as the test set. For each text in this set, we obfuscate its date in the same way as in the
ifrst experiment, while all D − 1 other texts keep their temporal gold information. The model
is trained with the D − 1 training texts for 100 iterations and then for another 100 iterations
with the combined training and test set (see the method Gibbs1 in Yao, Mimno, and
McCallum [
          <xref ref-type="bibr" rid="ref49">49</xref>
          ]). The results are compared with the predictions made by the Topic over Time model
[
          <xref ref-type="bibr" rid="ref45">45</xref>
          ] which is often used as a baseline for latent variable models with a temporal component.8
The results in Tab. 4 show that our model is slightly, but not significantly better than ToT
(p-value of a paired directed Wilcoxon test: 0.26). While ToT occasionally assigns all texts
from one period to the same date range, our model better captures the temporal dynamics.
This impression is confirmed when calculating the ranking accuracy (ours: 60%; ToT: 33%)
for the data in Tab. 4.
        </p>
        <p>7This limitation is due to time constraints. We choose three authors from the start, middle and end of each
period; see the first column of Tab. 4.</p>
        <p>
          8We use 150 topics and all hyperparameter settings as described in the original paper. The predicted time
slot is that with the highest posterior argmaxt ∑ind log p(t|ψzi ), with the additional constraint that ld ≤ t ≤ ud,
in order to make a fair comparison with the model presented in this paper; see Sec. 2 in Wang and McCallum
[
          <xref ref-type="bibr" rid="ref45">45</xref>
          ].
        </p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Features</title>
        <p>Getting a more realistic picture of how words are diachronically distributed in standardized
languages is an important aim of this paper. This section therefore compares the linguistic
expressiveness of empirical corpus distributions with those inferred by our model. Using
posterior estimates of the variational parameters (i.e. ωd′t = Bdt+τdt etc.) based on those cases
∑Tu Bdu+τdu
in which the model assigns words to unigrams, we obtain the conditional probabilities p(x|d)
of a word x given a text d by marginalizing the latent citations and temporal assignments:</p>
        <p>D
p(x|c) = ∑
c</p>
        <p>T D
∑ p(c|d)p(t|c)p(x|t) = ∑
t c</p>
        <p>T
∑ ξd′cωc′tϕ U ′tx
t
(3)
We expect that the diachronic trajectories of this conditional distribution difer from the corpus
distribution of a word x when the use of x in later texts is mainly due to literary influences.</p>
        <p>
          In order to quantitatively support the claim that the inferred distributions yield a more
realistic description of the actual language use, we address the problem of predicting lexical
stability [see 37]. While substantial parts of the vocabulary of Romance languages can be
derived from precursors in (Vulgar) Latin by applying rules of regular sound change [
          <xref ref-type="bibr" rid="ref36 ref37">36, 37</xref>
          ],
there are important individual words such as equus ‘horse’ or whole classes of words such as
the vocabulary of war that do not have derivatives in the Romance languages. Apart from
various socio-cultural factors (on which see e.g. Campbell [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], 244f. and especially Vincent
[
          <xref ref-type="bibr" rid="ref41">41</xref>
          ] on Latin vocabulary that was “submerged” in classical works), the frequency of use in the
spoken language is a determining factor for the survival or obsolescence of a word [
          <xref ref-type="bibr" rid="ref34">32</xref>
          ]. If the
inferred distributions better capture the actual use than the corpus distributions, they should
better be able to predict the survival of Latin words in the Romance languages.9
9A factor we ignored as non-essential for the present purpose, but which should be taken into account in
        </p>
        <p>
          As the etymological information in Wiktionary is incomplete and noisy [
          <xref ref-type="bibr" rid="ref48">48</xref>
          ], we collect all
Latin words that are recorded as etyma of Romance words in Meyer-Lübke [
          <xref ref-type="bibr" rid="ref26">25</xref>
          ], a standard
reference work of Romance etymologies. Although scholarly research has revised some decisions
made in this work, it is still considered as a largely complete collection of surviving Latin
etyma (see e.g. Stefenelli [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ], 568) so that words not recorded there can be assumed not to have
derivatives in Romance languages. From among the 10,308 words in our vocabulary, 2691, i.e.
26.1%, have such a derivative.10
        </p>
        <p>We aggregate the empirical and inferred distributions by various ranges of years, z-standardize
these binned values and use them as input features for a feed-forward neural network with four
hidden units and softplus activations. The neural network is trained on the binary prediction
task whether or not a Latin word has derivatives in any Romance language.11</p>
        <p>Figure 3a shows the F-scores (y-axis) depending on the sizes of the temporal bins applied
(x-axis). While the F-scores generally decrease with increasing sizes of the temporal bins, the
F-score of the inferred distributions is consistently higher than that of the empirical ones. The
drop of the F-score is especially obvious when using the empirical distributions with a bin size
of 30 years instead of the unbinned distributions (“all”). The failure of the model that uses the
empirical distribution is due to its low recall in these cases. In order to better understand the
behaviour of the predictor, we collect all inherited words that were labelled correctly using the
inferred, but wrongly using the empirical distributions, calculate their empirical and inferred
distributions and smooth these distributions with a Gaussian kernel. Figure 3c contrasts the
means (plus/minus one standard deviation) of the two groups. The plot shows that the inferred
distribution transfers probability mass from occurrences in (late) classical texts to the (early)
Middle Ages (∼ 8th c.+), i.e. to a period in which the Romance languages are generally
assumed to develop. A similar efect can be observed for words which are only predicted
correctly when using the empirical distributions (see Fig. 3d). Apparently, the mixture model
has missed efects of word reuse in these cases, as it assigns too much weight to occurrences in
the early Middle Ages. Finally, when examining distributions of inherited words detected by
neither classifier, it becomes apparent that many of them are popular in classical and medieval
texts, but rare in the Late Antiquity and the Transitional Period (see e.g. the plots for expecto
‘expect’ in Fig. 3b). Although the mixture model draws up the distributions for the critical
phase of the early Middle Ages, this efect is not strong enough to make the classifier label
such words as inherited.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Summary</title>
      <p>Diachronic corpora are indispensable tools for studying linguistic developments and
intellectual lineages in premodern societies. Depending on the degree of standardization which the
corpus language has undergone as well as on the amount of text reuse, linguistic distributions
extracted from diachronic corpora can be misleading because the language usage of
authorifuture, more detailed studies, is the fact that, starting from the early Middle Ages, for a growing number of
authors their mother tongue is not a Romance language but belongs to another family (mostly the Germanic
one).</p>
      <p>10Note that this number only covers Romance words derived by regular sound change, but not, for example,
borrowed words.</p>
      <p>
        11Meyer-Lübke [
        <xref ref-type="bibr" rid="ref26">25</xref>
        ] does not consistently report all Romance derivatives of a given Latin word, so that we
could not formulate this problem as a multi-class prediction task. – Apart from a simple neural network, we
also tested flat ML models such as logistic regression, but found our approach to perform better.
(a) F-scores of predicting the lexical stability (b) Normalized and smoothed empirical and
of Latin words in Romance languages; val- inferred distributions of the word expecto
ues reported without grouping (“all”) and ‘expect’
grouped by the number of years per
temporal bin.
(c) Inferred correct, empirical wrong
(d) Inferred wrong, empirical correct
tative, frequently cited works can conflate with that of their literary successors. This paper
introduces a latent variable model that captures such literary influences while simultaneously
accounting for uncertainties in the temporal assignments. While the latter aspect is only of
limited importance for Latin, the corpus language discussed in this paper, it is certainly
relevant for many ancient corpora whose temporal structure is more disputed. Our discussion has
shown that the model retrieves meaningful intellectual lineages and structures of word reuse
(see Sec. 5.2) and performs on par with latent variable models specifically designed for
capturing temporal topical trends (Sec. 5.3). In addition, the discussion of etymological derivations
in Sec. 5.4 indicates that the linguistic distributions generated by the model are better able
to describe certain aspects of language development than plain corpus distributions. Future
extensions should incorporate a component that smoothes the temporal distributions [see e.g.
11], and they should consider non-temporal influence factors such as the geographic origin or
genre of a text, as was proposed by Perrone et al. [
        <xref ref-type="bibr" rid="ref35">33</xref>
        ]. Given this outcome, we are planning
to apply the mixture model on text traditions of ancient South Asia whose intellectual and
diachronic structures are still not fully understood.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We thank Sabine Tittel for her help with digital resources for Romance languages and the
three anonymous reviewers for their insightful comments. The authors were partly funded by
the German Federal Ministry of Education and Research, FKZ 01UG2121.
[19]</p>
      <p>M. Rosen-Zvi, T. Griffiths, M. Steyvers, and P. Smyth. “The Author-topic Model for
Authors and Documents”. In: Proceedings of the 20th Conference on Uncertainty in AI.
2004, pp. 487–494.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Adamik</surname>
          </string-name>
          . “
          <article-title>The Periodization of Latin. An Old Question Revisited”</article-title>
          .
          <source>In: Latin Linguistics in the Early 21st Century</source>
          . Ed. by
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Haverling</surname>
          </string-name>
          .
          <source>Uppsala Universitet</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>640</fpage>
          -
          <lpage>652</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. von</given-names>
            <surname>Albrecht</surname>
          </string-name>
          .
          <source>Geschichte der römischen Literatur</source>
          . Vol.
          <volume>1</volume>
          . Berlin: de Gruyter,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Laferty</surname>
          </string-name>
          . “
          <article-title>Dynamic Topic Models”</article-title>
          .
          <source>In: Proceedings of the 23rd International Conference on Machine Learning</source>
          .
          <year>2006</year>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Campbell</surname>
          </string-name>
          . Historical Linguistics. Edinburgh: Edinburgh University Press,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Clackson</surname>
          </string-name>
          . “
          <article-title>Classical Latin”</article-title>
          . In:
          <article-title>A Companion to the Latin Language</article-title>
          . Ed. by
          <string-name>
            <given-names>J.</given-names>
            <surname>Clackson</surname>
          </string-name>
          . Maiden, MA: Blackwell Publishing,
          <year>2011</year>
          , pp.
          <fpage>236</fpage>
          -
          <lpage>256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Cohn</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          . “
          <article-title>The Missing Link - A Probabilistic Model of Document Content and Hypertext Connectivity”</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          .
          <year>2001</year>
          , pp.
          <fpage>430</fpage>
          -
          <lpage>436</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>E. Colledge.</surname>
          </string-name>
          “
          <article-title>James of Voragine's “Legenda Sancti Augustini” and its Sources”</article-title>
          .
          <source>In: Augustiniana 35.3/4</source>
          (
          <year>1985</year>
          ), pp.
          <fpage>281</fpage>
          -
          <lpage>314</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Dietz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bickel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Schefer</surname>
          </string-name>
          . “
          <article-title>Unsupervised Prediction of Citation Influences”</article-title>
          .
          <source>In: Proceedings of the 24th ICML</source>
          .
          <year>2007</year>
          , pp.
          <fpage>233</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Dinkova-Bruun</surname>
          </string-name>
          .
          <article-title>“Medieval Latin”</article-title>
          . In:
          <article-title>A Companion to the Latin Language</article-title>
          . Ed. by
          <string-name>
            <given-names>J.</given-names>
            <surname>Clackson</surname>
          </string-name>
          . Maiden, MA: Blackwell Publishing,
          <year>2011</year>
          , pp.
          <fpage>284</fpage>
          -
          <lpage>302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Erosheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fienberg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Laferty</surname>
          </string-name>
          . “
          <article-title>Mixed-membership Models of Scientific Publications”</article-title>
          .
          <source>In: Proceedings of the National Academy of Sciences 101.suppl 1</source>
          (
          <year>2004</year>
          ), pp.
          <fpage>5220</fpage>
          -
          <lpage>5227</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Frermann</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Lapata</surname>
          </string-name>
          .
          <article-title>“A Bayesian Model of Diachronic Meaning Change”</article-title>
          .
          <source>In: Transactions of the Association for Computational Linguistics</source>
          <volume>4</volume>
          (
          <year>2016</year>
          ), pp.
          <fpage>31</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Griffiths</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Steyvers</surname>
          </string-name>
          . “
          <article-title>Finding Scientific Topics”</article-title>
          .
          <source>In: tional Academy of Sciences 101.Suppl</source>
          .
          <volume>1</volume>
          (
          <issue>2004</issue>
          ), pp.
          <fpage>5228</fpage>
          -
          <lpage>5235</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Griffiths</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Steyvers</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Tenenbaum</surname>
          </string-name>
          . “Topics in Semantic Representation.”
          <source>In: Psychological Review 114.2</source>
          (
          <issue>2007</issue>
          ), pp.
          <fpage>211</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>O.</given-names>
            <surname>Hellwig</surname>
          </string-name>
          . “
          <article-title>Dating and Stratifying a Historical Corpus with a Bayesian Mixture Model”</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <source>In: Proceedings of LT4HALA</source>
          .
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Herman</surname>
          </string-name>
          . Vulgar Latin. University Park, Pennsylvania: Pennsylvania State University Press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Houghton</surname>
          </string-name>
          .
          <source>The Latin New Testament</source>
          . Oxford: Oxford University Press,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Joseph</surname>
          </string-name>
          .
          <source>Eloquence and Power: The Rise of Language Standards and Standard Languages. London: Frances Pinter</source>
          ,
          <year>1987</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kawamae</surname>
          </string-name>
          . “
          <article-title>Trend Analysis Model: Trend Consists of Temporal Words, Topics, and Timestamps”</article-title>
          .
          <source>In: Proceedings of the fourth ACM International Conference on Web Search and Data Mining</source>
          .
          <year>2011</year>
          , pp.
          <fpage>317</fpage>
          -
          <lpage>326</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Kempshall</surname>
          </string-name>
          . “
          <article-title>The virtues of rhetoric: Alcuin's “Disputatio de</article-title>
          rhetorica et de uirtutibus””.
          <source>In: Anglo-Saxon England</source>
          <volume>37</volume>
          (
          <year>2008</year>
          ), pp.
          <fpage>7</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lease</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Baldridge</surname>
          </string-name>
          . “
          <article-title>Supervised Language Modeling for Temporal Resolution of Texts”</article-title>
          .
          <source>In: Proceedings of the 20th ACM CIKM</source>
          .
          <year>2011</year>
          , pp.
          <fpage>2069</fpage>
          -
          <lpage>2072</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>B.</given-names>
            <surname>Laurioux</surname>
          </string-name>
          . “
          <article-title>Cuisiner à l'antique: Apicius au Moyen Âge”</article-title>
          .
          <source>In: Médiévales</source>
          <volume>26</volume>
          (
          <year>1994</year>
          ), pp.
          <fpage>17</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>“A Computational Model of Text Reuse in Ancient Literary Texts”</article-title>
          .
          <source>In: Proceedings of the 45th ACL</source>
          .
          <year>2007</year>
          , pp.
          <fpage>472</fpage>
          -
          <lpage>479</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Karsdorp</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          .
          <article-title>“A Statistical Foray into Contextual Aspects of Intertextuality”</article-title>
          .
          <source>In: Proceedings of the Workshop on Computational Humanities Research (CHR</source>
          <year>2020</year>
          ). Ed. by
          <string-name>
            <given-names>F.</given-names>
            <surname>Karsdorp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McGillivray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nerghes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Wevers</surname>
          </string-name>
          .
          <year>2020</year>
          , pp.
          <fpage>77</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>B.</given-names>
            <surname>McGillivray</surname>
          </string-name>
          .
          <article-title>Methods in Latin Computational Linguistics</article-title>
          . Vol.
          <volume>1</volume>
          .
          <string-name>
            <surname>Brill</surname>
          </string-name>
          <article-title>'s Studies in Historical Linguistics</article-title>
          .
          <source>Leiden: Brill</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [25] W. Meyer-Lübke.
          <article-title>Romanisches etymologisches Wörterbuch</article-title>
          . Heidelberg: Winter,
          <year>1935</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>D.</given-names>
            <surname>Mimno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          , and
          <string-name>
            <surname>B. E. Engelhardt.</surname>
          </string-name>
          “
          <article-title>Posterior Predictive Checks to Quantify Lack-of-fit in Admixture Models of Latent Population Structure”</article-title>
          .
          <source>In: Proceedings of the National Academy of Sciences 112.26</source>
          (
          <year>2015</year>
          ),
          <fpage>E3441</fpage>
          -
          <lpage>e3450</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [27]
          <string-name>
            <surname>R. M. Nallapati</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>E. P.</given-names>
          </string-name>
          <string-name>
            <surname>Xing</surname>
            , and
            <given-names>W. W.</given-names>
          </string-name>
          <string-name>
            <surname>Cohen</surname>
          </string-name>
          . “
          <article-title>Joint Latent Topic Models for Text and Citations”</article-title>
          .
          <source>In: Proceedings of the 14th ACM SIGKDD</source>
          .
          <year>2008</year>
          , pp.
          <fpage>542</fpage>
          -
          <lpage>550</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nehrdich</surname>
          </string-name>
          .
          <article-title>“A Method for the Calculation of Parallel Passages for Buddhist Chinese Sources Based on Million-scale Nearest Neighbor Search”</article-title>
          .
          <source>In: Journal of the Japanese Association for Digital Humanities 5.2</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>132</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nokel</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Loukachevitch</surname>
          </string-name>
          . “
          <article-title>Accounting N-grams and Multi-word Terms can Improve Topic Models”</article-title>
          .
          <source>In: Proceedings of the 12th Workshop on Multiword Expressions.</source>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <year>2016</year>
          , pp.
          <fpage>44</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>P.</given-names>
            <surname>Olivelle</surname>
          </string-name>
          .
          <source>The Early Upaniṣads. Annotated Text and Translation</source>
          . Oxford: Oxford University Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ouvrard</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Verkerk</surname>
          </string-name>
          . “Collatinus &amp;
          <article-title>Eulexis: Latin &amp; Greek Dictionaries in the Digital Ages”</article-title>
          . In: Digital Classics III:
          <article-title>Re-thinking Text Analysis</article-title>
          . Center for Hellenic Studies/Harvard University,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pagel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. D.</given-names>
            <surname>Atkinson</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Meade</surname>
          </string-name>
          . “
          <article-title>Frequency of Word-use Predicts Rates of Lexical Evolution throughout Indo-European history”</article-title>
          .
          <source>In: Nature</source>
          <volume>449</volume>
          .7163 (
          <year>2007</year>
          ), pp.
          <fpage>717</fpage>
          -
          <lpage>720</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>V.</given-names>
            <surname>Perrone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hengchen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vatri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Q.</given-names>
            <surname>Smith</surname>
          </string-name>
          , and
          <string-name>
            <surname>B. McGillivray.</surname>
          </string-name>
          “GASC:
          <article-title>Genre-Aware Semantic Change for Ancient Greek”</article-title>
          .
          <source>In: Proceedings of the 1st International Workshop on Computational Approaches to Historical Language Change</source>
          .
          <year>2019</year>
          , pp.
          <fpage>56</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>J. B. Solodow. Latin</given-names>
            <surname>Alive</surname>
          </string-name>
          .
          <article-title>The Survival of Latin in English and the Romance Languages</article-title>
          . Cambridge: Cambridge University Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>A.</given-names>
            <surname>Stefenelli</surname>
          </string-name>
          . “
          <article-title>Lexical Stability”</article-title>
          .
          <source>In: The Cambridge History of the Romance Languages. Volume I:</source>
          Structures. Ed. by
          <string-name>
            <given-names>M.</given-names>
            <surname>Maiden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Smith</surname>
          </string-name>
          , and
          <string-name>
            <given-names>A.</given-names>
            <surname>Ledgeway</surname>
          </string-name>
          . Cambridge: Cambridge University Press,
          <year>2011</year>
          , pp.
          <fpage>564</fpage>
          -
          <lpage>584</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Stone</surname>
          </string-name>
          . “
          <article-title>What Plagiarism was not: Some Preliminary Observations on Classical Chinese Attitudes Toward What the West Calls Intellectual Property”</article-title>
          .
          <source>In: Marquette Law Review</source>
          <volume>92</volume>
          (
          <year>2008</year>
          ), p.
          <fpage>199</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>G. Trompf.</surname>
          </string-name>
          “
          <article-title>The Concept of the Carolingian Renaissance”</article-title>
          .
          <source>In: Journal of the History of Ideas 34.1</source>
          (
          <issue>1973</issue>
          ), pp.
          <fpage>3</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>F.</given-names>
            <surname>Tutrone</surname>
          </string-name>
          . “
          <string-name>
            <surname>Lucretius</surname>
          </string-name>
          Franco-Hibernicus:
          <article-title>Dicuil's Liber de Astronomia and the Carolingian Reception of De Rerum Natura”</article-title>
          .
          <source>In: Illinois Classical Studies 45.1</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>224</fpage>
          -
          <lpage>252</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>N.</given-names>
            <surname>Vincent</surname>
          </string-name>
          . “
          <article-title>Continuity and Change from Latin to Romance”</article-title>
          . In: Early and
          <string-name>
            <given-names>Late</given-names>
            <surname>Latin</surname>
          </string-name>
          . Continuity or Change? Ed. by
          <string-name>
            <given-names>J.</given-names>
            <surname>Adams</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Vincent</surname>
          </string-name>
          . Cambridge: Cambridge University Press,
          <year>2016</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>J. Wackernagel. Altindische Grammatik. I.</given-names>
            <surname>Lautlehre</surname>
          </string-name>
          . Göttingen: Vandenhoek und Ruprecht,
          <year>1896</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>H. M.</given-names>
            <surname>Wallach</surname>
          </string-name>
          . “Topic Modeling:
          <article-title>Beyond Bag-of-words”</article-title>
          .
          <source>In: Proceedings of the 23rd ICML</source>
          .
          <year>2006</year>
          , pp.
          <fpage>977</fpage>
          -
          <lpage>984</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Blei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Heckerman</surname>
          </string-name>
          . “
          <article-title>Continuous Time Dynamic Topic Models”</article-title>
          .
          <source>In: Proceedings of the Twenty-Fourth Conference on Uncertainty in Artificial Intelligence</source>
          .
          <year>2008</year>
          , pp.
          <fpage>579</fpage>
          -
          <lpage>586</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <surname>A. McCallum.</surname>
          </string-name>
          “
          <article-title>Topics over Time: A Non-Markov Continuous-time Model of Topical Trends”</article-title>
          .
          <source>In: Proceedings of the 12th ACM SIGKDD International conference on Knowledge Discovery and Data Mining</source>
          .
          <year>2006</year>
          , pp.
          <fpage>424</fpage>
          -
          <lpage>433</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and X.</given-names>
            <surname>Wei</surname>
          </string-name>
          . “
          <string-name>
            <surname>Topical</surname>
          </string-name>
          N-grams:
          <article-title>Phrase and Topic Discovery, with an Application to Information Retrieval”</article-title>
          .
          <source>In: Proceedings of the Seventh ICDM</source>
          .
          <year>2007</year>
          , pp.
          <fpage>697</fpage>
          -
          <lpage>702</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>D.</given-names>
            <surname>Whitelock</surname>
          </string-name>
          . After Bede.
          <source>Newcastle: Bealls</source>
          ,
          <year>1978</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Yarowsky</surname>
          </string-name>
          . “
          <article-title>Computational Etymology and Word Emergence”</article-title>
          .
          <source>In: Proceedings of The 12th Language Resources and Evaluation Conference</source>
          .
          <year>2020</year>
          , pp.
          <fpage>3252</fpage>
          -
          <lpage>3259</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mimno</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. McCallum.</surname>
          </string-name>
          “
          <article-title>Efficient Methods for Topic Model Inference on Streaming Document Collections”</article-title>
          .
          <source>In: Proceedings of the 15th ACM SIGKDD</source>
          .
          <year>2009</year>
          , pp.
          <fpage>937</fpage>
          -
          <lpage>946</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>