<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Antwerp, Belgium
£ marijn.koolen@gmail.com(M. Koolen);rik.hoekstra@di.huc.knaw.n(lR. Hoekstra)
ç https://marijnkoolen.com/(M. Koolen)
ȉ</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Detecting Formulaic Language Use in Historical Administrative Corpora</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>MarijnKoolen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rik Hoekstra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DHLab, KNAW Humanities Cluster</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>KNAW Huygens Institute</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Historical administrative corpora are 昀椀lled with jargon and formulaic expressions that were used consistently across many documents. Governmental decisions, notarial deeds and o昀케cial charters o昀琀en contain 昀椀xed expressions to ensure that the same legal aspects in di昀erent documents had the same interpretation. Such formulaic expressions can be used to identify speci昀椀c elements of a document. For instance, a deed has di昀erent formulas to indicate whether it concerns the sale of property or the transferal of rights. In this paper we explore formulas as a methodological devise to structure the text of an administrative corpus and make the information contained in it better accessible. We use a datadriven method to detect potential formulaic expressions in historical corpora, that can deal with spelling variation and change and recognition errors introduced in the digitisation process. We apply this exploratory technique on a corpus of almost 300,000 eighteenth-century resolutions of the States General of the Dutch Republic and 昀椀nd many formulaic expressions that capture relationships between the political actors involved and the decisions that were made. A 昀椀rst analysis suggests that many formulas can be used to add metadata to individual resolutions on various elements of the proposals and decisions that are part of each resolution.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;formulaic expressions</kwd>
        <kwd>text reuse</kwd>
        <kwd>document structure</kwd>
        <kwd>information extraction</kwd>
        <kwd>text analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The Resolutions of the States General of the Dutch Republic (1576-1796) is a digitised archive
containing an estimated 1 million decisions made by the States General (SG) during their daily
meetings. It is 昀椀lled with administrative jargon and formulaic expressions that were used
consistently, tens of thousands of times across a 220 year period, in resolutions with a very 昀椀xed
structure. These formulaic expressions were used to signal speci昀椀c elements in the text, so that
anyone relying on the resolutions for their day-to-day work could easily 昀椀nd back requests,
decisions and agreements by looking for these 昀椀xed phrasings, which also made sure that similar
decisions and agreements had similar interpretations.</p>
      <p>In this paper we explore formulas as a methodological device to structure the text of the
archive and make the information contained in it better accessible. The secretaries of the
meetings used a 昀椀xed structure and 昀椀xed expressions to signal the opening of a new resolution,
that always started with a proposal submitted to the SG. Each resolution ends with a decision
paragraph, which also starts with a formulaic expression, followed by the details of what was
agreed upon and what should happen next. For instance, to signal the SG reached an
agreement on what should be done in response to a proposition, they used the formula ‘Waer op
gedelibereert zijnde, is goetgevonden ende verstaen, ...’ (ENO:n which has been agreed and
understood .... A number of examples of this phrase are shown in Figur1e). This phrase recurs
tens of thousands of times in the resolutions and signals that an agreement and decision were
reached that are detailed in the following paragraph.</p>
      <p>The formulas thus not only help us structure the material, but also to add metadata to the
individual resolutions and make them better accessible for analysis.</p>
      <p>Our experience in working with other historical collections prompted a set of questions: Are
such formulaic expressions used in other administrative corpora? And what other domains and
document genres contain formulaic expressions? From an information access perspective, is
the textual repetition of an administrative corpus like the Resolutions di昀erent from textual
repetition in corpora in other domains?</p>
      <p>To get a better idea of the relevance of formulas, we 昀椀rst compare repetitive text
characteristics in corpora of di昀erent domains, to establish whether administrative texts like the
resolutions are of a di昀erent quality from other types of text. The second topic of this paper is
the methodology we use for algorithmically identifying potential formulas in the resolutions,
and a discussion of their use. In this paper we con昀椀ne ourselves to these points, but we realise
that formulas have relevance for a number of wider humanities research questions. We discuss
these at the end in Section6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>2.1. Formulaic Language</title>
        <p>Our work touches on three strands of research: 1) formulaic language use, 2) text reuse
detection and 3) dealing with variation-rich text.</p>
        <p>
          The use of formulaic expressions is mostly studied in the 昀椀elds of linguistic4s7,[26] and
language learning 1[
          <xref ref-type="bibr" rid="ref13 ref5">4, 7, 42, 13</xref>
          ]. Formulaic language is typically de昀椀ned as 昀椀xed word
combinations, with o昀琀en non-literal meanings, that are used to improve 昀氀uency and reduce
misunderstanding [47, 41, 46]. Poß and Wouden studied formulaic expressions consisting of anything
more than one word asExtended Lexical Units, stored as single a entry in a speaker’s mental
lexicon.
        </p>
        <p>We found two studies that investigated the use of formulas in text corpora. Karsdorp
identi昀椀ed and classi昀椀ed formulaic opening and closing expressions of Dutch folk tales. Repetitive
patterns in the 昀椀rst and last 昀椀ve words of folk tales are detected and are found to be predictive
of the genre of a folk tale. That is, the opening formula o昀琀en signals that something is a joke,
saga or fairy tale. In the resolutions, the opening formulas are similarly indicative of what kind
of proposal (petition, report, declaration, etc.) is discussed in the resoluti3o7n].m[ anually
annotated the use of formulaic expressions in seventeenth and eighteenth century Dutch letters
and found that more experienced writers tend to use more 昀椀xed expressions. They suggest
that this indicates that formulas are partly used to reduce cognitive e昀ort.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Text Reuse Detection</title>
        <p>
          Text reuse detection has been studied extensively in the context of plagiarism detection. The
annual PAN competitions, starting in 2009, have been a main driver for developing algorithms
for plagiarism detection and text reuse detection29[
          <xref ref-type="bibr" rid="ref24 ref45">, 27, 28, 43, 24</xref>
          ].
        </p>
        <p>
          Most research on text reuse detection focuses on modern texts, o昀琀en digital born and using
modern language, which has rules for spelling and syntax. Detecting text reuse becomes more
complicated for long serial archives covering historic documents from an extensive historical
period in which the language used had no consistent spelling, spelling changed over time, and
the digitisation of those documents introduces text recognition erro4r5s, [
          <xref ref-type="bibr" rid="ref41">44</xref>
          ].
        </p>
        <p>
          In addition to plagiarism detection, textual repetition has been studied extensively in the
context of text alignment, collation and comparison10[
          <xref ref-type="bibr" rid="ref15">, 40, 15</xref>
          ], and text reuse [45, 38]. But there
is remarkably little previous work focusing on the identi昀椀cation of formulaic expressions. We
found several digital humanities studies regarding structure in text in gene1r,a2l, [
          <xref ref-type="bibr" rid="ref32 ref37">31, 36, 45,
38, 39</xref>
          ] and some more speci昀椀c studies for technical text [
          <xref ref-type="bibr" rid="ref9">8</xref>
          ] and legal arguments3[
          <xref ref-type="bibr" rid="ref12 ref2">2, 11, 12</xref>
          ]. In
all these cases, the object of study is texrtepetition and the use of isolated speci昀椀c terminology
(noun phrases) rather than the use oformulaic phrases. To the best of our knowledge, outside
of linguistics, humanities scholars have not written much about formulaic language use.
Probably, only serial use of textual features make formulas useful for study. Scholars who have to
read through them tend to see them as repetitive textual features without relevant information
content.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Issues with variation-rich text</title>
        <p>One of the big challenges of text analysis on corpora of historical texts is that they are rich in
spelling variation. Many historical languages had no standard spelling and changed in spelling
over time. Moreover, many texts extracted from digitised documents contain text recognition
errors. These issues together lead to possibly many di昀erent spellings of the same word or
phrase.</p>
        <p>
          This challenge can to some extent be addressed by normalising the spelling. This maps
spelling variants of words to a standard or ‘normal’ spelling of the word. VAR5,D126][ is
a lexicon-based technique that was originally developed for historical English but ported to
di昀erent historical languages. TICCL [
          <xref ref-type="bibr" rid="ref39">34, 35</xref>
          ] was developed originally to automatically
normalise very large collections of 19th and 20th century Dutch. A di昀erent approach is to use
fuzzy string matching and searching starting from a list of known phrase2s2][. There are
recent techniques based on deep neural networks, like PIE23[], that can be trained to lemmatise
variation-rich languages, resulting in ‘normalised’ lemmas. However, this also reduces
morphological variation that can be meaningful in distinguishing between expressions. Another
drawback is that this requires a large amount of training material of linguistically annotated
text.
        </p>
        <p>
          Since we are using the same corpus as 2[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], we took inspiration from their fuzzy searching
approach, but since it requires knowing the formulas in advance and the technique becomes
very slow when a large number of formulas is used for searching, we decided to use a simpli昀椀ed
approach of detecting common word n-grams and using character n-gram indexing to 昀椀nd
orthographically similar spellings.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Formulaic expressions and their use</title>
      <p>We narrow down our object of research with a more precise but still pre-theoretical de昀椀nition
of formulaic expressions and their context. A literature search has not given us any de昀椀nition
of formulaic expressions beyond the notions olefxical bundles and idiomatic expressions in
common language use 7[, 26]. Lexical bundles are o昀琀en noun phrases and examples of domain
speci昀椀c terminology. We take a information theoretical perspective, and need a de昀椀nition that
is applicable to di昀erent corpora and that helps to identify formulaic expressions from large
amounts of text. Given the nature of the corpus of resolutions and the corpus-speci昀椀city of its
formulaic expressions, we need a de昀椀nition that takes into account that formulas tend to be
longer phrases, though not necessarily complete clausal units, that can incorporate and give
context to variable elements like names of persons, organisations and locations or dates.</p>
      <sec id="sec-3-1">
        <title>3.1. Characteristics of Formulaic Expressions</title>
        <p>We de昀椀ne formulaic expression as a multi-word phrase (anextended lexical unit) that is reused
o昀琀en across documents in a collection , with minimal word variation, but with potentially high
variation in spelling.1 They may contain variable elements, spans consisting of e.g. entity
names or dates. In the resolutions, a phrase might express that a certain type of proposal was
submitted by a person, whose name is a variable element in the formula. As far as we can
tell, this de昀椀nition captures the formulas found by Rutten and Wal as well. What constitutes
a formulaic expression might di昀er across domains, genres or corpora. We will discuss this
further at the end of this paper in Section6, but note here that this de昀椀nition does not yet
give us proper criteria for deciding what is a formula and what is not. Therefore, our research
design is exploratory, rather than descriptive or explanato4r,ypp[.91-92]. It serves to give us
a better understanding of the phenomenon of formulaic expressions and how we can develop
methods to study them, rather than to o昀er a precise description of how they are used or an
explanation of how they emerge or evolve.</p>
        <p>The study of formulaic language use has identi昀椀ed a number of reasons for speakers and
authors to use formulaic expression4s2[]. Within the domain of legal and administrative texts,
the most relevant are precision of communicated information (e.g. “Cleared for takeo昀” signals
permission to enter a runway and commence takeo昀), and signalling the structure of discourse
(e.g. “on the other hand” signals an opposition)1[3, p.46].</p>
        <p>The formulaic expressions that are used over and over in many historical legal and
administrative documents, serve as precise referents to the information that the document is to
communicate. They prevent di昀erences in interpretation, but also structure the information of the
document. For instance, in a corpus of proclamations, a standard opening phrase signals where
each proclamation starts. Notarial deeds o昀琀en have template text to indicate the role of the
actors involved in the deed and that the contract is a legally binding agreement between the
actors. In this way, formulaic expressions let us detect this structure.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Textual repetition across domains and genres</title>
        <p>We compare the corpus of resolutions against a set of historical and modern text corpora from
various domains, to get insight in how textually repetitive it is and whether that is related to
the domain and genre of administrative and legal texts.</p>
        <p>We use the following document collections to study the amount of textual repetition (T1a)b.le
There are 昀椀ve corpora consisting of mostly administrative documents:
• The Resolutions corpus contains 286,871 printed resolutions of the SG in the period
17051796, with a Character Error Rate (CER) of 3%.
• The Notarial Deeds of the city archive of Amsterdam contain legal transactions. We expect
deeds to have at least a few formulaic expressions stating the nature of the transaction,
the parties involved, and that it has been signed o昀 by a notary. We estimate the CER to
be in around 10%.
1The latter aspect is part of the de昀椀nition to account for historical variations in spelling of what is semantically or
pragmatically the same phrase.
• The collection oDfutch medieval charters contains manual transcriptions of handwritten
charters, which we expect to contain mostly formal language with low CER.
• The Mandate and Police books from the city state of Bern cover the same period (and more)
and also contain administrative and legal documents where we expect formal language
use, but in a di昀erent language (German) and with a higher rate of text recognition errors
(CER of around 20%). The two types of books represent di昀erent administrative
subgenres so we treat them as separate corpora.</p>
        <p>We compare these against 昀椀ve corpora with mostly free-form text from various domains,
including two with historic Dutch language and three corpora with modern Dutch:
• The general missives of the Verenigde Oostindische Compagnie (VOC, Dutch East India
Company) consist of long business correspondences between the o昀케ces of the VOC in
Asia and Amsterdam, which we expect are less formal and more free-form than the
resolutions. These are handwritten documents, for which we could not obtain accurate CER
information, but we estimate it to be in the range of 15-20%.
• The Dutch Newspaper corpus of the National Library of the Netherlands contains over
700,000 articles from dozens of Dutch language newspapers in the 18th century. This
corpus was OCR’ed around 2006 and has a CER of 15-20%. The articles are free-form and
cover many topics, so we expect low repetition.
• Dutch Wikipedia consisting of articles that are in principle free-form, but occasionally,
entire databases of e.g. sports clubs, television shows or plant species are algorithmically
turned into a set of article stubs using a template article. Although later manual edits
of a template-based article tend to transform the template text to more free-form prose,
some template phrasings may remain.
• Dutch novels, a set of 10,921 recently published Dutch novels (text extracted from epubs).</p>
        <p>We expect these to be free-form with little repetition.
• Book reviews is a set of 472,810 online Dutch book reviews9[, 21] from seven di昀erent
reviewing platforms. Reviews are also free-form, although book reviews represent a
very narrow domain with potentially many stock phrases, so we expect some form of
repetition.</p>
        <p>
          A Vocabulary Growth Curve [
          <xref ref-type="bibr" rid="ref3">3, 6</xref>
          ] shows how much the frequency of vocabulary terms grows
with respect to the fraction of terms in a collection that have been seen only once. By iterating
over all paragraphs in a corpus in a random order, we count the total number of tersmesen
so far (i.e. term tokens) and at 昀椀xed points—e.g. once every 1000 words—divide the size of the
vocabulary ( ) over the number of hapax legomena 1( ), e.g. terms that have occurred
( )
only once so far. A higher value of1( ) means more of the term frequency mass is taken by
terms that occur more than once.
        </p>
        <p>
          The vocabulary growth curves are shown in Figu2refor frequencies of word n-grams with
Ā ∈ [
          <xref ref-type="bibr" rid="ref1 ref3 ref6">1, 3, 5</xref>
          ]. The administrative corpora are shown with dashed lines, while the more free-form
corpora are shown with dotted lines. Further, corpora with high CER are shown with dense
lines (little horizontal space between the symbols), and the rest with more widely separated
symbols.
        </p>
        <p>For single words, the modern Dutch corpora have higher curves, meaning they have
relatively more terms that occur more frequently, than the historic corpora based on algorithmic
text recognition. This is not surprising, given the spelling variation and recognition errors in
historic corpora. Both phenomena increase the size of the vocabulary and thereby lead to more
mass at to the hapax legomena. Novels and reviews have the highest curves. We speculate that
novels tend to use mostly common vocabulary to be easy to read by a large audience, while
book reviews are a speci昀椀c domain and genre, so use a relatively narrow vocabulary. Wikipedia
contains articles about a huge range of topics, so it is understandable that it has a longer tail of
hapax legomena. The medieval charters use a very limited vocabulary and have a lot of term
repetition. As it is based on manual transcription, we assume the rate of recognition errors
to be much lower than for the corpora based on OCR or HTR. The low/high CER distinction
corresponds to a clear di昀erence, with all low CER corpora having much higher curves.</p>
        <p>For word 3-grams, the resolutions and Bern police and mandate books have curves that fall
o昀 very little, signalling that, although they have many single term hapax legomena, they have
relatively many word 3-grams that occur more than once, compared to the other corpora. The
curve for online book reviews overtakes the resolutions curve a昀琀er around 500,000 tokens,
suggesting indeed that additional reviews introduce relatively few new 3-grams and that reviews
are therefore relatively similar to each other. For word 5-grams, the top curves are those of the
charters, Bern police and mandate books, the resolutions, notarial deeds and the book reviews.
With the exception of the latter, they are all in the domains of legal and administrative texts.</p>
        <p>These curves support our intuition that texts in the legal domain have relatively many
repeated phrases, despite the spelling variation and character recognition mistakes.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Modelling Formulaic Expressions</title>
      <p>How frequent should a phrase be to be considered a formulaic expression? There can be
formulaic expressions that are borrowed from other domains or genres, that are used with low
frequency in the collection in which formulaic expressions are analysed. We leave these out of
the scope of this paper, as we want to focus on expressions that are frequent enough that they
can be used as metadata that cover most of the resolutions. Identifying borrowed formulas
requires knowledge or analysis of external resources.</p>
      <p>We want to 昀椀nd word sequences that occur frequently. The simplest way would be to count
the frequencies of word n-grams for some range oĀ,f similar to the word n-gram analysis of
Section 3.2. However, the number of n-gram types grows rapidly Āasincreases, and the vast
majority of these occur only once or twice. The corpus of resolutions has 735,919 distinct
words, so the number of word 1-gram types is the same, but foĀr= 4, the number of n-gram
types is 23,985,191. We exploit the fact that frequently occurring phrases can only consist of
words that individually occur at least as frequently as the phrases themselves. That is, phrases
that occur 100 times in the collection must consist of words that occur at least 100 times. To
椀昀nd phrases that occur at least ĂℎĄ ą Ą ÿă Ā= 100, we can exclude all candidate phrases that
contain words with a corpus frequency below this phrase frequency threshold. Furthermore,
the words within a phrase also co-occur with each other at least 100 times within a window
that is equal to the word length of the phrase.</p>
      <p>With these observations in mind, we developed a naive algorithm for identifying candidate
formulaic expressions. In the pre-processing phrase, candidate phrases of 昀椀xed length are
detected, a昀琀er which their contexts of preceding and following words are clustered and analysed
to extend the partial formulas and identify their start and end boundaries. The
parameterisation we arrive at is ad hoc and speci昀椀c this the corpus of resolutions, since we have no precise
de昀椀nition yet of what makes a phrase formulaic in a particular context. The goal is to explore
the corpus with a pre-theoretical notion of what we are looking for.</p>
      <p>Concretely, the pre-processing phase consist of the following steps:
1. Tokenise each resolution into sentences and sentences into words.
2. Iterate over the corpus and count frequencies of individual words
3. Iterate over the corpus a second time, and replace each word with a variable to&lt;kVeAnR&gt;
if it either has 1) a term frequencyĆ Ąÿ Ą ă( ) &lt; ĂℎĄ ą Ą ÿă Ā, or 2) a co-occurrence
frequency āā Ą ă(, ) &lt; ĂℎĄ ą Ą ÿă Āwith at least one of its remaining
neighbouring = 5 words ∈ ( − , + ) on either side.
4. Slide a 5-word window over each individual sentence, extract the 5-word window as a
phrase if it contains no variable token, and count the frequency of each extracted phrase.
The process is demonstrated in detail in AppendiAx.</p>
      <p>In the extension phase, we reduce the set of candidate phrases to a set of formulas in two
steps. First, we use fuzzy string matching to 昀椀nd clusters of candidate phrases that are spelling
variations of each other. In the second step, we gather the contexts around each occurrence of
a cluster of phrases, and count how o昀琀en the phrases are preceded and followed by the same
sequence of words.</p>
      <sec id="sec-4-1">
        <title>4.1. Clustering phrase spelling variants</title>
        <p>Many of the common word 5-grams are spelling variants of each other. We cluster them by
indexing these word 5-grams as vectors of character 1-skip-2-grams. That is, we consider not
only 2 adjacent characters, but also pairs of characters that are separated by another character.</p>
        <p>Starting from the most frequent word 5-gram phrases, we query the index to 昀椀nd candidate
variants using cosine similarity. Further details are provided in AppenBd. ix</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Extending partial formulas</title>
        <p>
          Next, we build frequency lists of the 8 words preceding and following the 昀椀xed length phrase
and use transition probabilities to identify extensions that have a probability close to 1 of
preceding or following the phrase. This is inspired by probabilistic language models based on
Hidden Markov Models3[
          <xref ref-type="bibr" rid="ref17">0, 17</xref>
          ]. We note that there might be formulas shorter than 5 words.
These can be detected by starting with shorter 昀椀xed length phrases.
        </p>
        <p>Because the preceding (pre昀椀x) and following (post昀椀x) contexts can include clusters of
spelling variants as well, we use the same fuzzy matching algorithm as used for clustering
the phrases. We then split all 8-word contexts into sequences of words and calculate transition
probabilities for pre昀椀x and post昀椀x contexts separately, starting from the 昀椀xed length phrase
to the word immediately preceding or following it, and from that word to the next word, etc.
Words that occur in multiple pre昀椀x or post昀椀x contexts thereby have a higher transition
probability. Words that have a probability below 0.1 are considered to be not part of the formulaic
expression and are replaced by &lt;aVAR&gt; token. Once all transition probabilities have been
computed, we traverse the transition model starting from the 昀椀xed length phrase and consider
preceding words part of the formulaic expression if the probability is above 0.9. Once the
probability drops below 0.9 but is still above 0.1, we assume to have reached a common context of
the formula that is not part of the formula itself. We repeat the same process for the post-phrase
context, again, computing transition probabilities starting from the 昀椀xed phrase.</p>
        <p>This process is described in more detail in AppendBi x.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>We start with the 23,141 candidate phrases that we got from using a minimum frequency
threshold of 100. Clustering variant phrases reduces this to 12,880 clusters. The most frequent phrase
is considered the representative variant. Extending these phrases with preceding and following
words with a transition probability above 0.9 results in 11,497 candidate formulas and 52,348
common extensions (with cumulative transition probabiliti0e.s1 ≤ ĂĆĄ Āą &lt; 0.9).</p>
      <p>The 10 most frequent phrases are shown in Tabl2e, together with their corpus frequencies
and the formulas that were derived from analysing their contexts. The phrase ‘&lt;START&gt;
ontfangen een missive van’ is the most frequent 昀椀xed-length phrase that is also the most frequent
formula (there is no extension that has a transition probability above 0.9). A less frequent and
partially overlapping phrase is ‘ontfangen een missive van den’, which in the extension step is
extended to ‘&lt;START&gt; ontfangen een missive van den’, that is, the ‘&lt;START&gt;’ token is added
to it. The 昀椀rst two formulas therefore also partially overlap, but the second is an extension of
the 昀椀rst. The word ‘den’ (EN: the) is the most common continuation of the 昀椀rst formula, but
other common continuations are names of persons, so the second formula is less ‘formulaic’
than the 昀椀rst. Phrases 3 and 4 lead to the exact same formula, as do phrases 5, 7, 8 and 9. Phrase
10 is also a partial overlap with these phrases, but because it includes the word ‘dat’ (tEhNat:)
at the end — which is again a common follow up, but not the only one — it is not extended to
the same formula.</p>
      <p>We can further cluster these formulas by re-categorising the less frequent formulas that are
extended or reduced versions of more frequent formulas as commexotnensions. If we perform
this re-categorisation, the list of 11,497 formulas is reduced 7,153 formulas. Eyeballing the list
of remaining formulas reveals there are still many spelling variants in the list. This shows that
spelling variant is a challenging problem that needs to be analysed in more detail.</p>
      <sec id="sec-5-1">
        <title>5.1. Analysis of formulas</title>
        <p>Almost 58% of the candidate formulas are longer than 5 words. Of course, our choice to start
from candidate phrases of 5 words ensures no formulas shorter than 5 words are found, but
the fact that more than half are extended shows that formulaic expressions in the resolutions
are long syntactic units consisting of more than compound nouns and noun phrases.</p>
        <p>The 10 formulas resulting from clustering the most frequent formulas are shown in T3a.ble
The last column describes how we can use these formulas to identify meaningful elements in
the running text, and how they help in classifying these elements. It is worth noticing that
many formulas precede of follow a named entity, suggesting that formulas were frequently
used to assert the relationship of that entity to the proposition or decision. Several formulas
also contain verbs h(as been read, has been agreed and understood, to report back) that signal
speci昀椀c actions. In future work, we will analyse the types of relationships and actions that
these formulas express.
waar op gedelibereert zynde is
goetgevonden en verstaan
which, upon deliberation, has
been agreed and understood
en van alles alhier ter
vergaderinge rapport te doen
&lt;END&gt;
and to report back on
everything, here in the meeting
en andere haar hoogh mogende
gedeputeerden tot de
and other high and mighty
deputies of
de heeren gedeputeerden van
de</p>
        <p>the gentlemen deputies of the
in handen van de heeren</p>
        <p>in the hands of the gentlemen
haar hoogh mogende resolutie
van den
resolution of her high and
mighty of the
aan het hof van sijne</p>
        <p>at the court of his
10</p>
        <p>Is ter vergaderinge gelesen de
requeste van</p>
        <p>Has been read in this meeting,
the petition of
Applying the approach to other corpora is elaborated in AppenEd.ix</p>
        <p>Signal
this is the start of a resolution
and the start of the proposal
paragraph, the proposal
document type is missive
start of the decision paragraph
decision to start a investigative
committee that will report back
at a later date, signal that there
is a future resolution related to
this one
name preceding this formula is
a deputy, name following this
formula is an institution, the
deputy is a representative of the
institution
the name following this
formula is a province or institution
decision that the matter is
handed to a committee to
investigate, the name(s) following
this formula are the members of
this committee
what follows is a date of a
previous resolution, the previous
resolution is related to this
resolution
the name preceding this
formula is a representative of the
court of the name following this
formula
this is the start of a resolution
and the start of the proposal
paragraph, the proposal
document type is missive</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Challenges of Evaluation</title>
        <p>The detection approach above is exploratory, as we have no precise de昀椀nition of what a
formulaic expression is. We need clear criteria to determine if a phrase is formulaic before we can
precisely de昀椀ne the task of formula detection and quantitatively evaluate methods designed to
perform that task. For a proper evaluation of the detection method the entire corpus needs to
be annotated with all formulas. We could reduce the problem by focusing on a small sample
of resolutions and annotate anything that we think is a formula, but we would still need clear
criteria to decide what is a formula and what is not. Another alternative is to use simulation
and generate text and insert arti昀椀cially generated formulaic expressions to fully control the
characteristics of formula, including their length, variation and frequency of occurrence.</p>
        <p>Intuitively, a quantitative evaluation should consider precision and recall of detecting
formulaic expressions, with two di昀erent measures of recall. One is the fraction of di昀erent types
of formulaic expressions that are identi昀椀ed (type-based recall), and the other is the fraction of
occurrences of formulaic expressions that are detected (token-based recall).</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and Conclusions</title>
      <p>Formulas and their use have relevance to a number of both methodological and humanities
research questions. Because formulaic expressions and their use for text structuring are
understudied, for a large part we can only raise these research questions.</p>
      <p>Our 昀椀ndings of the use of repetitive phrases in various corpora suggests that repetitive
language use di昀ers strongly across domains and genres, with texts in administrative domains
containing more repetitive phrasing. It is not clear whether this is true for all or just a
speci昀椀c type of administrative texts. Our 昀椀ndings with detecting formulas suggest that this type
of serial government sources may contain many formulaic expressions that can be exploited
to structure texts and extract information. Further research has to shed light on the extent to
which this also holds for other administrative sources and how the composition and diversity of
collections relates to the statistical properties of formulas. A second type of research questions
centres on the use of formulas. We do not know whether comparable administrative archives
from di昀erent periods have the same rate of repetitive phrases and formulas. We observed
that formulas emerged, changed and disappeared over time, but not at what rate and what the
causes were. We do not know if there was an increase in adoption of formulaic expressions in
the resolutions. There could have been changes in legal or administrative customs, more
general language and cultural changes or perhaps even in昀氀uences of speci昀椀c scribes. The switch
to the use of printing would lead us to assume that there was less variation in the formulas
used, but it is too early to test this assumption. It is also an open question where the formulas
originate, and whether formulas are reused across (administrative) domains. A formula
borrowed from another domain might not be used o昀琀en in a corpus, in which case the method
and de昀椀nition we developed in this paper do not su昀케ce. So further discussion is needed on
what constitutes a formulaic expression, and what its generic and context-speci昀椀c elements
are. As for the content of formulas, we can only make some remarks about those that occur in
the resolutions. We have been able to spot a great number of formulas but it is hard to give a
precise de昀椀nition of what a formula is, or to establish if we can discern constituent elements
and if it makes sense to divide up formulas in these constituent parts. In this paper we describe
a methodology in development, that iteratively gathers formulas from our corpus. This works
well for identifying the most frequently used formulas and their variations. Further steps have
to make clear what the optimal rate of formula detection is, how to categorise the di昀erent
formulas, and if they are su昀케cient, to identify and localise di昀erent logical elements in the
text and whether this works for all resolutions. A further point of discussion is how much of
the methodology can be used for other corpora. We believe that the general methodology is
suitable for comparable text corpora.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Acknowledgments</title>
      <p>This work is part of the REPUBLIC project (2019-2024), a Research Infrastructure project funded
by the Dutch Research Council (NWO, Grant number 175.217.024).</p>
      <p>D. S. Carvalho, V. D. Tran, K. Van Tran, V. D. Lai, and M.-L. Nguyen. “Lexical to
DiscourseLevel Corpus Modeling for Legal Question Answering”.TInen:th International Workshop
on Juris-Informatics (JURISIN). 2016.
[28]
[42]</p>
      <p>A. Wray. Formulaic language and the lexicon. Eric, 2002.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Appendix</title>
    </sec>
    <sec id="sec-9">
      <title>A. Identifying candidate partial formulas</title>
      <p>We demonstrate the processing steps to identify candidate partial formulas using the following
example sentences:</p>
      <p>Ontfangen een Missive van het Collegie ter Admiraliteyt in Zeelandt, geschreven
te Middelburgh den negentienden deser loopende maandt, houdende, in gevolge en
tot voldoeninge van haar Hoogh Mogende Resolutie van den vyfden der voorlede
maandt, der zelver advis op het verzoeck van Burgermeesters en Scheepenen van
het hooge en laage Zas van Gent.</p>
      <p>With word tokenisation, we add&lt;START&gt; and &lt;END&gt; tokens so that for words at the start or
end of a sentence, this boundary is included in its context. Certain phrases only appear at or
near the start or end of a sentence, and adding boundary tokens allows us to keep track of such
cases.</p>
      <p>A昀琀er 昀椀ltering out low frequency words and words that do not meet the co-occurrence
frequency threshold, we end up with the following list:
[‘&lt;START&gt;’, ’ontfangen’, ’een’, ’missive’, ’van’, ’het’, ’collegie’, ’ter’, ’admiraliteyt’,
’in’, ‘&lt;VAR&gt;’, ’geschreven’, ’te’, ‘&lt;VAR&gt;’, ’den’, ‘&lt;VAR&gt;’, ’deser’, ’loopende’, ’maandt’,
’houdende’, ’in’, ‘gevolge’, ’en’, ’tot’, ’voldoeninge’, ’van’, ’haar’, ’hoogh’, ’mogende’,
’resolutie’, ’van’, ’den’,&lt;‘VAR&gt;’, ’der’, ’voorlede’, ’maandt’, ’der’, ’zelver’, ’advis’, ’op’,
’het’, ‘&lt;VAR&gt;’, ’van’, ‘&lt;VAR&gt;’, ’en’, ‘&lt;VAR&gt;’, ’van’, ’het’, ‘&lt;VAR&gt;’, ’en’, ‘&lt;VAR&gt;’, ‘&lt;VAR&gt;’,
’van’, ‘&lt;VAR&gt;’, ‘&lt;END&gt;’]</p>
      <p>The example above shows that most of the words in the sentence occur frequently and
cooccur with each other frequently. This results in a large number of candidate phrases. In the
resulting sentence, we extract candidate phraseĂs( , +5by identifying sequences of word
tokens (i.e. non-variable tokens) of length|Ă| = 5 and count their frequency over the entire
corpus.</p>
      <p>As little is known in advance about the relationship between frequencies of words, word
co-occurrence and formulas, we have few meaningful clues for setting a minimum phrase
frequency. The corpus of resolutions has 286,871 resolutions and over 58 million words, but we
have no reliable estimates of the total number of formulaic expressions that are used and how
o昀琀en they occur. In identifying the start of proposition paragraphs of resolution2s,2][ use a
list of 32 formulas, most of which are between 5 and 10 words long. But we do not know if
these are all the formulas used for openings of proposition and decision paragraph. We also do
not know which other elements of a resolution are expressed in formulas.</p>
      <p>To get some insight, we experimented with minimum frequencies of di昀erent orders of
magnitude, ranging from 1 to 10,000. The results are shown in Tab4l,ewith for each frequency
threshold, the size of the vocabulary (number of distinct words), the number of distinct
cooccurring word-pairs within a 5 word window, the number of distinct 5-word phrases (phrase
types) a昀琀er pre-processing, the total number of phrases (phrase tokens) and the number of
phrases that meet the frequency threshold.</p>
      <p>Frequency thresholds of 1 and 10 lead to large vocabularies and huge numbers of phrases.
If they are to be made useful as metadata labels or boundary signals, we need to analyse each
of them in their context manually, so these thresholds result in an unmanageable number of
phrases. On the other extreme, a threshold of 10,000 leads to a very small vocabulary of 605
highly common words and only 140 highly frequent phrases. The most common phrase is
&lt;START&gt; Ontfangen een Missive van’ (EN: &lt;START&gt; Received a missive of), which is also the
start of the example sentence above, with a frequency of 138,682. Note that there are variant
spellings of this phrase that are not included in the count. With such a high frequency, it is
clear that this phrase is part of a 昀椀xed formula, but because we made 昀椀xed-length phrases, it
is not clear whether this the entire formula.</p>
      <p>It is quite likely that some formulas contain words with a frequency below 10,000, and
because we know so little yet about the usage of formulas, it is better to be conservative and
choose a lower threshold. For the rest of this paper, we pick a minimum frequency threshold
ĂℎĄ ą Ą ÿă Ā = 100. This gives us a large set of 23,141 frequent phrases that are candidate
formulaic expression or part of them.</p>
    </sec>
    <sec id="sec-10">
      <title>B. From candidate phrases to formulaic expressions</title>
      <p>The process of transforming the set of candidate phrases to a set of formulas has two main
steps. First, we use fuzzy string matching to 昀椀nd clusters of candidate phrases that are spelling
variations of each other. In the second step, we gather the contexts around each occurrence of
a cluster of phrases, and count how o昀琀en the phrases are preceded and followed by the same
sequence of words.</p>
      <sec id="sec-10-1">
        <title>B.1. Clustering variant phrases</title>
        <p>We index each word 5-gram phrase as a vector of character 1-skip-2-grams. That is, we
consider not only 2 adjacent characters, but also pairs of characters that are separated by another
character. The reason to include a single skip is that spelling variants o昀琀en have di昀erences in
characters in the middle of a word, which results in multiple ngram mismatches and thus few
matches when no skips are used.</p>
        <p>Starting from the most frequent word 5-gram phrases, we query the index to 昀椀nd candidate
variants using cosine similarity. Further details are provided in AppenBd. iWxe limit the
candidate set to phrases that di昀er in length with the query phrase by at most 2 characters, based
on the assumption that much longer or shorter phrases are unlikely to be spelling variants. We
椀昀lter the candidate phrases by checking that the words in each position 1..5 of the candidate
phrase di昀er in length no more than 2 characters with their aligned words in the query phrase.
This avoids matching a query phrase ‘op gedelibereert zynde is goetgevonden’ with its partial
overlap ‘gedelibereert zynde is goetgevonden en’. The latter is typically the next 5-word phrase
following the former, but because the query phrase has a short 昀椀rst word and the candidate
phrase has a short last word, their ngram similarity is high. Using the length restriction on
aligned words 昀椀lters out such erroneous matches.</p>
      </sec>
      <sec id="sec-10-2">
        <title>B.2. Extending candidate formulas</title>
        <p>
          In the extension step, we build frequency lists of the 8 words preceding and following the 昀椀xed
length phrase and use transition probabilities to identify extensions that have a probability
close to 1 of preceding or following the phrase. This is a similar approach to probabilistic
language modelling1[
          <xref ref-type="bibr" rid="ref10">9</xref>
          ], where the probably that a word is followed by word , is calculated
as:
        </p>
        <p>This models the prediction of the next word only the current word. A natural extension is
to model the prediction based on all words preceding it:
( | ) =
( , )</p>
        <p>( )
( | 1, 2, ..., ) =
( 1, 2, ..., , )
( 1, 2, ..., )
(1)
(2)</p>
        <p>For extending phrases, we use this model, where the probability(ĂℎĄ) of a phrase ĂℎĄ
consisting of words , ..., is ( , ..., ). The transition probability from a phrasĂeℎĄ to a
(ĂℎĄ, ) ( ,ĂℎĄ)
word is (ĂℎĄ) and from a word to ĂℎĄ is (ĂℎĄ) .</p>
        <p>We split all 8-word contexts into sequences of words and calculate transition probabilities
for pre昀椀x and post昀椀x contexts separately, starting from the 昀椀xed length phrase to the word
immediately preceding or following it, and from that word to the next word, etc. Words that
occur in multiple pre昀椀x or post昀椀x contexts thereby have a higher transition probability. Words
that have a probability below 0.1 are considered to be not part of the formulaic expression
and are replaced by a&lt;VAR&gt; token. Once all transition probabilities have been computed, we
traverse the transition model starting from the 昀椀xed length phrase and consider preceding
words part of the formulaic expression if the probability is above 0.9. Once the probability
drops below 0.9 but is still above 0.1, we assume to have reached a common context of the
formula that is not part of the formula itself. We repeat the same process for the post-phrase
context, again, computing transition probabilities starting from the 昀椀xed phrase.</p>
        <p>We extend the 昀椀xed length phrase with up to 8 words preceding it and up to 8 words
following it, which means we can identify formulas o8f + 5 + 8 = 21 words in a single pass. If the full
8-word path preceding or following the phrase has a cumulative transition probability close to
1, we repeat this extension process with the extended phrase to identify the boundary of the
formula.</p>
        <p>An example of extending the phrase ’gevolge en tot voldoeninge van’ using transition
probabilities is shown in Figure3. On the le昀琀 side, the pre昀椀x context is shown, with the word ’in’
being the only word directly preceding the partial phrase. This means that ’gevolge en tot
voldoeninge van’ is not the full formulaic expression, but that ’in’ is also part of it. The
expression ’in gevolge en tot voldoeninge van’ is also a syntactically more comprehensible phrase,
meaning in consequence and ful昀椀lment of . There are multiple possible words preceding it. Two
common ones are the verbs ’hebbende’ (EN:having), which precedes the phrase ’in gevolge
en tot voldoeninge van’ in 38% of the occurrences of the phrase, and ’houdende’
(EmNa:intaining) which precedes the phrase in 47% of its occurrences. The remaining occurrences of
the phrase are preceded by a variety of other words. Each of these two verbs have multiple
possible pre昀椀xes, but are themselves not part of the formula according to the de昀椀nition above,
because their transition cumulative transition probabiliti1e.0s ∗(0.47 = 0.47 and 1.0∗0.38 = 0.38
respectively) do not meet the threshold o0f.9.</p>
        <p>On the right side, the post昀椀x context is shown, with two variants of continuations, neither
of which is part of the formulaic expression itself. One common continuation is ’der selver
resolutie’ and the other is ’haar hoog mogende resolutie’. Note that although neither path to the
word ’resolutie’ has itself a cumulative transition probability above00. 3.99(∗ 0.99 ∗ 0.99 = 0.38
for the former and0.59 ∗ 0.99 ∗ 0.99 ∗ 0.96 = 0.56 for the latter), their combined probabilities add
up to 0.94. This meets the 0.9 probability threshold, thereby being an example of a formulaic
expression that can have a variable middle part. In some cases that variable part is ’der selver’
and in others it is ’haar hoogh mogende’.</p>
        <p>The 昀椀nal formula is determined by extending the 昀椀xed length phrase with preceding and
following words that have a cumulative probability of at least 0.9. In the case of the phrase
’gevolge en tot voldoeninge van’, we extend with the pre昀椀x ’in’ to ’in gevolge en tot voldoeninge
van’</p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>C. The impact of spelling change</title>
      <p>One of the big hurdles is spelling change. Finding phrases that are orthographically similar to
each other is not hard, but phrases o昀琀en contain highly frequent, short function words that
require only few character edits to transform one function word into another. This makes
it di昀케cult to distinguish cases where two function words are variant spellings of each other
from cases where they represent di昀erent words and therefore signal that these phrases have
di昀erent meanings.</p>
      <p>
        We experimented with both classic Word2Vec2[
        <xref ref-type="bibr" rid="ref6">5</xref>
        ] and fastText [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] CBOW embeddings2 to
identify variant spellings of words, choosing the latter as it has better performance on a test
set of target words and their variants. FastText uses character-level embeddings that are more
2For both models we use their implementations in Gensim 4.303[], see https://radimrehurek.com/gensim/
suitable for detecting spelling variations. We use embeddings based on the assumption that
variants occur in the same or similar contexts so should end up in the same region in the
embedding space. This works for spelling variants that are used interchangeably in a single time
window. An example in the Resolutions corpus is the word ’en’ (EaNn:d) and its variant ’ende’.
There is a period in the 18th century when their uses overlap, as can be seen in the temporal
frequency distributions of the variant phrases ’gevolge en tot voldoeninge van’
(EfoNll:owing and in ful昀椀lment of ) and ’gevolge ende tot voldoeninge van’ (see the le昀琀 side of Figur4e).
The two versions ‘en’ and ‘ende’ share many contexts, so their word embeddings are similar.
Using a combination of orthographic similarity (edit distance) and embedding similarity, we
can identify word pairs in the corpus that can be linked as variants. However, experiments
on 昀椀nding variant spellings of words in the pre- and post-context of phrases have shown that
for short function words, spelling change is a major hurdle for both Word2Vec and FastText.
But for spelling changes where one variant is used in only one period and the other only in
another, non-overlapping period, their contexts can also have di昀erent spellings. This results
in two sets of contexts that also have no or little overlap. Hence, two spelling variants used
in di昀erent time periods may end up in di昀erent regions in the embedding space. An example
is the use ’ae’ in the early 18th century in words like ’aen’ (ENto: and ’haer’ (EN:her), which
changed to using ’aa’ from around 1717, when they switched to writing ’aan’ en ’haar’. These
words o昀琀en appear together, as in the common phrase ’aen haer hoogh mogende te’ (ENt:o
her high and mighty at). Because the spelling change for ’aen’ occurs at the same time as the
spelling change for its common contextual term ’haer’, the variants ’aen’ and ’aan’ have little
contextual overlap, so word embeddings consider them as di昀erent words.
      </p>
    </sec>
    <sec id="sec-12">
      <title>D. The impact of resolution length</title>
      <p>One of the characteristic of resolutions that we can study with our list of formulas is the fraction
of a resolutions text is made up of formulaic expressions, and which part is not. There are many
very short resolutions based on a received missive (starting with the formula ‘Ontfangen een
Missive van’) that merely states who wrote the missive, when and where they wrote it, but that
do not contain any proposal or request that the SG had to make a decision on. Such resolutions
end with the formula ‘Waar op geen resolutie is gevallen.’ (EONn: which no decision was made.</p>
      <p>We therefore expect there to be a relationship between the length of a resolution and the
amount of non-formulaic content. Longer resolution provide more detail of the proposition or
of the decision or both. We assume that these details are only given when they are deemed
relevant and necessary. The details vary across resolutions, therefore lead to less formulaic
text.</p>
      <p>The relationship between the length of resolutions and the fraction of words that are part
of formulaic expressions is shown in Figur5e. There is a clear relationship: short resolutions
tend to have a larger fraction of formulaic content. As resolutions get longer, a larger fraction
of the words they contain are below one of the two frequency thresholds.</p>
    </sec>
    <sec id="sec-13">
      <title>E. Formulas in Other Corpora</title>
      <p>To check if this approach generalises to other corpora, we use the same detection process on
the corpora of Notarial Deeds, Bern Manate books and t he Dutch Wikipedia.</p>
      <sec id="sec-13-1">
        <title>E.1. Mandate books of State of Bern</title>
        <p>Because of the high Character Error Rate (CE∼R0.2) and the much smaller size of the corpus
(3.8 million words compared to 58 million of the Resolutions), we used word 4-grams instead
of 5-grams. The procedure found a handful of candidates, two of which could be extended to
• The most common phrase is ’Schultheiß und Rath der Statt Bern’, which occur 907 times
in 505 di昀erent spellings. It refers to the head o昀케cial and the council of the city of Bern.
• The second most frequently identi昀椀ed phrase is ’An alle Deütsch und Weltsche’, which
occurs 559 times in 443 di昀erent spellings, and is extended to the formula ‘An alle Deütsch
und Weltsche Herren Amtleüth’ (ENT:o all German and X gentlemen o昀케cials ). This is a
formula to signal that the following statute pertains to the o昀케cials of both the French
and German speaking parts of the city state Bern, and therefore signals the start of a
statute.</p>
      </sec>
      <sec id="sec-13-2">
        <title>E.2. Notarial deeds from Amsterdam municipality</title>
        <p>The notarial deeds corpus also has a high CER of 15-20%. The most frequently found phrase
are:
• ‘als getuijgen hier overgestaen’ (ENs:tanding here as witnesses: this is part of a formulaic
phrase in the opening paragraph of a notarial deed to indicate who act as witnesses in
formalising the transaction. This is useful in identifying the starting paragraphs of deeds
that are spread across pages.
• ‘H Schaef N P’: the name one of the Amsterdam notaries, who is the o昀케cial responsible
for ensuring the transaction is legal.</p>
        <p>• ‘J de Winter N P’: the name of another Amsterdam notary.</p>
        <p>These phrases have similar potential in identifying meaningful structural elements in the
running text, such as where deeds start and end, and where certain elements of deeds are
located within their text.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Altmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Köhler</surname>
          </string-name>
          . “Forms and Degrees of Repetition in Texts”. IFno:rms and Degrees of Repetition in Texts. De Gruyter Mouton,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Auslander</surname>
          </string-name>
          . “On Repetition”.
          <source>InP:erformance Research 23</source>
          .
          <fpage>4</fpage>
          -
          <lpage>5</lpage>
          (
          <year>2018</year>
          ), pp.
          <fpage>88</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Baayen</surname>
          </string-name>
          . “
          <article-title>The e昀ects of lexical specialization on the growth curve of the vocabulary”</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>In: Computational Linguistics 22.4</source>
          (
          <issue>1996</issue>
          ), pp.
          <fpage>455</fpage>
          -
          <lpage>480</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E. R.</given-names>
            <surname>Babbie</surname>
          </string-name>
          .
          <article-title>The practice of social research</article-title>
          .
          <source>Cengage learning</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5] [6] [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baron</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Rayson</surname>
          </string-name>
          . “
          <article-title>VARD2: A tool for dealing with spelling variation in historical corpora”</article-title>
          .
          <source>In:Postgraduate conference in corpus linguistics</source>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>In: Available atzipfr</article-title>
          . r-forge. r-project. org/materials/zipfrtutorial. pdf
          <source>[last accessed1 June</source>
          <year>2019</year>
          ] (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Biber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Johansson</surname>
          </string-name>
          , G. Leech,
          <string-name>
            <given-names>S.</given-names>
            <surname>Conrad</surname>
          </string-name>
          , E. Finegan, and R. QuirkL.
          <article-title>ongman grammar of spoken and written English</article-title>
          . Vol.
          <volume>2</volume>
          . Longman London,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Boguraev</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Kennedy</surname>
          </string-name>
          . “
          <article-title>Technical Terminology for Domain Speci昀椀cation and Content Characterisation”</article-title>
          .
          <source>In:Scie</source>
          .
          <year>1997</year>
          . doi:
          <volume>10</volume>
          .1007/3-540-63438-x\_5.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Boot</surname>
          </string-name>
          .
          <article-title>“A Database of Online Book Response and the Nature of the Literary Thriller</article-title>
          .” In: Dh.
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Cannon</surname>
          </string-name>
          . “
          <article-title>OPCOL: An Optimal Text Collation Algorithm”</article-title>
          .
          <source>ICn:omputers and the Humanities</source>
          (
          <year>1976</year>
          ), pp.
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>T.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.-X.</given-names>
            <surname>Tran</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.-L.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          . “
          <article-title>Lexical-Morphological Modeling for Legal Text Analysis”</article-title>
          .
          <source>InJ:SAI International Symposium on Arti昀椀cial Intelligence</source>
          . Springer,
          <year>2015</year>
          , pp.
          <fpage>295</fpage>
          -
          <lpage>311</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Conklin</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Schmitt</surname>
          </string-name>
          . “
          <article-title>The processing of formulaic language”</article-title>
          .
          <source>IAnn: nual Review of Applied Linguistics</source>
          <volume>32</volume>
          (
          <year>2012</year>
          ), pp.
          <fpage>45</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Cowie</surname>
          </string-name>
          .
          <article-title>Phraseology: Theory, analysis, and applications</article-title>
          .
          <source>OUP Oxford</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>R. Haentjens</given-names>
            <surname>Dekker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Van</given-names>
            <surname>Hulle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Middell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Neyt</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Van Zundert. “</surname>
          </string-name>
          <article-title>Computersupported collation of modern manuscripts: CollateX and the Beckett Digital Manuscript Project”</article-title>
          .
          <source>In:Digital Scholarship in the Humanities 30.3</source>
          (
          <issue>2015</issue>
          ), pp.
          <fpage>452</fpage>
          -
          <lpage>470</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>I. Hendrickx</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Marquilhas</surname>
          </string-name>
          . “From Old Texts to Modern Spellings:
          <article-title>An Experiment in Automatic Normalisation</article-title>
          .” InJ:.
          <source>Lang. Technol. Comput. Linguistics 26.2</source>
          (
          <issue>2011</issue>
          ), pp.
          <fpage>65</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>F.</given-names>
            <surname>Jelinek</surname>
          </string-name>
          .
          <article-title>Statistical methods for speech recognition</article-title>
          . MIT press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          . “
          <article-title>Bag of tricks for e昀케cient text classi昀椀cation”</article-title>
          .
          <source>In: arXiv preprint arXiv:1607.01759</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Martin</surname>
          </string-name>
          . “Speech and
          <string-name>
            <given-names>Language</given-names>
            <surname>Processing</surname>
          </string-name>
          :
          <article-title>An introduction to speech recognition, computational linguistics and natural language processing”U.Ipn-: per Saddle River</article-title>
          , NJ: Prentice Hall (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>F.</given-names>
            <surname>Karsdorp</surname>
          </string-name>
          . “
          <article-title>Het is groen en lee昀琀 nog lang en gelukkig. Classi昀椀catie van volksverhaalgenres op basis van formules”</article-title>
          .
          <source>InT: ijdschri昀琀 voor Nederlandse Taal-en Letterkunde 129.4</source>
          (
          <issue>2014</issue>
          ), pp.
          <fpage>274</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Koolen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Boot</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. J. van Zundert. “</surname>
          </string-name>
          <article-title>Online Book Reviews and the Computational Modelling of Reading Impact”</article-title>
          .
          <source>InP:roceedings of the Workshop on Computational Humanities Research (CHR</source>
          <year>2020</year>
          ). Vol.
          <volume>2723</volume>
          .
          <year>2020</year>
          , p.
          <fpage>0073</fpage>
          . url:http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2723</volume>
          /lon g13.
          <source>pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Koolen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoekstra</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Nijenhuis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sluijter</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. van Gelder</surname>
          </string-name>
          , R. van Koert,
          <string-name>
            <given-names>G.</given-names>
            <surname>Brouwer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Brugman</surname>
          </string-name>
          . “
          <article-title>Modelling Resolutions of the Dutch States General for Digital Historical Research</article-title>
          .” In:Colco.
          <year>2020</year>
          , pp.
          <fpage>37</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [19] [21] [22] [23]
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          , Á. Kádár, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          . “
          <article-title>Improving Lemmatization of Non-Standard Languages with Joint Learning”</article-title>
          .
          <source>In:Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers). Minneapolis, Minnesota: Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>1493</fpage>
          -
          <lpage>1503</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          -1153. url: https://www .aclweb.org/anthology/N19-115.
          <fpage>3</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>V. T.</given-names>
            <surname>Martins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fonte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. R.</given-names>
            <surname>Henriques</surname>
          </string-name>
          , and D. d. Cruz. “
          <article-title>Plagiarism detection: A tool survey and comparison”</article-title>
          . In: (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          . “
          <article-title>Distributed representations of words and phrases and their compositionality”</article-title>
          .
          <source>IAn:dvances in neural information processing systems</source>
          <volume>26</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>M.</given-names>
            <surname>Poß</surname>
          </string-name>
          and T. v. d. Wouden. “
          <article-title>Extended lexical units in Dutch”</article-title>
          .
          <source>InLO:T Occasional Series</source>
          <volume>4</volume>
          (
          <year>2005</year>
          ), pp.
          <fpage>187</fpage>
          -
          <lpage>202</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Eiselt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Barrón Cedeño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          . “
          <article-title>Overview of the 3rd international competition on plagiarism detection”</article-title>
          .
          <source>CInE:UR workshop proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          Vol.
          <volume>1177</volume>
          . CEUR Workshop Proceedings.
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          Stein. “
          <article-title>Overview of the 5th international competition on plagiarism detection”</article-title>
          .
          <source>CInL:EF Conference on Multilingual and Multimodal Information Access Evaluation. Celct</source>
          .
          <year>2013</year>
          , pp.
          <fpage>301</fpage>
          -
          <lpage>331</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , E. Andreas, and
          <string-name>
            <surname>A. B.-C. P. Rosso</surname>
          </string-name>
          . “
          <article-title>Overview of the 1st international competition on plagiarism detection”</article-title>
          .
          <source>In3:rd PAN Workshop</source>
          . Uncovering Plagiarism,
          <source>Authorship and Social So昀琀ware Misuse</source>
          .
          <year>2009</year>
          , p.
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>L. R.</given-names>
            <surname>Rabiner</surname>
          </string-name>
          . “
          <article-title>A tutorial on hidden Markov models and selected applications in speech recognition”</article-title>
          .
          <source>In:Proceedings of the IEEE 77.2</source>
          (
          <issue>1989</issue>
          ), pp.
          <fpage>257</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Raney</surname>
          </string-name>
          .
          <article-title>“A Context-Dependent Representation Model for Explaining Text Repetition E昀ects”</article-title>
          .
          <source>In: Psychonomic Bulletin &amp; Review 10.1</source>
          (
          <issue>2003</issue>
          ), pp.
          <fpage>15</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rashidi-Tabrizi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Mussbacher</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Amyot</surname>
          </string-name>
          . “
          <article-title>Legal Requirements Analysis and Modeling with the Measured Compliance Pro昀椀le for the Goal-Oriented Requirement Language”</article-title>
          . In: 2013 6th International Workshop on Requirements Engineering and
          <string-name>
            <surname>Law (RELAW). Ieee</surname>
          </string-name>
          ,
          <year>2013</year>
          , pp.
          <fpage>53</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rehurek</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Sojka</surname>
          </string-name>
          . “
          <article-title>Gensim-python framework for vector space modelling”</article-title>
          .
          <source>In: NLP Centre</source>
          , Faculty of Informatics, Masaryk University, Brno,
          <source>Czech Republic 3.2</source>
          (
          <issue>2011</issue>
          ), p.
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          Springer.
          <year>2008</year>
          , pp.
          <fpage>617</fpage>
          -
          <lpage>630</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <source>In: Proceedings of coling</source>
          <year>2014</year>
          ,
          <article-title>the 25th international conference on computational linguistics: System demonstrations</article-title>
          .
          <year>2014</year>
          , pp.
          <fpage>52</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Radzikowska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Michura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fiorentino</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Clement</surname>
          </string-name>
          . “
          <article-title>Visualizing Repetition in Text”</article-title>
          .
          <source>In:Digital Studies/Le champ numérique 1</source>
          .3 (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>G.</given-names>
            <surname>Rutten</surname>
          </string-name>
          and
          <string-name>
            <surname>M. J. van der Wal.</surname>
          </string-name>
          “
          <article-title>Functions of epistolary formulae in Dutch letters from the seventeenth and eighteenth centuries”</article-title>
          .
          <source>In:Journal of Historical Pragmatics 13.2</source>
          (
          <issue>2012</issue>
          ), pp.
          <fpage>173</fpage>
          -
          <lpage>201</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [34] [35] [38] [39]
          <string-name>
            <given-names>H.</given-names>
            <surname>Salmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Paju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rantala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nivala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vesanto</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          . “
          <article-title>The reuse of texts in Finnish newspapers</article-title>
          and journals,
          <fpage>1771</fpage>
          -
          <lpage>1920</lpage>
          :
          <article-title>A digital humanities perspective”</article-title>
          .
          <source>In: Historical Methods: A Journal of Quantitative and Interdisciplinary History 54.1</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>14</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Samkova</surname>
          </string-name>
          . “
          <article-title>Repetition and Intertextuality as Modalities of Text Structuring and perception”</article-title>
          .
          <source>In:Facta Universitatis, Series: Linguistics and Literature</source>
          <volume>0</volume>
          (
          <year>2016</year>
          ), pp.
          <fpage>95</fpage>
          -
          <lpage>105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [44] [45] [46]
          <string-name>
            <given-names>D.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Colomb</surname>
          </string-name>
          . “
          <article-title>A data structure for representing multi-version texts online”.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <source>In: International Journal of Human-Computer Studies 67.6</source>
          (
          <issue>2009</issue>
          ), pp.
          <fpage>497</fpage>
          -
          <lpage>514</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <given-names>N.</given-names>
            <surname>Schmitt</surname>
          </string-name>
          .
          <article-title>Formulaic sequences: Acquisition, processing, and use</article-title>
          . Vol.
          <volume>9</volume>
          . John Benjamins Publishing,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <given-names>N.</given-names>
            <surname>Schmitt</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Carter</surname>
          </string-name>
          . “
          <article-title>Formulaic sequences in action”</article-title>
          .
          <source>InFo:rmulaic sequences: Acquisition, processing and use (</source>
          <year>2004</year>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [43]
          <string-name>
            <surname>E. Stamatatos.</surname>
          </string-name>
          “
          <article-title>Intrinsic plagiarism detection using character n-gram pro昀椀les”</article-title>
          .
          <source>Itnh:reshold 2.1</source>
          ,
          <issue>500</issue>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vesanto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Salmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nivala</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Salakoski</surname>
          </string-name>
          .
          <article-title>“A system for identifying and exploring text repetition in large historical document corporaP”</article-title>
          .
          <source>rIonc:eedings of the 21st Nordic Conference on Computational Linguistics</source>
          .
          <year>2017</year>
          , pp.
          <fpage>330</fpage>
          -
          <lpage>333</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vesanto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nivala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rantala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salakoski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Salmi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          . “
          <article-title>Applying BLAST to text reuse detection in 昀椀nnish newspapers</article-title>
          and journals,
          <fpage>1771</fpage>
          -
          <lpage>1910</lpage>
          ”.
          <source>InP:roceedings of the NoDaLiDa 2017 Workshop on Processing Historical Language</source>
          .
          <year>2017</year>
          , pp.
          <fpage>54</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Wood</surname>
          </string-name>
          . “
          <article-title>Uses and functions of formulaic sequences in second language speech: An exploration of the foundations of 昀氀uency”</article-title>
          .
          <source>In:Canadian Modern Language Review 63.1</source>
          (
          <issue>2006</issue>
          ), pp.
          <fpage>13</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>