<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Aarhus, Denmark
∗Corresponding author.
£ simon.gabay@unige.ch(S. Gabay);thibault.clerice@inria(T.f.rClérice)
ç https://cv.hal.science/simon-gab(aSy. Gabay);https://cv.hal.science/thibault-cle(Tr.icCelérice)
Ȉ</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>The birth of French orthography. A computational ⋆ analysis of French spelling systems in diachrony</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Inria Centre de Recherche de Paris</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>The 17th c. is crucial for the French language, as it sees the creation of a strict orthographic norm that largely persists to this day. Despite its significance, the history of spelling systems remains however an overlooked area in French linguistics for two reasons. On the one hand, spelling is made up of micro-changes which requires a quantitative approach, and on the other hand, no corpus is available due to the interventions of editors in almost all the texts already available. In this paper, we therefore propose a new corpus allowing such a study, as well as the extraction and analysis tools necessary for our research. By comparing the text extracted with OCR and a version automatically aligned with contemporary French spelling, we extract the variant zones, we categorise these variants, and we observe their frequency to study the (ortho)graphic change during the 17th century.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Computational linguistics</kwd>
        <kwd>History of orthography</kwd>
        <kwd>Information extraction</kwd>
        <kwd>Corpus building</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        One of the main technical challenges for carrying out such a study relies on the existence
of important amounts of data, in order to guarantee quantitatively the reliability of the results.
Unfortunately, such corpora of classical French do not exist for two reasons. On the one hand,
as Cl. Vachon 5[5, p. 32, n. 31] bitterly experienced, text editors got into the habit of
standardising the language of that 2e2r,a16[], which makes its study particularly complicated, if
not impossible to use for graphematic studies. On the other hand, the few corpora that have
been created, like that of Cl. Vachon, but also others like thaRtésoefauthCeorpus Français
Préclassique et Classique (RCFC) [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ] are not, or not in full, available to researchers. Given the
ever-increasing quantities of data necessary for computational studies, it is however dubious
that these two corpora, even freely accessible, would in any case remain insufÏcient for the
most recent approaches proposed in NLP.
      </p>
      <p>This paper proposes to return to the history of the Fvrêetnemchent graphique (“graphic
clothing”) in a computational way. We introduce a two-step approach: first, a unique
corpus creation pipeline meticulously extracts spelling information from digital facsimiles. This
pipeline includes a layout analysis model to distinguish text from paratext on the page, an OCR
model that retains the historical charſa›c,tpeivro‹tal to written French, and a linguistic
normaliser that “translates” historical French into its contemporary counterpart at the sentence
level. In the second step, we analyse the created corpus using a comparison algorithm that
matches the extracted historical text with its modern equivalent at the character level. This
enables us to pinpoint significant variations, categorise these diferences, and uncover detailed
trends throughout the 17th century. This methodological framework not only enhances our
understanding of historical French orthography, but also proposes a new approach for
computational linguistic studies of spelling variation.</p>
    </sec>
    <sec id="sec-2">
      <title>2. State of the art</title>
      <p>
        Corpus building from OCR has long been a task in digital humanities and corpus
linguistics. Initially deemed unsuitable for historical sources i4n4]1,9O93C[R gained credibility in
the late 1990s for corpus building, including XML TEI formalisation in commercial projects
such as thePatrologia Latina Database, and for Ancient Greek scripts in the 201409s].[Most
project using TEI, such as the First1KGreek project, relied on manual formalisation of the text’s
logical structur4e0][, as manual work was considered essential for accuracy. The advent of
user-friendly OCR and HTR technologies has spurred interest in automatic document
formalisation (ADF), primarily focused on facsimile formalisa5t2io]nan[d noisy text removal with
tools based on vocabularies such as SegmOn2t5o], [which standardises the identification of
paratextual zones (running titles, footnotes, etc.). Few projects, however, have utilised font,
geometric, and textual features to reconstruct or emulate the original text structure from
borndigital PDFs or OCR outputPsa.perXML [
        <xref ref-type="bibr" rid="ref50">51</xref>
        ] demonstrated such transformation but was
limited to the ACL Anthology structure. Gr5o0b]idan[d Grobid Dictionarie3s5,[
        <xref ref-type="bibr" rid="ref33">34</xref>
        ] employed
geometric, font, and textual features to produce XML TEI output, though they were specific
to scientific papers and dictionaries. In 2022, visual features outperformed linguistic ones in
document formalisation, with YOLO models using the SegmOnto controlled vocabulary
surpassing LayoutLM models in multilingual sett4i1n]g. sR[ecently, research has started on OCR
output formalisation for corpus building with a controlled vocabulary and a training dataset
for models 3[
        <xref ref-type="bibr" rid="ref2 ref44">3, 45</xref>
        ]. Lastly, the Layout Analysis Dataset with SegmOnto (LA1D3]aSa)ll[owed
a much finer granularity in the analysis, and a significant improvement of the entire pipeline
for the automatic creation of files encoded in XML-TEI that goes beyond facsimile approach
and closer to reproducing the logical structure of the text.
      </p>
      <p>
        Linguistic Normalisation (LN) has a long history, dating back to the 8107’]s, [but has
developed itself as derived task from Machine Translation (MT) in the beginning of the 2010’s,
usually to improve downstream tasks in the pipeline such as linguistic ann5o4t].atLiNon [
share important similarities with MT, and therefore relies on the same methods, but with a
slightly diferent objective: to “translate” a source into another state of the language, usually
more recent (16th c. Germa→n contemporary German), rather than into another language
(Italian→ German). Resources existed first for Slovene, German, English, Hungarian, Spanish,
Swedish, Portuguese8[], but several studies have recently improved both resou21r]ceasn[d
techniques for historical French, first comparing rule-based, statistical and neural23m]e,thods [
and then alignment-based and neural MT-approa6c]h.es [
Computational scriptology is based on the notion socfripta, coined by Remacle4[
        <xref ref-type="bibr" rid="ref7">8</xref>
        ] and
widely used in Romanistics to to distinguish a spoken language (the dialect) and a written
language (thescripta). The first studies on dialectometry date back from the early 70’s with the
pioneer work of Jean Séguy, who invented the tdeiramlectométrie [
        <xref ref-type="bibr" rid="ref52">53</xref>
        ], on the distance between
dialects in vast corpo4r8a].[ Since then, two main schools, based in Salzbu3r0g]a[nd
Groningen [
        <xref ref-type="bibr" rid="ref42">43</xref>
        ], have advanced research on the topic, but relying mainly on geographical data to
localise dialects. In parallel to these research, Cl. Vachon has changed the approach, switching to
corpus-based research, using historical data to study semi-automatically th5e5s]p,ealnlding [
more recently, J.-B. Camps has shifted the method, using unsupervised stylometry to categorise
medievalscriptae [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ]. Regarding modern French, alternative studies have proposed
alignementbased approaches to compare the historical source and an automatically normalised version to
detect the evolution of spellin24g]so[r to categorise documen2t8s][.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Corpus building</title>
      <p>3.1. Data
For practical reasons, a first corpus of limited size (c. 600 texts) spanning the 17th c. was
produced with our pipeline. The data comes from Gthalelica digital library and contains only
French-language documents. For our experiment, we have selected only plays, which ofer
medium size documents (compared to novels, potentially much longer), and linguistically
homogeneous data (spelling can influenced by the genre, such as legal documents which tend to
use more “archaic” traits and may involve Latin phrases).</p>
      <sec id="sec-3-1">
        <title>3.2. Method</title>
        <p>Our pipeline allows us to extract data, enrich it and store it in a standard for1m).at (cf. fig.
Firstly, we apply a layout analysis model specialised in theatrical data trained for the occasion,
then we use an OCR model prepared for this study which preserves thse (l‹oſ›n).gBased on
the layout analysis we convert ALTO files to TEI files. Only textual data which contains text of
the work (paragraph, speech, verse, etc.), and not linked to the structure of the book (running
title, page number, quire marks, etc.) is extracted and normalised automatically, before to be
reintroduced into the TEI file.</p>
        <p>
          Layout analysis. Based on the resultsTable 1
of Najem-Meyer and Romanello41[] and Training and evaluation data for Layout
Analthe initial evaluation of YOLO region sryeesgsise-ntacroass sitnhgele dadtaosceutms.ent Einach thiemagdeatarseept-.
menter against Kraken’3s6[] as a region
segmenter with YALTAi1[
          <xref ref-type="bibr" rid="ref1">2</xref>
          ], we proposed Train Dev Test
to evaluate the ability of YOLO4v78] [ Images 497 61 62
to detect regions in our 17th c. prinMtainZone-Sp 1738 219 187
corpus. For this purpose, we anno-NumberingZone 384 49 45
tated one random image from each digi-RunningTitleZone 373 43 41
        </p>
        <p>DigitizationArtefactZone 189 23 24
tised version of our corpus, which couldQuireMarksZone 183 29 15
include empty pages (e.g., bookbinding,MainZone-Head 159 28 33
cover) and full pages. This resultedDMraoipnCZaopniet-aSlZp-oCnoentinued 113564 2136 1283
in a corpus of 620 images for trainingG,raphicZone-Decoration 130 16 22
evaluation, and testing. Our final corMM-aaiinnZZoonnee--LEgntry 5889 65 119
pus comprises 32 null pages (without anM-ainZone-P 41 10 12
notations) and a variety of annotatioMMnaasir,ngZinoTneex-tPZ-oCnoen-MtinauneudscriptAddendum 2218 31 72
with a majority of speech-related tagMsarginTextZone-Notes 30 1 0
(pMaariantZeoxnteu:aSPl-,rMealiantZeodneo:bSjPe#cCtosnt(Neiu.ngmu.b,eedr)-, SSMttaaamminppZZZooonnne-eeO-Stthicekrer 112394 331 522
ingZone, RunningTitleZone), a smaller TitlePageZone 8 2 4
number of logical structuring featurGMersaaipnhZiocnZeo-nInecipit 54 01 10
such as scene titleMsa(inZone:Head) and MainZone-Signature 1 0 0
cast listsMa(inZone:Entry), as well as a
few paragraphs and poetic excerpts, mainly found in incipits or prefaces of the books
(MainZone:P), as seen in tab1..
(a) Groundtruth.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Optical character recognition. Since Table 2</title>
        <p>
          YOLO is well integrated within YALTATir,aining and evaluation data for OCR.
which in turn works seamlessly with
Kraken, we decided to use the latter tDoataset Century Language Books Lines
train a new OCR model that includes thTerain/Dev 16 French 7 17817
longs (‹ſ›). Kraken, unlike other OCR sys- Train/Dev 17 French 19 20267
tem, avoids the integration of a strong laTnra-in/Dev 16 Latin 12 10648
guage model which in turn, for our purTT-eesstt 1176 FFrreenncchh 1100
pose, allows for keeping more variationsT.est 18 French 10
This new model, derived from CATMuS
Print 2[
          <xref ref-type="bibr" rid="ref5">6</xref>
          ], uses three datasets for fine-tunin1g9,[
          <xref ref-type="bibr" rid="ref19 ref28">20, 29</xref>
          ] and one evaluation datas2e7t] [
(cf. tab.2). We evaluate on a test set that includes data spanning three centuries (from the
16th to the 18th) and comprises one page from 10 diferent documents for each period.
TEI Document production. Document formalisation follows a logical approach based on
the ALTO output produced by Kraken and YALTAi, rather than a neural one. Each region is
processed in reading order, with regions not matMcahiinnZgone being ignored, except for the
“default” region, which handles orphan lines. The default region is placedfwin&gt;t(o“foar&lt;me
work”) tag, which is typically excluded from our text export processes. Regions marked as
#Continued are logically merged with previous ones. Each line is prepended by albT/E&gt;I &lt;
(line beginning) tag to facilitate back-to-document correction capabilities. Hyphenisation is
resolved by removing hyphen but keeping thleb&lt;/&gt; tag at its pla2ceW.hile machine learning
is employed for initial region detection, the formalisation process itself does not involve any
2The cases that may pose a problem (e.lgu.i-mesme → luimesme, eng. “himself”) represent less than 0.1% of the
corrected hyphenations.
learned behaviour. Metadata are systematically integratedteiinHtehadee&lt;r&gt;, using
information automatically retrieved from the catalogue of the French National Library via the ark
ID.
        </p>
        <p>Linguistic normalisation. All documents are processed via a normaliser previously
traine3d. Only text contained inp&gt;&lt;and &lt;sp&gt; (“speech”) elements are kept for
normalisation, because a specific spelling variation occurring in the running title, for instance, would be
repeated every two pages and potentially alter artificially the result of the scriptometric
analysis. The text is split into sentences (ending by a full stop, an exclamation or a question mark)
or subsentences (ending by a colon or a semicolon), all storesdeign&gt;a(“&lt;arbitrary segment”)
element, with the source text oinri&lt;g&gt; (“original form”) and the automatically normalised
text in r&lt;eg&gt; (“regularization”). The normalised version is evaluated against a dictionary of
modern French to control the quality of the final product.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Experimental Setup and Evaluation</title>
        <p>Layout analysis. We evaluate two possible setups: both use fine-tuning with the original
YOLOv8L models and an input image size of 960 pixels (higher than the default). One setup uses
only the dataset produced in the context of this paper, while the other merges this dataset with
the larger LADaS dataset (5,000 images). We train both setups for 100 epochs with otherwise
default parameters.</p>
        <p>Since our study focuses exclusively on MtahienZone, which contains the primary text and
excludes all paratextual elements (such as decorations, page numbers, and running titles), we
have concentrated our evaluation on this specific zone. Overall, when considering all classes,
we found that integrating our data with the LADaS corpus yields improved results (0.768 vs.
0.8). However, for the most critical clasSspesa(ndSp-continued), the model trained
exclusively with theatrical data produces slightly better outcomes. As previously mentioned, these
are the classes essential for our study.
3https://huggingface.co/rbawden/modern_french_normalis.ation
Text recognition. To fine-tune and Table 4
adapt the CATMuS Print OCR model tCoharacter and word error rates for both models.
the allographic variation of rosu/lnodng Models Characters Errors CER WER
s, we modified the classifier codec (-- No fine-tuning 38394 924 2.41 11.06
resize new mode) and used a standard Fine-tuning 38394 649 1.69 8.34
learning rate of 0.0001, along with a batch
size of 32. This logical approach ensures</p>
        <p>Table 5
the model is fine tuned to the specific ty-Character and word error rates for both models.
pographic variations without relying on
any learned behaviour during the formal%i-errors CER (part) Errors Correct Generated
sation process. We compare this approach 8.78% 0.14% 57 SPACE
to a model without fine-tuning, trained 7.55% 0.13% 49 ' ’
from scratch, with the same architecture 6.62% 0.11% 43 s ſ
(cf. tab.4), revealing the superiority of the 23..2737%% 00..0055%% 2118 ſ– fØ
approach with fine tuning. 2.62% 0.04% 17 ’ '</p>
        <p>Most of the errors are errors related to2.16% 0.04% 14 Ø SPACE
poor segmentation of the text (cf.5t)a,ibn. 22%% 00..0033%% 1133 .1 IØ
which there should be a space that is miss- 1.85% 0.03% 12 ◌́ Ø
ing from the prediction – a classic error for1.69% 0.03% 11 , .</p>
        <p>1.54% 0.03% 10 0 o
historical prints. The prediction errors re1-.54% 0.03% 10 t r
garding two types of apostrophes (curved 1.54% 0.03% 10 ◌̂ Ø
or straight) are of little concern because
they do not afect the result from a linguistic point of view and are due to poor data
preparation that is easily correctable. The confusion between thse arnoduntdhe longs is likely
attributable to the fine-tuning process and the absence of thseinlotnhge base model.
Linguistic normalisation To evaluate
the results of the normalisation, we
compare the prediction of the normaliser with
a dictionary of contemporary French to
obtain a Word Accuracy (WAcc). Results
are satisfactory (cf. fig3.), with a median
above 90%. Texts with a WAcc under
80% are removed to avoid using unreliable
data.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Result dataset</title>
        <p>The final dataset is made of around 80,000
pages for 620 documents. While theFigure 3: Word error rate for the corpus.
number of unit is uneven over the years
(cf. fig. 5a), the accumulated tokens are progressing evenly (c5fb. )fig. .An example of our TEI
encoding is presented in fig4..
&lt;sp&gt;
&lt;ab&gt;
&lt;seg&gt;
&lt;orig&gt;SGANARELLE.&lt;/orig&gt;
&lt;reg&gt;SGANARELLE.&lt;/reg&gt;
&lt;/seg&gt;
&lt;seg&gt;
&lt;orig&gt;Promettez-moy donc, Seigneur Geronimo, de me parler avec toute ſorte de franchiſe.&lt;/orig&gt;
&lt;reg&gt;Promettez-moi donc, Seigneur Geronimo, de me parler avec toute sorte de franchise.&lt;/reg&gt;
&lt;/seg&gt;
&lt;/ab&gt;
&lt;/sp&gt;
&lt;sp&gt;
&lt;ab&gt;
&lt;seg&gt;
&lt;orig&gt;GERONIMO.&lt;/orig&gt;
&lt;reg&gt;GERONIMO.&lt;/reg&gt;
&lt;/seg&gt;
&lt;seg&gt;
&lt;orig&gt;Ie vous le promets.&lt;/orig&gt;
&lt;reg&gt;Je vous le promets.&lt;/reg&gt;
&lt;/seg&gt;
&lt;/ab&gt;
&lt;/sp&gt;</p>
        <p>(b) Accumulated tokens over the years. Year of
(a) Number of bibliographical units per year. Thperinting is used for the date, any document
withbin around 1720 represent printed books fromout a precise date are removed from the plot.
Towithin the 17th century but with unclear or kime n-s are taken from the original OCR documents,
precise printing dates. only from theMainZones.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation of spelling variation</title>
      <sec id="sec-4-1">
        <title>4.1. Method</title>
        <p>
          We use the ABA 4[
          <xref ref-type="bibr" rid="ref5">6</xref>
          ] tool to precisely identify the portions of words which difer between
the original version and the normalised version, and group similar diferences, for example
having the same historical-linguistic origin, or the same type of operations in terms of addition,
deletion or modification of characters. Eaocrhi&lt;g&gt; and &lt;reg&gt; of the corpus is split into words,
the punctuation is removed, and then the original and normalised versions are aligned at the
word level using the Needleman-Wunsc42h][algorithm, using the Levenshtein distan39c]e [
between each pair of words in the samseeg&lt;&gt; in the original and normalised ver4s.ion
        </p>
        <p>Secondly, for each of the aligned word pairs, the original version and the normalised version
are aligned at the character level, still using the Needleman-Wunsch algorithm, but using a
specific substitution matrix to allow not only identical letters to be aligned, but also letters
considered close in (pre)classical French and contemporary French (presence/absence of diacritic,
ligatures…). For example, while identical letters benefit from a substitution score of 4, letters
difering only in accent or cedilla benefit from a score of 2, asſ›daon‹d ‹s› or ‹s› and ‹ß› for
example. Other pairs of letters benefit from a score of 1, such as ‹u› and ‹v›, ‹s› and ‹z› or even
‹n› and ‹m ›. Conversely, a score of -1 is assigned to pairs of distinct letters not subject to such
exceptions, as well as to the deletion or insertion of a character.</p>
        <p>This execution of the
NeedlemanWunsch algorithm to obtaiTnable 6: Prefix similarity matrix for the original and
character-level alignment is tniholer-mparleisveiodusvebrsoixononofth‹eApoopſttrime›a.l paTthhe taorrcoawlcsuliantdeictahtee
lustrated in the matrix in t6a,b.similarity between two prefixes, one from the word on
where each number representsthe firstOrnowt,hitsheopottihmearl frpoamth,thgerewenordin dinicatthees feirqsutalciotyl-,
the similarity score of the besruetmdni.ndicates substitution, and blue indicates deletion.
pttaaaurhollnpiieeggfixtcnnbaooermmomsrteefteohnnwia‹ttfsArl.oipbgiFmoofnnooxdſmrutt.iencroenaedIxt›ttoaibmibnastbpegneatlptdiewrnww,ee‹techetAoeeinhcpndiohôs‹eAbtdſtb›rtpbeaebhoos›iyexnttreôpA ↓↓↓↓↓A↘-320141 ↓↓↓↓p→↘754638 ↓↓↓o↘→→8791270 ſ↓↓↓→→→768169 t↓↓↘→→→111058213 ↓r↘→→→→11-4716172 e↘→→→→→ 2-36112161
and ‹Apô›, we must consider the
best alignment between ‹Apo› and ‹Apô› (which has a score of 10) then make an inserTt,ion of
which has a score of -1, which provides a total score of 9. If we had preferred to first consider
the best alignment between ‹Aſ›paond ‹Ap›, which has a score of 6, then deleteôt,hwehich
has a score of -1, we would have obtained an alignment with a score of 5, therefore lower than
optimal. In case of insertion or deletion during this alignment step, we use the ¤ character
in order to obtain two words of the same length in both the original and normalised version.
Thus, at the end of this second alignment step, the AwpoorTdtre in the original version is
matched withapô¤tre in a normalised version to obtain character-by-character alignment.</p>
        <p>Finally, for each word in the corpus, its original and normalised versions are analysed,
char4Some subtleties are brought to this adjustment, suetchanads&amp; which are considered equivalent.
acter by character, to detect, in the case of diferent characters at the same position, the
normalisation rule that applies, or to signal that no existing rule was identified when
appropriate. 72 rules were defined based on the bibliography and the diferences observed in the gold
FreEMnorm parallel corpu2s1[]. For example, the rulReamist letter is detected if an ‹i›, a ‹j›,
an ‹u› or a ‹v› is present in the associated original word respectively to a ‹j›, an ‹i›, a ‹v› or an
‹u› in the normalised version.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Results</title>
        <p>Based on the alignments obtained using the
Needleman-Wunsch algorithm and the
detections of the 72 rules mentioned earlier, our
analysis reveals four distinctive patterns of
historical spelling changes. The principle underlying
this analysis is straightforward: if a
normalisation rule is detected less frequently, it indicates
that the historical spelling it targets is
becoming less prevalent in the corpus. To examine its
evolution throughout the century, we
normalize the total number of rule applications tFoigitusre 6: Disappearance of ‹gn›.
percentage within each text. For instance, the etymological spelling ‹gn›, founcdogi-n form
noitre (&lt;lat. cognoscere), is less and less replace by ‹nn› (tocdonanyaître, eng. “to know”),
signifying the slow disappearance of this spelling (c6f)..fig.</p>
        <p>(a) Substitution of ‹es› by ‹é›.</p>
        <p>(b) Substitution of ‹as› by ‹â›.</p>
        <p>Pattern A: constant rate. Using ABA, it is possible to detect more complex traits of
historical graphic systems than the specific use of a single letter (e.g. ‹u› vs ‹v› as a vowel) or a group
of letters (‹gn› vs ‹nn›), such as the presence of a diacritical letter to change the sound-value of
the letter to which it is added (e.g. vowel + ‹s›). In historical French, the phoneme [e] is thus
regularly noted with the grapheme ‹es› where today we useesta‹ét›v(s état, eng. “state”), and
the phoneme [ä] is noted ‹as› where we now find ‹â›pa(sturage vs pâturage, eng “pasture”). If
counting the presence of ‹v› followed by consonant (‹vne› = [yn]) to identify the historical use
of ‹v› is enough, it is impossible to count the occurrences of ‹es› to measure the presence of a
diacriticasl(‹esponge›→ ‹éponge›, eng. “sponge”, but ‹espagno→l›‹espagnol› and not
‹épagnol›, eng. “spanish”): the transition from a complex grapheme (such as a digraph) to a simple
grapheme requires an alignment at the character level of the original text and its normalised
version, and then the deduction of the spelling change from the diference between the two.
(a) Accumulation of occurrences of diferent (b) Theoretical progression of two similar
varispellings: two similaers(→é, as→â) and one ants over time, which start at diferent times,
diferent (ct→t). Data are scaled to base 100 but progress at the same speed, according to
to be comparable. the constant rate hypothesis.</p>
        <p>In our corpus, we detect a clear decrease in the use of complex graphemes with a diacritical
s, whether the latter is combined with ‹e› (cf.7afig). or with ‹a› (cf. fig. 7b). Interestingly,
the propagation of these two new spellings (vowel+accent) does occur at a very similar speed
(cf. fig. 8a), recalling Kroch’s constant rate hypothesis (c8bf.)5fig,. of which researchers have
already found traces in syntac5t9i]ca[nd phonological18[] change.
(a) Apparition of the contemporary use of ramist(b) Disappearance of the etymological
combinaletters. tion ‹ct›.
5“When one grammatical option replaces another with which it is in competition across a set of linguistic contexts,
the rate of replacement, properly measured, is the same in all of 3t8h]em.” [
Pattern B: abrupt change. On the basis of such observations, it is however possible to go
further and date the moment when a break occurs in the scribal practice, to date the moment
when the spelling changes. To do so, we can use binary segmentation58(,B4S]), [an algorithm
using a forward stepwise method, to identify change-point detection. This method has already
been used in diachronic linguistic to study the sudden introduction of new lexi3c7a]l. items [</p>
        <p>
          One of the main discoveries of our study is the extremely abrupt nature of certain changes,
which take place at very high speed, such as the disappearance between 1668 and 1672 of
Ramist letters (cf. fig9.a), as proposed by Christophe Plantin in the 16t1h1c].a[nd defended
by Pierre Corneille in his forewoarudlecteur of 1663 [
          <xref ref-type="bibr" rid="ref13">14</xref>
          ]. A similar phenomenon, although
slightly less abrupt, exists for the disappearance of the etymological ‹c› followed by ‹t› (e.g.
‹faict›&lt;factum, todayfait, eng. “fact”) at the end of the 1630s (cf. 9figb.).
        </p>
        <p>(a) Slow decrease of the losn(g‹ſ›).</p>
        <p>(b) Increase of the acute accent.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Pattern C: correlation. L. Biedermann</title>
        <p>Pasques proposed as one of the parameters for
spelling change the type case available: “the
typographical use of the ligature has slowed
down, in our opinion, the regular replacement
of silents by an accent”7[, p. 92]. It is indeed
faster to compose the woredſtoit with the
ligature (e+ſt+o+i+t=4 characters) than without
(e+ſ|s+t+o+i+t=5 characters). One could argue
that switching to the accented letter Faiglusore 11: Correlation ‹ſ›/acute accent.
requires only four characters (é+t+o+i+t=4 characters), but if the ligatures are present in
number in the printer’s type case, the accented characters are less so.</p>
        <p>Our working hypothesis is as follows: as ligatures are largely composed osf a(‹ſl›)o,ng
we should obtain a correlation between the use osf (tchf.i sfig. 10a) and the acute accent
(cf. fig. 10b). We evaluate the correlation between the evolution of the two phenomena over
time, and obtain a Pearson product-moment correlation coefÏcient of 0.365 pw-ivtahluae of
4.88e-20, which indicate a good correlation (cf1. 1fig).
(a) Increase of the diaeresis.</p>
        <p>(b) Increase of the confusion ‹en|m›/‹an|m›.</p>
        <p>Pattern D: innovation. Finally, it is important to note that, in this slow movement of
standardisation that we are drawing, innovations also appear. These innovations concern a lot
diacritics, some of which are exploding in number like the diaeresis (1c0fa.)fig:.scriptors tend
to add them more and more on one of the two hiatus vowel, especially with the sequence ‹ue›
(louër orloüer, todaylouer, eng. “to rent/to praise”). We also note a great hesitation regarding
the notation of nasal vowels (cf.1fig2b.), especially [ã], for which we can use ‹en|m› or ‹an|m›
such asaventure vs avanture (todayaventure, eng. “adventure”).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and further work</title>
      <p>The spelling of the 17th c. is changing throughout the century, and at the beginning of the 18th
century the standardisation process is very advanced (etymological letters, use of diacritical
letters, historical use of ‹u› and ‹i›, etc.), as studies on other languages, such a31s]Polish [
or English 5[], have been able to demonstrate.We still observe, however, a certain instability,
which concerns more minor hesitations than anything else (notation of nasal vowels, hiatus
vowel, etc): although the standardisation process is advanced, it is not yet finished.</p>
      <p>
        Among all the changesTable 7
in spelling, all those obM- ain spelling change and their dating using change-point detection.
served seem to follow the Rule Example Change-point
traditional shape of the
scurve, and no “anomalies” ‹‹cctq›u→e›‹→t›‹c› eaxvpecloqiucet→→aexvepcloit 11663328
have yet been found as it tilde → vowel hõme→homme 1637
have been the case else- ‹gn› → ‹nn› incognu→inconnu 1654
where [
        <xref ref-type="bibr" rid="ref30">31</xref>
        ]: on the con- ‹‹éass››→→‹‹éâ›› epsatsalte→→éptâalte 11666609
trary, we even think we Distinction of ramist letters vniuers→univers 1670
fond new evidences sup- ‹eu›→‹u› asseurance→assurance 1675
porting Kroch’s constantSSuupppprreessssiioonn ooff ceatyllmigoralopghiiccalleltettetrer vnroapyc→e→vrnaoice 11668838
rate hypothesis.
      </p>
      <p>The idea of a slow change, which spreads over a long per5i]o,sde[ems to be confirmed by our
analyses (cf. tab7.). However, the velocity of change varies greatly from one phenomenon to
another, with sometimes slow shifts over decades, or sometimes abrupt ruptures whose cause
is not entirely clear, and which would be interesting to discover.</p>
      <p>As for the reasons for the change, a lot of work still needs to be done, particularly in trying
to find features that could predict the cha3n2g].e O[ne of them, the identity of the printers,
would be interesting to evaluate, unfortunately the data is not always available, particularly
for the 18th century, which will pose a problem for the future of this study. Nevertheless, some
indications suggest that it would be important to review the hypothesis that sees printing as a
vector of change5][: the limitations imposed by the type case of printers could for instance be
a hindrance to change.</p>
      <p>A more precise modelling of these changes is therefore on the agenda for our future research.
Whether it concerns the identification of possible reasons for these changes, their more precise
dating (in particular by integrating confidence intervals), or the addition of new data for the
18th century. The improvement of all data extraction and enrichment tools has already begun,
and should thus allow the creation of an even larger and more precise corpus.</p>
    </sec>
    <sec id="sec-6">
      <title>Data and code</title>
      <p>All the data and code is available on our GitHubhrtetppos://github.com/DEFI-COLaF/Theat
reLFSV2.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>Merci (dans l’ordre alphabétique) à Jean Barré, Alexandre Bartz, Rachel Bawden, Philippe
Gambette et Benoît Sagot pour leur aide. À nos relecteur·trices aussi pour leurs excellentes
remarques.</p>
    </sec>
    <sec id="sec-8">
      <title>Funding References</title>
      <p>This paper has been funded by the DEFI Inria COLCaoFrpus et Outils pour les Langues de France
and the FNS-Spark projecNt°220833.</p>
      <p>[1] Académie françaiseC.ahiers de remarques sur l’orthographe françoise pour estre
examinez par chacun de Messieurs de l’Academie, avec des observations de Bossuet, Pellisson, etc.
Charles Joseph Marty-Laveaux. Paris: Jules Gay, 1863. uhrtlt:ps://books.google.ch/boo
ks?id=u5Y5AQAAIAAJ.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Amatuzzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ayres-Bennett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gerstenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schøsler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Skupien-Dekens</surname>
          </string-name>
          .
          <article-title>“Changement linguistique et périodisation du français (pré)classique: deux études de cas à partir des corpus du RCFC”</article-title>
          .
          <source>IJno:urnal of French Language Studies 30.3</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>301</fpage>
          -
          <lpage>326</lpage>
          . doi:
          <volume>10</volume>
          .1017/s0959269520000058.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Baddeley.L'</surname>
          </string-name>
          <article-title>Ortographie française au temps de la Réforme</article-title>
          .
          <source>Genève: Droz</source>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          . “Estimating Multiple Breaks One at a Time”E.cIonn:
          <source>ometric Theory 13.3</source>
          (
          <issue>1997</issue>
          ), pp.
          <fpage>315</fpage>
          -
          <lpage>352</lpage>
          . doi:
          <volume>10</volume>
          .1017/s0266466600005831.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Basu</surname>
          </string-name>
          . “
          <article-title>Early Modern Studies: the Digital Turn”</article-title>
          . In: ed. by L.
          <string-name>
            <surname>Estill</surname>
            ,
            <given-names>D. K.</given-names>
          </string-name>
          <string-name>
            <surname>Jakacki</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ullyot</surname>
          </string-name>
          . Toronto: Iter Press,
          <year>2016</year>
          . Chap. “
          <article-title>Ill shapen sounds, and false orthography”': A Computational Approach to Early English Orthographic Variation</article-title>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Bawden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Poinhos</surname>
          </string-name>
          , E. Kogkitsidou,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gambette</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sagot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          . “Automatic Normalisation of Early Modern French”L.
          <string-name>
            <surname>RInE:C 2022 - 13th Language Resources</surname>
          </string-name>
          and Evaluation Conference.
          <source>European Language Resources Association</source>
          . Marseille, France,
          <year>2022</year>
          , pp.
          <fpage>3354</fpage>
          -
          <lpage>3366</lpage>
          . url: https://inria.hal.
          <source>science/hal-03540</source>
          .
          <fpage>226</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Biedermann-PasquesL.es Grands</surname>
          </string-name>
          <article-title>Courants orthographiques au XVIIe siècle et la formation de l'orthographe moderne, Impacts matériels, interférences phoniques</article-title>
          ,
          <source>théories et pratiques</source>
          (
          <volume>1606</volume>
          -
          <fpage>1736</fpage>
          ). Tübingen: Max Niemeyer Verlag,
          <year>1992</year>
          . do1i0:.
          <volume>1515</volume>
          /97831109385 93.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bollmann</surname>
          </string-name>
          . “
          <article-title>Normalization of Historical Texts with Neural Network Models”</article-title>
          .
          <source>PhD thesis</source>
          . Bochum:
          <string-name>
            <surname>Ruhr-Universität Bochum</surname>
          </string-name>
          ,
          <year>2018</year>
          . urhlt:tps://www.linguistics.rub.de/fo rschung/arbeitsberichte/22..pdf
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.-B.</given-names>
            <surname>Camps</surname>
          </string-name>
          . “
          <article-title>Manuscripts in Time and Space: Experiments in Scriptometrics on an Old French Corpus”</article-title>
          .
          <source>InP:roceedings of the Second Workshop on Corpus-Based Research in the Humanities CRH-2</source>
          . Ed. by A. U. Frank,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ivanovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mambrini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Passarotti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Sporleder</surname>
          </string-name>
          . Vienna, Austria,
          <year>2018</year>
          , pp.
          <fpage>55</fpage>
          -
          <lpage>64</lpage>
          . uhrtl:tps://hal.science/hal-01695 8.
          <fpage>99</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>CatachH</surname>
          </string-name>
          .istoire de l'orthographe française. Paris: Honoré Champion,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Catach</surname>
          </string-name>
          and
          <string-name>
            <surname>J. Golfand. “</surname>
          </string-name>
          <article-title>L'orthographe plantiniennDe”e</article-title>
          .
          <source>GInu:lden Passer</source>
          <volume>50</volume>
          (
          <year>1973</year>
          ), pp.
          <fpage>19</fpage>
          -
          <lpage>69</lpage>
          . url: https://www.dbnl.org/tekst/%
          <source>5C%5Fgul005197301%5C%5F01/%5C%5Fg ul005197301%5C%5F01%5C%5F0003.php.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Clérice</surname>
          </string-name>
          . “
          <article-title>You Actually Look Twice At it (YALTAi): using an object detection approach instead of region segmentation within the Kraken enginJeo”u</article-title>
          .rInna:l of Data Mining &amp; Digital
          <string-name>
            <surname>Humanities</surname>
          </string-name>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .46298/jdmdh.9806.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Clérice</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Janès</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Scheithauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bénière</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Sagot</surname>
          </string-name>
          . “
          <article-title>Layout Analysis Dataset with SegmOnto”</article-title>
          .
          <source>IDnH:2024 - Annual conference of the Alliance of Digital Humanities Organizations. Alliance of Digital Humanities Organizations (ADHO)</source>
          . Washington,
          <string-name>
            <given-names>D.C.</given-names>
            ,
            <surname>United</surname>
          </string-name>
          <string-name>
            <surname>States</surname>
          </string-name>
          ,
          <year>2024</year>
          . urhl:ttps://inria.hal.
          <source>science/hal-04513</source>
          .
          <fpage>725</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>CorneilleL.e Théâtre de P. Corneille</surname>
          </string-name>
          . Paris: G. de Luyne,
          <volume>1663</volume>
          . urlh:ttps://gallica.bnf.f r/ark:/12148/bpt6k71442.p
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dees</surname>
          </string-name>
          . “
          <article-title>Dialectes et scriptae à l'époque de l'ancien françaRise”v</article-title>
          .uIen:de
          <source>Linguistique Romane</source>
          <volume>49</volume>
          (
          <year>1985</year>
          ), pp.
          <fpage>87</fpage>
          -
          <lpage>117</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Duval</surname>
          </string-name>
          . “Les éditions de textes du XVIIe siècle”M.Iann: uel de la philologie de l'édition. Ed. by
          <string-name>
            <given-names>D.</given-names>
            <surname>Trotter</surname>
          </string-name>
          . Berlin, Boston: De Gruyter,
          <year>2015</year>
          , pp.
          <fpage>369</fpage>
          -
          <lpage>394</lpage>
          .
          <year>1d0</year>
          .
          <year>o1i5</year>
          :
          <volume>15</volume>
          /978311030 2608-
          <fpage>017</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Fix</surname>
          </string-name>
          . “
          <article-title>Automatische Normalisierung - Vorarbeit zur Lemmatisierung eines diplomatischen altisländischen Textes”</article-title>
          .
          <source>TInei:l 3 Beiträge zum dritten Symposion Tübingen</source>
          <volume>17</volume>
          . -
          <fpage>19</fpage>
          .
          <year>Februar 1977</year>
          . Berlin/Boston: Max Niemeyer Verlag,
          <year>1980</year>
          , pp.
          <fpage>92</fpage>
          -
          <lpage>100</lpage>
          .
          <source>ddooi:i:10.1515/9 783111438788</source>
          .92.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fruehwald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gress-Wright</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Wallenberg</surname>
          </string-name>
          . “
          <article-title>Phonological Rule Change: The Constant Rate Efect”</article-title>
          .
          <source>InP:roceedings of the 40th Annual Meeting of the North East Linguistic Society</source>
          . Cambridge, MA: GLSA Publications,
          <year>2013</year>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>230</lpage>
          . urhlt:tps://www.resea rch.ed.ac.uk/files/14416788/Fruewald%5C
          <source>%5FGress%5C%5FWright%5C%5FWallenberg %5C%5FPhonological%5C%5FRule%5C%5FChange</source>
          .p.df
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay.</surname>
          </string-name>
          Fondue-fr-print-
          <volume>16</volume>
          .
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.11526150.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay.</surname>
          </string-name>
          Fondue-fr-print-
          <volume>17</volume>
          .
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.11526040.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay.</surname>
          </string-name>
          FreEM-corpora/FreEMnorm:
          <article-title>FreEM norm Parallel corpus</article-title>
          .
          <source>Version 1.0.0</source>
          .
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.5865428.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          . “
          <article-title>Pourquoi moderniser l'orthographe? Principes d'ecdotique et littérature du XVIIe siècle”</article-title>
          .
          <source>In:Vox Romanica 73.1</source>
          (
          <issue>2014</issue>
          ), pp.
          <fpage>27</fpage>
          -
          <lpage>42</lpage>
          . doi:
          <volume>99</volume>
          .125005/vox201410027.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Barrault</surname>
          </string-name>
          . “
          <article-title>Traduction automatique pour la normalisation du français du XVIIe siècle”</article-title>
          .
          <article-title>InA:ctes de la 6e conférence conjointe Journées d'Études sur la Parole (JEP, 33e édition), Traitement Automatique des Langues Naturelles (TALN, 27e édition), Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL, 22e édition)</article-title>
          . Volume 2
          <article-title>: Traitement Automatique des Langues Naturelles</article-title>
          . Nancy, France,
          <year>2020</year>
          , pp.
          <fpage>213</fpage>
          -
          <lpage>222</lpage>
          . url:https://aclanthology.org/
          <year>2020</year>
          .jeptalnrecital-.
          <source>taln.20</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bawden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gambette</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Poinhos</surname>
          </string-name>
          , E. Kogkitsidou, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Sagot</surname>
          </string-name>
          . “
          <article-title>Le changement linguistique au XVIIe s. : nouvelles approches scriptométriqueCsM”</article-title>
          .LIFn:
          <fpage>2022</fpage>
          - 8e
          <string-name>
            <surname>Congrès Mondial de Linguistique Française</surname>
          </string-name>
          . Vol.
          <volume>138</volume>
          .
          <article-title>SHS Web of conferences</article-title>
          . Orléans, France: EDP Sciences,
          <year>2022</year>
          , pp.
          <year>02006</year>
          .
          <volume>1</volume>
          -
          <fpage>14</fpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .1051/shsconf/202213802006.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-B. Camps</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Pinche</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Jahan</surname>
          </string-name>
          . “
          <article-title>SegmOnto: common vocabulary and practices for analysing the layout of manuscripts (and mor1est)</article-title>
          <source>”I.nItner:national Workshop on Computational Paleography (IWCPICDAR</source>
          <year>2021</year>
          ). Lausanne, Switzerland,
          <year>2021</year>
          . url: https://hal.science/hal-03336 5.
          <fpage>28</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Clérice</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jacsont</surname>
          </string-name>
          , E. Leblanc,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jeannot-Tirole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Solfrini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dolto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Goy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Luján</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zaglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perregaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Janès</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sagot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bawden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Nédey</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Chagué</surname>
          </string-name>
          . “
          <article-title>Reconnaissance des écritures dans les imprimésH”</article-title>
          .uImn:anistica
          <year>2024</year>
          .
          <article-title>Association francophone des humanités numériques</article-title>
          . Meknès, Morocco, 2024h.tutrpl: s://hal.science/hal-0455745.
          <fpage>7</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Clérice</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>JanèFsO.NDUE-MLT-PRINT-</surname>
          </string-name>
          TEST-longS.
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .5281/z enodo.
          <volume>11526316</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gambette</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bawden</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Sagot</surname>
          </string-name>
          . “
          <article-title>Ancien ou moderne ? Pistes computationnelles pour l'analyse graphématique des textes écrits au XVIIe siLèicnlxe”8</article-title>
          .5In: (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .4000/linx.9346.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jeannot-Tirole</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>GoFoyn.</surname>
          </string-name>
          due-la
          <source>-print-16</source>
          .
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.1
          <fpage>1526160</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>H.</given-names>
            <surname>Goebl</surname>
          </string-name>
          . “Dialektometrie”. IQnu:antitative Linguistik/Quantitative Linguistics.
          <article-title>Ein internationales Handbuch/An International Handbook</article-title>
          . Berlin, New York: De Gruyter Mouton,
          <year>2005</year>
          , pp.
          <fpage>498</fpage>
          -
          <lpage>531</lpage>
          . doi:
          <volume>10</volume>
          .1515/9783110155785.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Górski</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Eder</surname>
          </string-name>
          . “
          <article-title>Modelling the Dynamics of Language Change: Logistic Regression, Piotrowski's Law, and a Handful of Examples in PolishJ”o</article-title>
          .
          <source>uInrn: al of Quantitative Linguistics 30.1</source>
          (
          <issue>2023</issue>
          ), pp.
          <fpage>125</fpage>
          -
          <lpage>151</lpage>
          . doi:
          <volume>10</volume>
          .1080/09296174.
          <year>2022</year>
          .
          <volume>2151208</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>L.</given-names>
            <surname>Hou</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Smith</surname>
          </string-name>
          . “
          <article-title>Modeling the Decline in English PassivizationPr”o.cInee:dings of the Society for Computation in Linguistics (SCiL) 2018</article-title>
          . Ed. by
          <string-name>
            <given-names>G.</given-names>
            <surname>Jarosz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J.</given-names>
            <surname>Pater</surname>
          </string-name>
          .
          <year>2018</year>
          , pp.
          <fpage>34</fpage>
          -
          <lpage>43</lpage>
          . doi:
          <volume>10</volume>
          .7275/r5zc812c.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>J.</given-names>
            <surname>Janès</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pinche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jahan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          . “
          <article-title>Towards automatic TEI encoding via layout analysis”</article-title>
          .
          <source>InF:antastic future 21, 3rd International Conference on Artificial Intelligence for Librairies</source>
          , Archives and Museums.
          <source>AI for Libraries, Archives, and Museums (AI4LAM)</source>
          . Paris, France,
          <year>2021</year>
          . urlh:ttps://hal.science/hal-03527 2.
          <fpage>87</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>M.</given-names>
            <surname>Khemakhem</surname>
          </string-name>
          . “
          <article-title>Standard-based Lexical Models for Automatically Structured Dictionaries”</article-title>
          .
          <source>PhD thesis</source>
          . Paris: Université de Paris,
          <year>2020</year>
          . uhrtlt:ps://theses.hal.
          <source>science/tel-0 3274454.</source>
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>M.</given-names>
            <surname>Khemakhem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Foppiano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          . “
          <article-title>Automatic Extraction of TEI Structures in Digitized Lexical Resources using Conditional Random FieldEsl”e</article-title>
          .cItnro:nic lexicography, eLex
          <year>2017</year>
          . Leiden, The Netherlands,
          <year>2017</year>
          . urhlt:tps://hal.science/hal-01508 8.
          <fpage>68</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kiessling</surname>
          </string-name>
          . “
          <article-title>Kraken - an Universal Text Recognizer for the HumanitieDsi”g</article-title>
          .
          <source>iItnal: Humanities Conference 2019 - DH2019. Utrecht, The Netherlands: Alliance of Digital Humanities Organizations (ADHO)</source>
          ,
          <year>2019</year>
          . do1i0:.
          <volume>34894</volume>
          /z9g2ex.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>C.</given-names>
            <surname>Klaussner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Vogel</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          . “
          <source>Detecting Linguistic Change Based on Word Co-occurrence Patterns”P.Irnoc:eedings of the 4th International Workshop on Computational History. Singapore</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>14</fpage>
          -
          <lpage>21</lpage>
          . url:https://ceur-ws.
          <source>org/</source>
          Vol-1992
          <source>/paper %5C%5F4.pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Kroch</surname>
          </string-name>
          . “
          <article-title>Reflexes of grammar in patterns of language changeL”</article-title>
          .
          <source>aInng: uage Variation and Change 1</source>
          .3 (
          <issue>1989</issue>
          ), pp.
          <fpage>199</fpage>
          -
          <lpage>244</lpage>
          . doi:
          <volume>10</volume>
          .1017/s0954394500000168.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>V.</given-names>
            <surname>Levenshtein</surname>
          </string-name>
          . “
          <article-title>Binary codes capable of correcting deletions, insertions, and reversals”</article-title>
          .
          <source>In: Soviet physics doklady 10.8</source>
          (
          <issue>1966</issue>
          ), pp.
          <fpage>707</fpage>
          -
          <lpage>710</lpage>
          . url: https://www.mathnet.ru/eng/da n31411.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>L.</given-names>
            <surname>Muellner</surname>
          </string-name>
          . “
          <article-title>Digital Classical Philology Ancient Greek and Latin in the Digital Revolution”</article-title>
          . In: Berlin/Boston: De Gruyter Saur,
          <year>2019</year>
          . Chap.
          <source>The Free First Thousand Years of Greek</source>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>17</lpage>
          . doi:
          <volume>10</volume>
          .1515/
          <fpage>9783110599572</fpage>
          -
          <lpage>002</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>S.</given-names>
            <surname>Najem-Meyer</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Romanello</surname>
          </string-name>
          . “
          <article-title>Page Layout Analysis of Text-heavy Historical Documents: a Comparison of Textual and Visual ApproachesP”r</article-title>
          .
          <source>oIcne:edings of the Computational Humanities Research Conference</source>
          <year>2022</year>
          . Antwerp, Belgium,
          <year>2022</year>
          , pp.
          <fpage>36</fpage>
          -
          <lpage>54</lpage>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2212.13924.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Needleman</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Wunsch</surname>
          </string-name>
          . “
          <article-title>A general method applicable to the search for similarities in the amino acid sequence of two proteinJso”u</article-title>
          .
          <source>rIna:l of Molecular Biology 48.3</source>
          (
          <issue>1970</issue>
          ), pp.
          <fpage>443</fpage>
          -
          <lpage>453</lpage>
          . doi:
          <volume>10</volume>
          .1016/
          <fpage>0022</fpage>
          -
          <lpage>2836</lpage>
          (
          <issue>70</issue>
          )
          <fpage>90057</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>J.</given-names>
            <surname>Nerbonne</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Heeringa</surname>
          </string-name>
          . “
          <fpage>31</fpage>
          . Measuring dialect diferences”
          <source>T.hIeno:ries and Methods: An International Handbook of Linguistic Variation</source>
          . Vol.
          <volume>1</volume>
          .
          <string-name>
            <surname>De Gruyter Mouton</surname>
          </string-name>
          ,
          <year>2010</year>
          , pp.
          <fpage>550</fpage>
          -
          <lpage>567</lpage>
          . doi:
          <volume>10</volume>
          .1515/9783110220278.550.
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>Netherlands</given-names>
            <surname>Historical Data Archive</surname>
          </string-name>
          ,
          <source>Nijmegen Institute for Cognition &amp; Information. Optical Character Recognition in the Historical Discipline: Proceedings of an International Workshop. St. Katharinen: Halbgraue Reihe zur Historischen Fachinformatik</source>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pinche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Christensen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Gabay</surname>
          </string-name>
          . “
          <article-title>Between automatic and manual encoding”</article-title>
          . In:
          <article-title>TEI 2022 conference : Text as data</article-title>
          . Newcastle, United Kingdom,
          <year>2022</year>
          . do1i0:.5281/zenod o.
          <volume>7092214</volume>
          . url: https://hal.science/hal-03780 3.
          <fpage>02</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [46]
          <string-name>
            <surname>J. Poinhos.</surname>
          </string-name>
          <article-title>ABA (Alignment-Based Approach)</article-title>
          .
          <source>Version 1</source>
          .
          <year>2020</year>
          . url:https://github.com/j ohnseazer/ab.a
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>D.</given-names>
            <surname>Reis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kupec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Daoudi</surname>
          </string-name>
          . “
          <article-title>Real-Time Flying Object Detection with YOLOv8”</article-title>
          . In: CoRR abs/2305.09972 (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .48550/arxiv.2305.09972.
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>L.</given-names>
            <surname>Remacle.Le Problème de l</surname>
          </string-name>
          <article-title>'ancien wallon</article-title>
          .
          <source>Liège: Presses universitaires de Liège</source>
          ,
          <year>1948</year>
          . url: http://books.openedition.org/pulg/.338
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>B.</given-names>
            <surname>Robertson</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Boschetti</surname>
          </string-name>
          . “
          <article-title>Large-Scale Optical Character Recognition of Ancient Greek”</article-title>
          .
          <source>In:Mouseion: Journal of the Classical Association of Canada</source>
          <volume>58</volume>
          (3
          <year>2017</year>
          ), pp.
          <fpage>341</fpage>
          -
          <lpage>359</lpage>
          . url: https://muse.jhu.edu/article/6791.81
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Lopez</surname>
          </string-name>
          . “
          <article-title>GROBID - Information Extraction from Scientific Publications”</article-title>
          .
          <source>In: ERCIM News. Scientific Data Sharing and Re-use 100</source>
          (
          <year>2015</year>
          ). urhlt:tps://inria.hal.sci ence/hal-01673305.
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>U.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Weitz</surname>
          </string-name>
          . “
          <article-title>Combining OCR Outputs for Logical Document Structure Markup</article-title>
          .
          <article-title>Technical Background to the ACL 2012 Contributed TasPkro”</article-title>
          .cIene:dings of the ACL-2012 Special Workshop on Rediscovering 50 Years of Discoveries. Ed. by
          <string-name>
            <given-names>R. E. Banchs. Jeju</given-names>
            <surname>Island</surname>
          </string-name>
          , Korea: Association for Computational Linguistics,
          <year>2012</year>
          , pp.
          <fpage>104</fpage>
          -
          <lpage>109</lpage>
          . url: https://aclanthology.org/W12-3.
          <fpage>212</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>H.</given-names>
            <surname>Scheithauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chagué</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          . “
          <article-title>Which TEI representation for the output of automatic transcriptions and their metadata? An illustrated proposition”</article-title>
          .
          <year>2022</year>
          . url: https://inria.hal.
          <source>science/hal-04001</source>
          .
          <fpage>303</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>J.</given-names>
            <surname>Séguy</surname>
          </string-name>
          . “
          <article-title>La dialectométrie dans l'Atlas linguistique de la GascogRneev”u</article-title>
          .eInd:e linguistique romane
          <volume>37</volume>
          (
          <year>1973</year>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>E.</given-names>
            <surname>Tjong Kim Sang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bollmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Boschker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Casacuberta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dietz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dipper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Domingo</surname>
          </string-name>
          , R. van der Goot, M. van
          <string-name>
            <surname>Koppen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ljubešić</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Östling</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Petran</surname>
            , E. Pettersson,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Scherrer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Schraagen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Sevens</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Tiedemann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Vanallemeersch</surname>
            , and
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Zervanou</surname>
          </string-name>
          . “
          <article-title>The CLIN27 Shared Task: Translating Historical Text to Contemporary Language for Improving Automatic Linguistic AnnotatioCno”</article-title>
          .
          <source>mInp:utational Linguistics in the Netherlands Journal</source>
          <volume>7</volume>
          (
          <year>2017</year>
          ), pp.
          <fpage>53</fpage>
          -
          <lpage>64</lpage>
          . url: https://clinjournal.org/clinj/article/vie w/68/61.
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Vachon</surname>
          </string-name>
          .
          <article-title>Le Changement linguistique au XVIe siècle: une étude basée sur des textes littéraires français</article-title>
          .
          <source>Strasbourg: ELiPhi</source>
          , Éditions de linguistique et de philologie,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          [56]
          <string-name>
            <surname>C. F. d</surname>
          </string-name>
          . VaugelasR.
          <article-title>emarques sur la langue françoise</article-title>
          .
          <source>Geneva: Droz</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          [57]
          <string-name>
            <surname>C. F. d</surname>
          </string-name>
          . VaugelasR.
          <article-title>emarques sur la langue françoise, utiles à ceux qui veulent bien parler et bien escrire</article-title>
          . Paris: Vve J.
          <string-name>
            <surname>Camusat</surname>
          </string-name>
          et P. Le Petit,
          <volume>1647</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          [58]
          <string-name>
            <given-names>L.</given-names>
            <surname>Vostrikova</surname>
          </string-name>
          . “
          <article-title>Detection of the disorder in multidimensional random-processes”</article-title>
          .
          <source>In: Doklady Akademii Nauk SSSR 259.2</source>
          (
          <issue>1981</issue>
          ), pp.
          <fpage>270</fpage>
          -
          <lpage>274</lpage>
          . url: http://mi.mathnet.ru/d an44582.
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          [59]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zimmermann</surname>
          </string-name>
          . “
          <article-title>An improved test of the constant rate hypothesis: late Modern American English possessive have”</article-title>
          .
          <source>InC:orpus Linguistics and Linguistic Theory</source>
          <volume>19</volume>
          .3 (
          <issue>2023</issue>
          ), pp.
          <fpage>323</fpage>
          -
          <lpage>352</lpage>
          . doi: doi:10.1515/cllt-2021
          <source>-003.8</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>