<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Paris, France
∗Corresponding author.
£ wouter.haverals@princeton.e(Wdu. Haverals)m;ike.kestemont@uantwerpen.b(Me. Kestemont)
ç https://whaverals.github(.iWo/. Haverals)h;ttp://mikekestemont.github.(iMo/. Kestemont)
ȉ</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>The Middle Dutch Manuscripts Surviving from the Carthusian Monastery of Herne (14th century): Constructing an Open Dataset of Digital Transcriptions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>WouterHaveral</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mike Kestemont</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Antwerp Center for Digital Humanities and Literary Critcism (ACDC), University of Antwerp</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Center for Digital Humanities (CDH), Princeton University</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for the Study of Literature in the Low Countries (ISLN), University of Antwerp</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>A substantial collection of Middle Dutch manuscripts survives from the Carthusian monastery of Herne (Hérinnes-lez-Enghien) in nowadays Belgium. During the latter half of the fourteenth century, Herne served as a signi昀椀cant literary hotspot in the region around Brussels, with a devoted community of monks deeply involved in the production of (vernacular) texts and manuscripts, o昀琀en as collaborative e昀orts. The corpus o昀ers abundant material for the (computational) exploration of authorship, translation, and scribal cultures in the premodern Low Countries. Yet, much of this material has remained digitally inaccessible. Here we describe the creation of an almost exhaustive, open-access dataset comprising diplomatic transcriptions of all known Middle Dutch Herne manuscripts, acquired through handwritten text recognition. Apart from rich codicological and textual metadata, we include a normalized text layer (with expanded abbreviations), as well as a linguistic annotation layer (with lemmas and part of speech tags). We conclude by discussing our work against current trends in medievalist scholarship. The dataset is released together with this paper and we encourage its re-use in future research.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Carthusians</kwd>
        <kwd>medieval literature</kwd>
        <kwd>handwritten text recognition</kwd>
        <kwd>scribal cultures</kwd>
        <kwd>orthography</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Herne and Middle Dutch literature</title>
      <p>A sizable corpus of Middle Dutch literature survives today, composed in the various
vernacular Germanic dialects that were spoken in the Low Countries during the middle ages
(ca. 1200-1450). These works were manually copied onto handwritten text witnesses, mostly
manuscripts or book rolls, which are nowadays o昀琀en only extant in a fragmentary or heavily
damaged state. In many cases, very little is known concerning the original provenance of such
books, including details about the scribes, patrons, and the date and place of composition.
Especially for vernacular manuscripts, it is challenging to connect the remaining pieces of evidence.
Against this backdrop, the Herne corpus stands out: an extensive collection of Middle Dutch
manuscripts that were all produced in a single environment, by a restricted number of
intensively collaborating scribes, during the limited time period of the second half of the fourteenth
century.</p>
      <p>Founded in 1314 as the 昀椀rst Carthusian monastery in the Low Countries, the monastery in
Herne (Fr. Hérinnes-lez-Enghien) is located roughly thirty kilometers southwest of Brussels.
The monastery actively liaised with cultural stakeholders in the city; as indicated by some
prologues, local citizens acted as patrons for speci昀椀c projects, while the monks engaged in
commercial (pro pretio) book production, delivering, or correcting copies for Brussels book
dealers, for instance. This heightened level of book production at Herne should be understood
against the Carthusian monastic rules that promoted 1si5l]e.nTchee[ monks spent most of
their time in individual cells and were restricted from verbal communication except during
speci昀椀c moments each week. As these rules of silence and minimal outside interaction precluded
Carthusians from traditional preaching, the (silent) production and dissemination of written
texts became their principal form of apostolate.</p>
      <p>
        The Carthusian orientation of the Herne monastery in昀氀uenced their text and book
production in speci昀椀c ways [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Firstly, because the monks worked in isolation, Carthusian libraries
are characterized by the presence of multiple copies of important texts (so that they could
consult these works simultaneously). Secondly, because the monks were silent, collaborations
were o昀琀en discussed in writing: many manuscripts contain marginal notes in which
collaborating scribes discussed the quality of a text or negotiated translations. Such detailed paratexts are
extremely rare for this period, providing an exceptional lens to study medieval scribal cultures.
Thirdly, while the (still young) monastery also produced many books for itself, an important
share of the manuscripts were in fact intended for lay readers beyond its walls. This helps
explain the fourth de昀椀ning characteristic of the Herne corpus, namely its focus on the vernacular.
The monastery acted as a hub for the translation and distribution of Middle Dutch literature,
part of its apostolate involved making Latin texts accessible to a lay audience. A remarkable
achievement, for instance, was the Herne Bible translation initiated aro2u4n,d8]1.3T6h0i[s
project became the 昀椀rst near-exhaustive vernacular Bible translation in prose in history, a
venture both novel and controversial in the fourteenth century. Many of the surviving prologues
discuss the backlash faced by the translators.
      </p>
      <p>
        While mostly religious texts from Herne survive, it is clear that the Carthusians had access
to a broad spectrum of contemporary material, including a variety of worl1d4l]y.
tFeoxrtlsit[erary history, Herne holds special importance as it is home to many rare works, including the
two earliest manuscripts of the world-renowned mystical poetess Hadewijch. Moreover, many
unica are only extant in Herne copies, such as strophic poems by Jacob van Maerlant or the
second part of the Middle Dutch adaptationSopfetcuhleum historiale [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Finally, Herne’s
history is intertwined with the biography of John of Ruusbroec, another famous mystical author
(and, a昀琀er Anne Frank, the most o昀琀en translated Dutch auth39o,r3)2[]. The Carthusians had
gained access to an unauthorized copy of an early work of Ruusbroec and asked him for
clari椀昀cations regarding some controversial passages. In 1362, despite his advanced age, Ruusbroec
journeyed on foot to Herne to discuss and clarify controversial passages with the monks. His
clari昀椀cations would eventually yield an important commenDtaatrbyoe(csken der verclaringhe).
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. The surviving corpus of manuscripts</title>
      <p>The attribution of the Herne corpus to this monastery only came about recently. Prior to 2002,
the lion’s share of the manuscripts discussed here were widely believed to have been produced
in Rooklooster (FrR.ouge Cloître), an Augustinian priory located on the eastern outskirts of
Brussels and established in 1367. In his pioneering PhD, Erik Kwakkel was able to show that
most of these books were, in fact, not copied in Rooklooster but in 2H6e]r.nHee[convincingly
argued that only later, a surprisingly large share of the Herne books ended up in the
possession of the Rooklooster monastery, which explains why so many of these manuscripts bare
contemporary ownership marks pointing to the priory. Interestingly, these books arrived in
Rooklooster via rather diverse routes and didn’t travel2i5n].bMulakn[y of these specimens
feature in a well-known book list (entDitieleDdietsche boeke die ons toebehoeren or “The Dutch
books that belong to us”), although it is still unclear whether this inventory has been drawn
up in Rooklooster or in Herne.</p>
      <p>Kwakkel was able to attribute the production of a cohesive collection of manuscripts to the
Herne charterhouse on the basis of a variety of arguments. Firstly, a number of scribal hands
can be localized in Herne beyond reasonable doubts, because they have worked on documents
that played a central role in the local community, such as the so-called Necrol o g),y Scribe (
who contributed to the local necrologies (e.g. Brussels, RL, 21536–40). Consequently, scribes
who co-produced manuscripts withcan also be con昀椀dently located in Herne, such as the
Speculum Scribe ( ; see Fig. 1). Secondly, it seems that the Herne community employed a
unique correction system, through placing a struck out delta in the margin. Given that this
correction system appears exclusive to Herne during this period, the production of many other
manuscripts can tentatively be attributed to the charterhouse.</p>
      <p>
        All in all, Kwakkel has been able to identify 13 hands who contributed to Herne manuscripts
in the second half of the fourteenth century. (These have been numbered 1 to 13, although
e.g. scribe 1 and 2 are also frequently referred toanads .) These manuscripts, however,
remain shrouded in anonymity: none of the local scribes are known by their personal names.
Despite the existence of numerous highly informative prologues, the monks 昀椀ercely avoided
self-identi昀椀cation in these texts and frequently referred to their work as a collective endeavor
(e.g., through the use of the 昀椀rst-person plural pronoun). Only a small number of Carthusians
lived in Herne at any given time; around 1390, for instance, we know the names of the 17
inhabitants who resided permanently in the community (both monks and lay brothers),
including two brothers, Petrus and Jacob Naghel, who both took up authoritative positions in
the community 2[6, p. 85]. Many scholars have attempted to identify Petrus Naghel as the
‘Bible translator of 136204’][, although this attribution has remained controversial – not in the
least because many of the arguments that apply to Petrus equally apply to his brother Jacob
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Some observations can nevertheless be drawn from the scant evidence that is available:
Petrus Naghel, for instance, died in 1395 and does cannot be identi昀椀ed as the hyper-productive
Speculum Scribe ( ), whose youngest (dated) booklet stems from 1402.
      </p>
      <p>
        The research project “Silent voices: A Digital Study of the Herne Charterhouse as a Textual
Community (ca. 1350-1400)”, funded by the Flemish Research Agency (Belgium), studies the
evolution of the scribal culture in the Herne monastery through computational text analysis
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. For this project, limited to the vernacular book production in the charterhouse, we rely
on diplomatic transcriptions, obtained through handwritten text recognition. In this paper,
we outline the construction of the underlying dataset, which is released in conjunction with
this paper. The composition of the full corpus can be found in 1T.aTblhee corpus is
semiexhaustive and encompasses almost all currently known Middle Dutch manuscripts associated
with Herne (i.e. the manuscripts which were produced there or which were at least in the
possession of the community). Non-vernacular manuscripts (e.g. the Latin codex Brussels,
RL, 21536-40) are excluded, as well as very short texts (such as the aforementioned book list).
Local scribes are identi昀椀ed through their Kwakkel number. Ghent, UL, 941 (Hadewijch MS
C) and Paris, Bibl. de l’Arsenal, 8224 are additionally included: even though these were
not produced in Herne, they stand in a close relationship to the corpus. Note that most of the
manuscripts can only be dated approximately.
      </p>
      <p>When considering the codices listed in T1a,bilteis important to note that these are typically
composite manuscripts, consisting of distinct codicological units, which were sometimes not
even bound together until post-medieval times. Kwakkel adopts the term “production unit” in
this context: “groups of quires that formed a material unity at the time of production. Such
quires were copied “in one go”, by either one or more scribes. A codex may contain several
production units2”7[, p. 13]. This distinction is crucial, because it is typically the production
units that can be unambiguously dated. The larger codex Vienna, ÖNB, Cod. 13.708, for
instance contains an older core that dates to 1393–1394, but also a younger extension that was
only added in 14022[6, p. 97–100]. The dataset presented here therefore contains detailed
metadata on the production unit which each manuscript page belongs to.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Scribal characteristics</title>
      <p>The contributions of the 13 local scribes vary in size and nature. Whereas the corpus contains
a few, extremely large scribal oeuvres, many hands only appear sporadicall2ys.hToawblsefor
each manuscript witness, the number of production units (containing Middle Dutch), folium
sides, number of glyphs that are available, alongside the number of unique scribal hands. In
handwritten text recognition (HTR), the size and diversity of the annotated training material
are known to have a meaningful impact on the model quality. We therefore anticipate lower
HTR performance for less common scribal hands. Note that this Carthusian corpus contains
several parallel transcriptions of the same works: these pose interesting cross-source
challenges for HTR. On the one hand, transcribing a new version of a text that has been observed
during training might increase the recognition accuracy. On the other hand, there’s also a risk
of the engine over昀椀tting to the previous copy and ignoring variations present in the new one.</p>
      <p>Another important factor is the intended readership of a codex which correlates with the
quality of its execution. Local scribes were able to adapt the quality of a codex and their writing
style as a function of the anticipated audience – for various scribes, Kwakkel distinguishes
between a low, medium or high ‘style’ of execution (see2Ffiogr. examples). The codex Brussels,
RL, 394–98, for instance, was intended for a (female) Benedictine sister in VoFrosrtes(tF)r,.and
thus explicitly commissioned for the outside world. This attractive manuscript o昀ers high
readability, with a generous leaf size, spacious layout, and fewer abbreviations to enhance content
accessibility, evidently to enhance the accessibility of its contents. Manuscripts which were
likely meant to function in-house only adopt much more abbreviations. At the other end of
the spectrum, we 昀椀nd the small-size miscellany Ghent, UL, 1374, created entirely from le昀琀over
low-grade parchment, which appears to have functioned as a personal booklet for the scribe.
The abbreviation frequency is particularly high in this manuscript with its densely packed
layout. Such aspects of readability and accessibility will inevitably in昀氀uence the quality of the
automatic transcriptions discussed below.
(a) High style: fol. 4r (cutout) in Brussels, RL(b,) Medium style: fol. 1r (cutout) in Brussels, RL,
2485, which was meant to function outside 3093-95, which likely functioned locally inside
of Herne. the charterhouse.</p>
      <p>
        A major challenge of the Herne corpus are the aforementioned abbreviations, which are a
hallmark of medieval handwritten cultures. Abbreviations (or brev2ig0r])awpherse[ used to
increase writing speed and conserve space. The Herne scribes adopted abbreviations profusely
(especially in comparison to other contemporary texts in the vernacular), which they borrowed
from Latin scribal cultures. Older editorial scholarship (e.g. in the Lachmannian tradition
[
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]) o昀琀en aimed to reconstruct the authorial texts, rather than acknowledge individual copies
as a scholarly object of study in their own right; in this paradigm, abbreviations were o昀琀en
treated as meaningless ‘accidentals’, that presented an obstacle, rather than an opportunity for
research. Recent work in computational philology has rehabilitated the value of abbreviations,
for instance, for scribal attributi2o2n,s17[], linguistic analysi3s8[], stemmatology1[] and
other innovative forms of material philol1o2g,1y9[
        <xref ref-type="bibr" rid="ref30">, 30</xref>
        ].
      </p>
      <p>In this project too, we have adopted a (hyper-)diplomatic transcription standards that
attempts to mimic the glyphs on the page, including the non-expanded abbreviations, as closely
as possible. We report on an automated approach to expanding these abbreviations below.
Throughout the project, we adhere to the diplomatic transcription MVN guidelines developed
by Boot and Brinkman3][, which have been published online with a digital edition framework
for the (popular, but proprietary) XML editing so昀琀wOaxryegen [url]. An illustration of this
fairly deep and complex encoding in practice can be found in3aF.igT.hese digital guidelines
are derived from those for the long-standing print editioMnisdedreileeseuwse
Verzamelhandschri昀琀en uit de Nederlanden , originally developed by Merte2n9s].[Table3 o昀ers a selection of
the most common abbreviation practices in the corpus, together with the character reference
used in the MVN-guidelines to identify glyphs, as well as a number of illustrative examples.
This table excludes multiple more speci昀椀c, yet not uncommon brevigraphs, suchea’san‘md
‘m n’ for ‘minne’ and (“love”) ‘minnen’ (“to love”) respectively or the dense logogr̄pācm’̄ f‘oxr
‘christus’.
(b) XML snippet of the manual transcription of the lines above, following the MVN guidelines.</p>
      <p>Note how abbreviated tokens are represented usincghotihcee element, which has aanbbr
and expan subnode. Under thaebbr node, abbreviation glyphs are e.g. encoded usingg- a
node with aref attribute; the expanded counterparts of such glyphs unedxepratnhneode are
enclosed byex tags.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Handwritten Text Recognition</title>
      <p>
        For the transcription of our corpus, we leveraged the technology of Handwritten Text
Recognition (HTR), speci昀椀cally, the Transkribus platfo2r1m, 3[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Initially, the digital images of all
manuscripts in the corpus (see Tab1)lewere uploaded to this platform. An essential subsequent
step in any HTR application involves layout analysis. This process involves identifying the
speci昀椀c region on each folio, demarcated by pixel coordinates, that contains the text. Additionally,
the recognition of baselines within these designated regions is required, which determine the
location of the text.
      </p>
      <p>In order to prevent the misinterpretation of elements on the parchment surface as textual
regions, these regions were manually outlined. Unlike the uniformity of blank paper, parchment
o昀琀en contains organic elements like hair follicles, gaps, and various other irregularities, which
an automated layout model might mistakenly interpret as text regions. By manually outlining
these regions, we managed to bypass this issue. Subsequent to the manual outlining of text
regions, baselines within these regions were identi昀椀ed. For this procedure, we used an automated
approach. Transkribus o昀ers the possibility of training dedicated models speci昀椀cally for
recognizing baselines. For the task of baseline detection, we used two di昀erent models. One large
baseline detection model was trained on a diverse collection of folia, for which the baselines
were manually added. This model obtained an error rate of 7.00% on the training set consisting
of 1,007 folia, and 6.35% on the validation set with 111 folia. In certain cases, it was
bene昀椀cial to train dedicated baseline models, especially for manuscripts with unique layouts, such as
Vienna, ÖNB, Cod. 13.708 and Saint Petersburg, BAN, O 256. The former manuscript
comprises closely packed dual-column text, while the latter contains numerous text lines blurred
by water damage. Regardless of whether the general or dedicated model was employed, all
baselines were manually reviewed. This process involved eliminating incorrectly recognized
baselines and adjusting the length of overly short or long baselines. While this phase was the
most labor-intensive in building the dataset, we consider it vital to assure the quality of the
layout analysis, as it forms the cornerstone for all subsequent steps in the process.</p>
      <p>
        Next, we produced sizable sample transcriptions for almost every document in the corpus
using Transkribus. The transcription standard we adhered to, following the granularity
levels of transcription described by Robinson and Solopova, can be characterized as ‘graphemic
reproductions’. This means that the spelling in the manuscripts is retained as is, and all
abbreviations are transcribed exactly as they appear on the page. However, we made no distinction
between letter forms (e.g., a long ‘ſ’ and a round ‘s’ are both transcrib1edWaesf‘os’l)l.owed
the Medieval Unicode Font Initiative (MUFI) for encoding and displaying special characters in
medieval texts1[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Furthermore, we consistently opted for the precomposed Unicode entities,
ensuring that combinations of a letter character and a diacritical mark are always treated as a
single entity.
      </p>
      <p>
        A signi昀椀cant advantage for our project was the ability to bootstrap our transcriptions from
certain existing editions, such as the manuscripts Ghent, UL,1143]7a4n[d Brussels, RL,
309395 (see below). We also utilized diplomatic editions of substantial sections of the manuscripts
Brussels, RL, 1805-8, Brussels, RL, KBR 2877-78, and Brussels, RL, 2879-2880 [
        <xref ref-type="bibr" rid="ref17 ref22">22, 17</xref>
        ]. These
previously existing transcriptions were incorporated into the recognized baselines in
Transkribus, using the text-to-baseline method, which is unfortunately currently deprecated.
Following this, large segments of each manuscript were manually transcribed, guaranteeing that
a minimum of 10% of the folia of each manuscript were manually transcribed. This approach
ensured that our sample was representative of all manuscripts in our corpus.
      </p>
      <p>In total, our full corpus contains 6,159 folios (both recto and verso-sides) of Middle Dutch text.
A昀琀er the manual transcription phase, we have 1,331 folios of Ground Truth text. The next step
1This is in line with the tradition in Middle Dutch schol2a9r]s,hwiphe[re scholars do not normally distinguish
between the various allographsro&gt;fa&lt;nd &lt;s&gt;. We do distinguish, on the other hand, betwuee&gt;na&lt;nd &lt;v&gt; and
do not normalize such instances, even if the medieval use of a glyph doesn’t match the modern use.
was to use these materials to train a “Grand” Middle Dutch model to automatically transcribe
the remaining 4,828 folios. A 90%-10% train-test split resulted in 1,197 folios for training and
133 folios for validation. By pooling together all this data and training this comprehensive
model, we achieved a Character Error Rate (CER) of 2.7% on the validation set (s4e)e. Table
This means that, on average, 2.7 out of every 100 characters are still incorrectly recognized
when comparing the HTR output to the ground truth transcriptions. This is an encouraging
result, bolstering our con昀椀dence in the model’s capability to automatically transcribe various
scribal hands, layouts, thematic contents, and spelling pro昀椀les.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Abbreviation resolution</title>
      <p>
        To analyze the Herne manuscripts using conventional text technologies, we need an automated
way to expand the abbreviations in the output of the HTR. Many brevigraphs are
ambiguous, however, and can be expanded in multiple ways, depending on the context in which they
occur (see Table3). We therefore approach abbreviation resolution as a sequence modelling
task in string transduction, in which an abbreviated line (or ‘sentence’) from a manuscript
acquired through HTR is converted into a line of normalized tokens with their correctly expanded
spelling. We operationalize this task at the token level, but need a model that can bene昀椀t from
contextual information from surrounding tokens to resolve ambiguous instances. In essence,
this task bears similarity to the traditional problem of lemmatization in the 昀椀eld of natural
language processing. Following similar work by Camps et5,a6l]., w[e resort to the PIE
lemmatizer, which is an e昀케cient lemmatizer for historic languages that abound in orthographic
variation2[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Crucially, this lemmatizer doesn’t adopt a classi昀椀cation approach but uses
recurrent neural architectures to generate a lemma, on a character-by-character basis. This means
that it can also transduce tokens into normalized forms which it hasn’t encountered during
training.
5.1. Abbreviation materials
Our resources include (partial) transcriptions of four codices where the abbreviations have been
manually resolved: (1) Ghent, UL, 1374 (Heber-Serrure MS b;yrhyming texts, but written
down in continuous prose with inter-verse punctua1t4io];n()2)[ Brussels, RL, KBR 2877-78
(Letters in Hadewijch MS B, scribe 10; prose formatted in two-column form22a]t;)(3[)
Brussels, RL, 2879-2880 (Letters in Hadewijch MS A; prose penned by multiple non-local scribes,
presented in a two-column forma2t2)];[ and (4) Brussels, RL, 3093-95 (miscellany; complete
transcription; multiple scribes; single-column p2roInsea).ddition to the Herne material, we
possess of transcriptions of multiple parallel versions of the thirteentMha-crteinjnt-utriylogy
by Jacob van Maerlant, which have been produced by dra. So昀椀e Moors (University of Antwerp)
adhering to the same MVN guidelines and markuurpl].[3 To establish how well the presented
abbreviation solver scales to other material, with a diversity of scribal hands and dialects, we
will also evaluate our procedure on this out-of-sample material. It should be noted, however,
that thMeartijn tradition, like most Middle Dutch texts, generally comprises far fewer
abbreviations.
      </p>
      <p>Below we report a series of experiments using a conventional setup, dividing the
accessible material into training, validation, and test splits. Additionally, we report on a series of
cross-domain and cross-scribal setups to stress-test the generalizability of our approach. We
train PIE only for the lemmatization task, using the default con昀椀guration setups, but a
relatively smaller layer size (to avoid over昀椀tting, in what is clearly a less complicated task than full
lemmatization). In Tab5l,ea random sample of pairs of abbreviated and expanded verse lines is
provided. Given our focus on the line level, certain lines may begin or end with an incomplete
token, a common occurrence in prose texts when words are split over two lines. Some lines
need no changes at all, especially inMtahretijn subset. Importantly, the lemmatizer shouldn’t
be greedy and must learn when to simply reproduce the input as output without alterations.</p>
      <p>Table6 displays descriptive summary statistics about the data per witness type: the number
of available line pairs, the number of identical pairs, the length of the abbreviated/expanded
lines (in characters), as well as the length di昀erence between the two. A number of
observations are important. The number of lines per Herne witness is fairly well balanced (25,017 lines
in total), but the cumulative numbeMraorftijn lines is much larger (note, however, that these
are all parallel copies of the same text). The abbreviation rate is also the lowest in this witness
category and more than half of the lines in fact need no abbreviation resolution at all, with very
low mean di昀erence scores overall. The manuscripts which presumably functioned locally in
Herne (Ghent, UL, MSS 1374 and Brussels, RL, 3093-3095) clearly have the highest
abbreviation density overall. The line length can also vary signi昀椀cantly, depending on the layout of a
manuscript page. MS 1374, for instance, with its extravagantly dense prose layout, combines a
very long line length with a high rate of brevigraphs: with some exaggeration, it would appear
nearlyall words in this codex are abbreviated.</p>
      <p>The evaluation is reported in Ta6b.lTehe results are high for the in-domain test set, which
was to be expected since there were no unknown tokens or targets. The few errors that persist
are confusions between ambiguous resolutions; -ee.rg.vs -aer for #apomod) or -m vs -n for
#bar. In many cases, the limited context of a single line likely prevented a correct
disambiguation. Only in a handful of cases, the system makes less forgivable errors, where it incorrectly
copies a substring of the input, for instance. Some of the errorsMinarttihjnetexts point to
abbreviatory practices that are not encountered in Herne,#suuscmhodasfor -as (instead of
-us).
2This last manuscript has been transcribed by Dr. Ine Kiekens in 2018.
3This corpus in fact includes two copies that originated from Herne – bo–twhhbiych will be excluded from the
Martijn corpus below.
Ghemaect en̄ .i. ādʼ makē : Al ne warē Ghemaect ende .i. ander maken : Al ne Ghent, UL, 1374
nᵗ die sakē : Alsoe waren niet die saken : Alsoe
dꝫ hi heme : Vʼswoʼ . dādʼ alsict vʼneme dat hi heme : Verswoer . dander alsict Ghent, UL, 1374
: Was ī ouʼspele gᵒt : verneme : Was in ouerspele grot :
Wes warachtich ende stout</p>
      <p>Wes warachtich ende stout</p>
      <p>Martijn (D)
en̄ ōdʼ dē dorē : Dꝫ seit die ewāgeliste ende onder den doren : Dat seit die Ghent, UL, 1374
wel : Om ewangeliste wel : Om
en̄ wōdʼleke liedekene
wʼEnde welcker du liets sneuen/
Dus ontbant my en̄ beschiet.
beghʼte altoes en̄ hare
Die alre meeste claerhᵗ die
ende wonderleke liedekene
werEnde welcker du liets sneuen/</p>
      <p>Brussels, RL, 3093-3095</p>
      <p>Martijn (D)
Dus ontbant my ende beschiet.</p>
      <p>Martijn (D)
begherte altoes ende hare
Die alre meeste claerheit die</p>
      <p>Brussels, RL, 2877-78
Brussels, RL, 2879-80</p>
    </sec>
    <sec id="sec-6">
      <title>6. Linguistic annotation</title>
      <p>Following the resolution of abbreviations in the original transcriptions, standard language
technology can be applied to the corpus. (Middle) Dutch studies have been blessed with rich
lexicographic resources developed by the Dutch Language Institute. We have applied a recently
developed lemmatizer and part-of-speech tagger on the basis of the PIE so2昀琀w8]a,rweh[ich
was trained on all of the available resources for Middle Dutch, including a subcorpus of
religious and mystical writings that should suit the Herne mater2i3a]l. wTehlel u[sed model
(dum) is publicly available from a framework developed by Clé9r]icteh[at makes it easy to
apply PIE models to new data. This linguistic enrichment layer allows scholars to abstract
over surface variation, which can be bene昀椀cial for studies that wish to abstract over scribal
accidentals (for instance, in the context of search), but also for applications in stylometry or
semantic analysis.</p>
      <p>In conjunction with this paper, all related materials are released on the Zenodo data
repository platform, accessible via the follolwinink.gThese materials include the transcriptions in
an XML scheme, metadata, and linguistic annotation layers. All these materials are available
under a CC-BY 4.0licence(Attribution 4.0 International). Users are allowed to share and adapt
these resources for any purpose, with the condition of providing adequate attribution to this
paper. However, due to unresolved intellectual rights (refer to the
link1s),inthTeaoblreiginal images will be stored in a separate “restricted access” Zenodo repository. They can only
be shared within the context of a scienti昀椀c collaboration under Belgian copyright law. Finally,
the corpus-wideTranskribus htr model is set to be made publicly available on the so昀琀ware
platform.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Discussion: an un-edition?</title>
      <p>In this dataset paper, we reported on the construction of a near-exhaustive collection of digital
transcriptions of the Herne manuscripts that is readily and openly available for computational
text analysis. Because the project’s focus lies with studying Herne as a scribal community, we
have chosen to produce a hyper-diplomatic rendering of the manuscripts. Apart from
theoretical considerations, this choice was also meant to help the HTR engine, because the target
transcriptions stay relatively closer to the glyphs on the page and, thus, require less contextual
interpretation and disambiguation. These choices set us apart from most of the prior Herne
editions that have been heavily text critical in orien4t,a2,t7io].nO[lder editions abstracted
over various manuscript versions of a text, in an attempt to reconstruct the original
authorial version. Critical editions typically normalize a text’s surface v3a3r]:iaetdiiotnor[s adapt
spellings (e.g. allographs) to modern orthographic conventions and they resolve abbreviations
in a manner that doesn’t allow to reconstruct the original glyphs on the page. For all these
reasons, critical editions do not optimally suit the computational analysis of scribal pro昀椀les.</p>
      <p>In a recent online semin4art,he traditional endeavour of the critical, closed, printed edition
4The Birth of the UnEdition led by Ben Albritton and Laura Morreale under the umbrella of the ‘DarkArchives 20/21’</p>
      <sec id="sec-7-1">
        <title>Edition</title>
      </sec>
      <sec id="sec-7-2">
        <title>Un-Edition</title>
        <sec id="sec-7-2-1">
          <title>Printed (in paper)</title>
          <p>Printed (in paper)</p>
          <p>Static (for eternity)
Critical (multi-witness)
Hide sources
Old philology</p>
          <p>Accuracy
Paywalled, copyrighted
Focus on text</p>
        </sec>
        <sec id="sec-7-2-2">
          <title>Digital (in the cloud)</title>
          <p>Collaborative (asynchronous)
Dynamic (until next bugfix)
Hyper-diplomatic (single-witness)
Foreground sources</p>
          <p>New philology
Accessibility</p>
          <p>Open science</p>
          <p>Focus on document
was put in a stimulating, yet perhaps somewhat undue opposition with the novel forms of
collaborative, open, digital edition – tentatively termed an “un-edition8”. wIne Tsuarbvley
some of the oppositions that were discussed at this event. Exaggerated by the global lockdown
conditions during the COVID pandemic, the digital medium has increased the need among
medievalists for digital editions that are widely and openly accessible, beyond the conventional
consultation in physical libraries. Increasingly, editions are a collaborative, asynchronous
endeavour. Static, printed editions, published for eternity, lose their relevance in the digital
sphere: apart from the paywalls that limit the accessibility, un-editions have the advantage
of being dynamic – they are open to constant improvement. Because of the ubiquity of digital
facsimiles (that are increasingly served over open protocols like IIIF) as well as the a昀ordances
of handwritten text recognition, the diplomatic edition is going through a clear renaissance. In
the words of Fisher in a thought-provoking article:
‘The provision of digital facsimiles has put the manuscripts, the very material base
of any editorial endeavor, into focus again. Several editions have been created that
engage primarily with individual manuscripts; others have posited a wide range
of variance as a central characteristic of medieval literature instead of relegating
variants to the footnotes of ahistorically normalized and regularized texts or
speculative reconstructions of archetypes and autho1r3i,tpie.sS’2[65].</p>
          <p>
            Fisher goes on to stress that the idea of a critical edition is not obsolete – on the contrary, in
his view the need for critical editions has only increased with the availability of
documentoriented transcriptions: ‘without a critical text the modern reader is at a loss to adjudicate on
the quality of the textual version picked up randomly on the inibtiedr.)n. eHte’ (nevertheless
notes at the same time that only a relatively small number of born-digital critical editions of
texts exists, although the necessary technology is in fact available. This observation touches
upon a more fundamental development in medieval studies, where scholars have come to prefer
availability over accuracy in recent years. This aspect is also crucial to the concept of an
unedition, which favors an openly available text over a more accurate, yet closed text. This might
inter-conference, on Thursday 18th March 2021, via Zoom. uSrele] [
be a controversial point but colleagues seem, overall, quite happy to strike a scholarly trade-o昀
between the poor textual quality of mega-collections like the Internet Archive or Google Books
and the vast serendipity which these textual reservoirs enable. The COVID pandemic has, in
this sense, fast-tracked scholarly developments that had been emerging over a much longer
period, in the wake of the material or “artefactual” phi3l7o]l.oCguyr[rently, many terms have
surfaced to describe the new forms of digital editorial scholarship: un-edition, proto-edition,
arti昀椀cial edition3[
            <xref ref-type="bibr" rid="ref4">4</xref>
            ].
          </p>
          <p>In conclusion, our project o昀ers a fresh approach to the study of the Herne manuscripts,
emphasizing the usefulness of hyper-diplomatic transcriptions and leveraging digital technology
to render them accessible to a broader scholarly community. The concept of the “un-edition”,
although still emerging and somewhat controversial, holds promise as it aligns with the
ongoing shi昀琀 in medieval studies from accuracy to availability. By making the original, abbreviated
transcriptions of our corpus available online, along with the fully expanded and lemmatized
texts, part-of-speech tags, and extensive metadata, we embrace this transformation. In this
sense, we hope our work underscores the dynamic nature of digital humanities, which, much
like the manuscripts we study, are in constant 昀氀ux.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>Thanks to Anouck Kuypers, Sam Verellen and Frans de Jonge for their work on the
transcriptions. The transcription of Brussels, RL, 3093-95 was contributed by Dr. Ine Kiekens. We
acknowledge the help of Renée Gabriël and Peter Boot in previous collaborations that relate to
the present paper. This work has been funded by the Flemish Research Agency (FWO) in the
context of the project “Silent voices: A Digital Study of the Herne Charterhouse as a Textual
Community (ca. 1350-1400)”. All (transcriptions) data and code for this paper are available
from Zenodo:https://zenodo.org/doi/10.5281/zenodo.10005.253</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Andrews</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Macé</surname>
          </string-name>
          . “
          <article-title>Beyond the Tree of Texts: Building an Empirical Model of Scribal Variation through Graph Analysis of Texts and StemmaLtitae”r</article-title>
          .
          <source>aIrny: and Linguistic Computing 28.4</source>
          (
          <issue>2013</issue>
          ), pp.
          <fpage>504</fpage>
          -
          <lpage>521</lpage>
          . doi:
          <volume>10</volume>
          .1093/llc/fqt03.2
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berteloot</surname>
          </string-name>
          , G. Claassens, and
          <string-name>
            <given-names>W.</given-names>
            <surname>Kuiper</surname>
          </string-name>
          , eGdusld.en Legende: De Middelnederlandse vertaling van de '
          <article-title>Legenda aurea' door Petrus Naghel, uitgegeven naar handschri昀琀 Brussel</article-title>
          , Koninklijke Bibliotheek,
          <volume>15140</volume>
          .
          <string-name>
            <surname>Deel</surname>
            <given-names>II</given-names>
          </string-name>
          . Brepols,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Boot</surname>
          </string-name>
          and H. BrinkmanR.ichtlijnen voor digitale diplomatische edities in de reeks Middeleeuwse Verzamelhandschri昀琀en uit de Nederlanden.
          <year>2020</year>
          . uhrtlt:ps://github.com/Hu ygensING/mvn-xml/blob/main/docu/Richtlijnen%5C%
          <article-title>20MVN%5C%20digitaa</article-title>
          .l.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K. D.</given-names>
            <surname>Bundel</surname>
          </string-name>
          . “
          <article-title>Van woerde tot woerde o昀琀 van synne te sinne . Petrus Naghel en het translatorium van de kartuis te Herne (ca</article-title>
          .
          <fpage>1350</fpage>
          -1400)”.
          <source>PhD thesis</source>
          .
          <source>KU Leuven</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>J.-B. Camps</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Clérice</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Pinche</surname>
          </string-name>
          . “
          <article-title>Noisy medieval data, from digitized manuscript to stylometric analysis: Evaluating Paul Meyer's hagiographic hypothDesigisit”a.lIn: Scholarship in the Humanities 36</article-title>
          .Supplement_
          <volume>2</volume>
          (
          <issue>2021</issue>
          ), pp.
          <fpage>ii49</fpage>
          -
          <lpage>ii71</lpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .1093/llc/fqa b033.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>J.-B. Camps</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Vidal-Gorène</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vernet</surname>
          </string-name>
          . “
          <article-title>Handling Heavily Abbreviated Manuscripts: HTR Engines vs Text Normalisation ApproachesD”</article-title>
          .
          <source>oIcnu:ment Analysis and Recognition - ICDAR</source>
          <year>2021</year>
          Workshops. Ed. by
          <string-name>
            <given-names>E. H. Barney</given-names>
            <surname>Smith</surname>
          </string-name>
          and
          <string-name>
            <given-names>U.</given-names>
            <surname>Pal</surname>
          </string-name>
          . Cham: Springer International Publishing,
          <year>2021</year>
          , pp.
          <fpage>306</fpage>
          -
          <lpage>316</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Claassens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Kuiper</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Berteloot</surname>
          </string-name>
          , eGdusl.den legende. De Middelnederlandse vertaling van de Legenda aurea door Petrus Naghel. Brepols,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G. H. M.</given-names>
            <surname>Claassens</surname>
          </string-name>
          and K. de Bundel. “Petrus Naghel, Übersetzer in HernZei”t.
          <source>sIcnh:ri昀琀 Für Deutsche Philologie</source>
          <volume>130</volume>
          (
          <year>2011</year>
          ), pp.
          <fpage>267</fpage>
          -
          <lpage>281</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Clérice</surname>
          </string-name>
          .
          <article-title>Pie Extended, an extension for Pie with pre-processing and post-processing</article-title>
          .
          <year>2020</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.3883589. url: https://doi.org/10.5281/zenodo.38835.89
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Coun</surname>
          </string-name>
          . “
          <source>The Bible Translator of</source>
          <volume>1360</volume>
          ”.
          <source>OInn:s Geestelijk Erf</source>
          <volume>85</volume>
          .
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          (
          <year>2014</year>
          ), pp.
          <fpage>3</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deschamps</surname>
          </string-name>
          , ed.The Vienna manuscript of the “Second part” of the “Spiegel historiael”: Ms. Vienna, Österreichische Nationalbibliothek,
          <volume>13</volume>
          .708.
          <article-title>Medieval manuscripts from the Low Countries in facsimile</article-title>
          .
          <source>Rosenkilke and Bagger</source>
          ,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Driscoll</surname>
          </string-name>
          . “
          <article-title>The Words on the Page: Thoughts on Philology, Old and NeCwre”a. tI-n: ing the Medieval Saga: Versions, Variability, and Editorial Interpretations of Old Norse Saga Literature</article-title>
          . Ed. by
          <string-name>
            <given-names>J.</given-names>
            <surname>Quinn</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Lethbridge</surname>
          </string-name>
          . Odense: University of South Denmark,
          <year>2010</year>
          , pp.
          <fpage>87</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Fischer</surname>
          </string-name>
          . “
          <article-title>Digital Corpora and Scholarly Editions of Latin Texts: Features and Requirements of Textual Criticism”</article-title>
          .
          <source>ISnp:eculum 92.S1</source>
          (
          <year>2017</year>
          ),
          <fpage>S265</fpage>
          -
          <lpage>s287</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gabriël</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          , edDse. Heber-Serrurecodex: Gent, Universiteitsbibliotheek, Hs. 1374.
          <article-title>Diplomatische editie bezorgd door Renée Gabriël en Mike Kestemont, met een dialectologische analyse door Amand Berteloot</article-title>
          . Middeleeuwse Verzamelhandschri昀琀en uit de Nederlanden XVII. Amsterdam: Huygens Instituut voor Nederlandse Geschiedenis en Cultuur van de Koninklijke Nederlandse Akademie van Wetenschappen,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.8385501. url: https://hbsr.mvn.huygens.knaw..nl
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Gaens</surname>
          </string-name>
          and J. de GrauweD. e kracht van de stilte. Geest &amp; geschiedenis van de kartuizerorde. Peeters Publishers,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>O. E.</given-names>
            <surname>Haugen</surname>
          </string-name>
          .
          <source>MUFI character recommendation v. 3</source>
          .0.
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>W.</given-names>
            <surname>Haverals</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          . “
          <article-title>From exemplar to copy: the scribal appropriation of a Hadewijch manuscript computationally exploredJ”o.uInrn:al of Data Mining &amp; Digital Humanities On the Way to the Future of Digital Manuscript Studies (202130).</article-title>
          .4d6o2i9: 8/jdmdh.10206.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>W.</given-names>
            <surname>Haverals</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          . “
          <article-title>Silent voices. A Digital Study of the Herne Charterhouse Scribal Community (ca</article-title>
          .
          <fpage>1350</fpage>
          -1400)”.
          <source>InQ: ueeste 27.2</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>186</fpage>
          -
          <lpage>195</lpage>
          . doi:
          <volume>10</volume>
          .5117/q ue2020.
          <volume>2</volume>
          .006.have.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Honkapohja</surname>
          </string-name>
          . “Digital Approaches to Manuscript Abbreviations:
          <article-title>Where Are We at the Beginning of the 2020s?” InD:igital Medievalist 14 (</article-title>
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .16995/dm.88.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>K.</given-names>
            <surname>Honkapohja</surname>
          </string-name>
          . “Digital Approaches to Manuscript Abbreviations: Where Are We Now?”
          <source>In: Digital Medievalist 5.1</source>
          (
          <year>2009</year>
          ). doi:
          <volume>10</volume>
          .16995/dm.88.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kahle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Colutto</surname>
          </string-name>
          , G. Hackl, and
          <string-name>
            <surname>G. Mühlberger.</surname>
          </string-name>
          “
          <article-title>Transkribus - A Service Platform for Transcription, Recognition and Retrieval of Historical Documen14ttsh”</article-title>
          .
          <source>IIAnP:R International Conference on Document Analysis and Recognition (ICDAR)</source>
          . Vol.
          <volume>4</volume>
          .
          <year>2017</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          . “
          <article-title>A computational analysis of the scribal pro昀椀les in two of the oldest manuscripts of Hadewijch's letters”</article-title>
          .
          <source>SIcnri:ptorium 69.1</source>
          (
          <issue>2015</issue>
          ), pp.
          <fpage>159</fpage>
          -
          <lpage>177</lpage>
          . doi:
          <volume>10</volume>
          .14 84/j.scri.
          <volume>5</volume>
          .103464.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          , G. de Pauw, R. van Nie, and
          <string-name>
            <given-names>W.</given-names>
            <surname>Daelemans</surname>
          </string-name>
          . “
          <article-title>Lemmatization for variationrich languages using deep learning”</article-title>
          <source>D.Iingi:tal Scholarship in the Humanities 32.4</source>
          (
          <issue>2016</issue>
          ), pp.
          <fpage>797</fpage>
          -
          <lpage>815</lpage>
          . doi:
          <volume>10</volume>
          .1093/llc/fqw034.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kors</surname>
          </string-name>
          .
          <article-title>De bijbel voor leken</article-title>
          .
          <source>Studies over Petrus</source>
          Naghel en de historiebijbel van 1361.
          <string-name>
            <surname>Brepols</surname>
          </string-name>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>E. Kwakkel. “</surname>
          </string-name>
          <article-title>A meadow without 昀氀owers. What happened to the Middle Dutch manuscripts from the Charterhouse Herne?</article-title>
          <source>Q”uIna:erendo 33</source>
          .1-
          <fpage>2</fpage>
          (
          <year>2003</year>
          ), pp.
          <fpage>191</fpage>
          -
          <lpage>211</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kwakkel</surname>
          </string-name>
          .
          <article-title>Die Dietsche boeke die ons toebehoeren</article-title>
          . De kartuizers van Herne en de productie van Middelnederlandse handschri昀琀en in de regio Brussel (
          <volume>1350</volume>
          -
          <fpage>1400</fpage>
          ) . Peeters,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>E. Kwakkel.</surname>
          </string-name>
          “
          <article-title>Towards a terminology for the analysis of composite manuscripts”</article-title>
          .
          <source>In: Gazette du livre médiéval 41</source>
          (
          <year>2002</year>
          ), pp.
          <fpage>12</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          , Á. Kádár, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          . “
          <article-title>Improving Lemmatization of Non-Standard Languages with Joint Learning”.PIrno:ceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</article-title>
          , Volume
          <volume>1</volume>
          (Long and Short Papers).
          <year>2019</year>
          , pp.
          <fpage>1493</fpage>
          -
          <lpage>1503</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          -1153.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mertens</surname>
          </string-name>
          .Richtlijnen voor de uitgave van Middeleeuwse Verzamelhandschri昀琀en uit de Nederlanden.
          <source>Hilversum: Verloren</source>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>S.</given-names>
            <surname>Moors</surname>
          </string-name>
          . “Een bijzondere Brabantse tekstgetuige van de Martijntrilogie van Jacob van Maerlant.
          <string-name>
            <surname>De fragmenten Lyon</surname>
          </string-name>
          , Bibliothèque municipale,
          <source>ms 6848”Q. uIene:ste 29.1</source>
          (
          <issue>2022</issue>
          ), pp.
          <fpage>36</fpage>
          -
          <lpage>65</lpage>
          . doi:
          <volume>10</volume>
          .5117/que2022.1.002.moor.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>G.</given-names>
            <surname>Muehlberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Seaward</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Terras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Ares</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bryan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Colutto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Déjean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Diem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fiel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Greinoecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Grüning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hackl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Haukkovaara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Heyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hirvonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hodel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jokinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kahle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kallio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Kleber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Labahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Lang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Laube</surname>
          </string-name>
          , G. Leifert, G. Louloudis,
          <string-name>
            <given-names>R.</given-names>
            <surname>McNicholl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Meunier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mühlbauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Philipp</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Pratikakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Puigcerver</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Putz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Retsinas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sablatnig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Sánchez</surname>
          </string-name>
          , P. Scho昀椀eld, G. S昀椀kas, C. Sieber,
          <string-name>
            <given-names>N.</given-names>
            <surname>Stamatopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Strauß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Terbul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Toselli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ulreich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Vidal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Walcher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Weidemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wurster</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Zagoris</surname>
          </string-name>
          . “
          <article-title>Transforming scholarship in the archives through handwritten text recognition: Transkribus as a case stuJdouyr”n</article-title>
          .
          <source>Ianl:of Documentation 75.5</source>
          (
          <issue>2019</issue>
          ), pp.
          <fpage>954</fpage>
          -
          <lpage>976</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>F. van OostromW</surname>
          </string-name>
          .ereld in woorden: Geschiedenis van de Nederlandse literatuur 1300-
          <fpage>1400</fpage>
          . Amsterdam: Prometheus,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pierazzo</surname>
          </string-name>
          .
          <source>Digital Scholarly Editing: Theories, Models and Methods. Routledge</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <article-title>Proto-editions: Historians and the ”Something between digital image and digital scholarly edition”</article-title>
          .
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.8107922. url: https://doi.org/10.5281/zenodo.810792 2.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>P.</given-names>
            <surname>Robinson</surname>
          </string-name>
          and
          <string-name>
            <surname>E. Solopova.</surname>
          </string-name>
          “
          <article-title>Guidelines for transcription of the manuscripts of the Wife of Bath's Prologue”</article-title>
          .
          <source>InTh:e Canterbury Tales Project. Occasional papers I (</source>
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>P.</given-names>
            <surname>Roelli</surname>
          </string-name>
          , edH.andbook of Stemmatology. History, Methodology,
          <string-name>
            <given-names>Digital</given-names>
            <surname>Approaches</surname>
          </string-name>
          . Berlin, Boston: De Gruyter,
          <year>2020</year>
          . do1i0:.
          <volume>1515</volume>
          /9783110684384.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>L.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          . “Material Philology”.
          <article-title>HIann:dbook of Pre-Modern Nordic Memory Studies</article-title>
          . Interdisciplinary Approaches. Ed. by
          <string-name>
            <given-names>J.</given-names>
            <surname>Glauser</surname>
          </string-name>
          , P. Hermann, and
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          . Berlin, Boston: De Gruyter,
          <year>2019</year>
          , pp.
          <fpage>210</fpage>
          -
          <lpage>216</lpage>
          .
          <year>do1i0</year>
          :.
          <volume>1515</volume>
          /
          <fpage>9783110431360</fpage>
          -
          <lpage>020</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Thaisen</surname>
          </string-name>
          . “
          <article-title>Standardisation, exemplars, and the Auchinleck manuscriTphte”</article-title>
          .
          <article-title>MInu:ltilingual Origins of Standard English</article-title>
          . Ed. by
          <string-name>
            <given-names>L.</given-names>
            <surname>Wright</surname>
          </string-name>
          . Berlin, Boston: De Gruyter Mouton,
          <year>2020</year>
          , pp.
          <fpage>165</fpage>
          -
          <lpage>190</lpage>
          . doi:
          <volume>10</volume>
          .1515/
          <fpage>9783110687545</fpage>
          -
          <lpage>006</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>G.</given-names>
            <surname>Warnar.Geert Ruusbroec: Literatuur En Mystiek In De Veertiende Eeuw</surname>
          </string-name>
          . Amsterdam: Athenaeum-Polak &amp; Van
          <string-name>
            <surname>Gennep</surname>
          </string-name>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>