<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modelling Resolutions of the Dutch States General for Digital Historical Research</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marijn Koolen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rik Hoekstra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ida Nijenhuis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ronald Sluijter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Esther van Gelder</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rutger van Koert</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gijsjan Brouwer</string-name>
          <email>gijsjan.brouwerg@di.huc.knaw.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hennie Brugman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Huygens Institute for the History of the Netherlands</institution>
          ,
          <addr-line>Amsterdam, Netherlands https://</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>KNAW Humanities Cluster</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <fpage>37</fpage>
      <lpage>50</lpage>
      <abstract>
        <p>The Resolutions of the Dutch States General (1576-1796) is an archive covering over two centuries of decision making and consists of a heterogeneous series of handwritten and printed documents. This archive has rich potential for historical research, but the heterogeneity and dispersion of information makes using it for research a challenge. In this paper we describe how we deal with the challenges of structuring and connecting the information contained within this archive, and how this results in a computational platform that allows users to explore and analyse the archive through many connected layers of metadata.</p>
      </abstract>
      <kwd-group>
        <kwd>Information Extraction</kwd>
        <kwd>Digital History</kwd>
        <kwd>Data Modelling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The Resolutions of the Dutch States General (1576-1796) constitute an archival
series that covers more than two centuries of continuous decision making and
consists of more than 500,000 pages, handwritten and printed resolutions, in
separate, chronologically ordered series. The Resolutions of the States General
in the Dutch Republic are a key resource to the political history of this period
as they contain all decisions made by the States General (SG), the central ruling
body in the Republic. It was designated as a key resource when in 1905 the
work of publishing the resolutions started [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The manual editing resulted in
two series of analogue publications of (a selection of) the resolutions, divided in
an old series (14 volumes running from 1576 { 1609), a new series (7 volumes,
1610 { 1625), and a digital edition (1626-1630).3. The resolutions reveal the
decision making process and are relevant for both high and low politics. They
allow researchers to answer many di erent research questions about politics
but not only politics - in the Dutch Republic and its position in the world. The
resolutions are also key to all the other records of the SG (about 1 mile) and form
a backbone with which these other records can be connected and contextualised.
      </p>
      <p>
        Many research questions require working your way through hundreds of large
volumes of text without adequate indexes. The relevant data are hidden and
scattered across millions of paragraphs of dense and repetitive text. Moreover,
di erent research questions require di erent selections, reorganisations and
reorderings of the records to bring together and connect the dispersed information.
It is therefore crucial to extract high-quality metadata from the corpus of
resolutions on various levels, including the meetings, dates, attendants, the individual
resolutions and their topics. Many archives and libraries have experimented with
giving access to their collections by means of their digitised inventories and some
have gone a step further, using existing indexes of serial collections [
        <xref ref-type="bibr" rid="ref2 ref5 ref9">9,2,5</xref>
        ]. But
these archival referential systems are too coarse for access beyond the document
level. Moreover, the existing scholarly apparatus consists of many more reference
systems and tools that can be put to good use. Centuries of dealing with these
complications have led to a number of convenient and often-employed structures
that are part of the printed culture but are often ignored in the translation to
digital access [
        <xref ref-type="bibr" rid="ref12 ref16">16,12</xref>
        ].
      </p>
      <p>In this paper, we describe our work on identifying and extracting several
of these reference systems and structure as layers of data and metadata to
enhance access to all the resolutions in an online computational environment that
supports a broad range of digital historical research. We combine established
information extraction techniques with a work ow in which we iteratively build
models of the structure of the corpus and the many standard phrases used in
the resolutions. With these phrase models, we can exploit both expert knowledge
and the fact that the resolutions contain highly repetitive language to improve
the extraction process.</p>
      <p>Our approach departs from the important condition that the extracted
information should both re ect the structure of the resource and the type of research
questions it is supposed to support. We illustrate this with a relatively simple
research problem for the following question:</p>
      <p>Do the resolutions over time re ect an increasing possibility for citizens
to put forward their concerns to the SG?</p>
      <p>To investigate this, we need to know what types of proposals and requests
were submitted that led to resolutions, when each was put forward, by whom and
what decision was reached. Moreover, answering whether access to the general
public changed over time requires that the set of resolutions is either complete
or at least representative. Just digitising the archive does not support answering
this question. A digitised archive consists only of a combination of images and
(roughly) digitised text, without structure and only very limited metadata.</p>
      <p>This paper discusses the way we deal with a number of challenges to
transform a digitised archive into an online publication which supports the structured
analysis required for such historical research:</p>
      <p>Text recognition challenges The corpus contains handwritten and printed
materials, and many di erent types of information (paragraphs, lists, tables,
marginalia). Moreover, the language use and spelling changed throughout
the covered period.</p>
      <p>Information extraction challenges The text recognition process results in
textual representations, but without handles for using the rich structure in
the analogue documents. Rich layers of metadata need to be derived from
the text. This requires a combination of text mining expertise and domain
knowledge to identify, extract and connect the dispersed information.
Linking challenges The structure of the records is complex, with various types
of materials that are linked to each other, and identifying this structure and
making it inter-operable is crucial to making the online publication usable.
Presentation challenges Users approach the digital archive with di erent aims
and background knowledge and with di erent tasks of reading, searching,
combining and extracting information. The information system should
support a broad range of methods to select, organise and order the records.</p>
      <p>
        These challenges are interconnected, as errors in text recognition in uence
the quality of information extraction, which in turn in uences the accuracy of
linking and the possibilities for users to interact with the information system
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In the rest of this paper we discuss how we deal with these challenges.
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Resolutions as a Historical Resource</title>
      <p>The volumes of the resolutions embody a continuous series of day-to-day
meetings of the Dutch States General as the central assembly of delegates from
(ultimately) seven sovereign provinces. Each session starts with a date and a list
of attendants and the president of the day, followed by summaries of a varying
number of proposals and requests (see Figure 1) about a wide variety of
subjects of both high and low politics, such as foreign policy, nance, army and
navy, pensions and patents, administrative and cultural issues. All summarised
issues contain a decision (resolution) of acceptance, rejection or deferral, pending
further investigation or requests for information. Both unresolved and decided
issues led to trails of resolutions and all of them had to be traceable for the
SG and other governing bodies, as the decisions had the force of law. Some
summaries contain copies of letters or other important incoming documents as
insertions.
2.1</p>
      <sec id="sec-2-1">
        <title>Supporting Research Questions</title>
        <p>The types of data and metadata layers we want to make operable are related to
the types of research questions and themes that we want to support:
Narrative analysis Long-term developments and shifts, both internationally
and within Dutch society (e.g. what was the relative position and wealth of
the province, how did the competition between navy and army develop, how
important were the di erent colonies, how did the SG deal with di erent
religious groups).</p>
        <p>Thematic analysis the development of certain themes/topics over time. Do
the resolutions over time re ect an increasing possibility for citizens to put
forward their concerns to the SG?
Content analysis Quantitative analysis with regard to what or who was
discussed, when and how often (e.g. nancial and economic policy, nominations
for o ceholders or army positions, petitions by citizens).</p>
        <p>Network analysis How can we trace the dividing line between formal and
informal politics, and politics behind the scenes? Performing serial research
into the attendance at meetings and in committees can answer questions
like: who worked together? Who were involved in decisions around speci c
topics and in larger policy issues and how were these persons related?
Linguistic analysis How did the language of decision-making develop and is
it possible to link transformation of the SG's language with its growing
administrative competence?</p>
        <p>
          Addressing these questions requires operationalising several structural
elements of the written and printed texts, into multiple layers of metadata to
organise and classify the resolutions and make useful selections. For instance, to
study changes in how petitions of citizens were treated, a researcher needs to
select resolutions related to petitions, categorise them according to what group
a proposer belonged to and order them temporally. For a network analysis, a
researcher needs to select the relevant resolutions, extract and classify the types
of parties involved (e.g. proposers, the attending SG members, committees
investigating and reporting on issues, and nanciers providing budget) and the
relationships between them. Adding metadata for these di erent aspects requires
extracting the relevant information from an estimated one million resolutions,
which requires an information extraction that is automated where possible, but
which needs to be informed by export knowledge and a human in the loop. On
top of that, each selection or reorganisation made by users creates a di erent
view of the data, the interpretation of which is in uenced by our decision in
creating the metadata layers, so it is important that our processing is transparent
and visible to the user [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Operationalising the Structure</title>
        <p>In many digitisation projects for e.g. newspapers or books, the text of individual
scans is recognised using a model trained on ground truth, then metadata about
the scan is added to give ways to organise the text. For newspapers these are
typically the name of the newspaper and the date and page number. For books,
this is more di cult. The bibliographic metadata from a library catalogue can
be added, but the structure of books is often not available as metadata. The
text is accessible only as a sequence of plain text pages and maybe paragraphs,
making it hard to gure out whether the book has parts or chapters or sections,
where these start and end, and whether they have titles or headings.</p>
        <p>For the volumes of resolutions, text recognition results in one document per
scan with the recognised words and their pixel coordinates. To access the
resolutions of a speci c date, the recognised text gives you few handles to go to
the right set of pages. Using string matching to locate dates is extremely error
prone because of the combination of recognition errors, linguistic variation and
frequent references to previous dates within a resolution. You have to select the
book with resolutions for the desired year, then browse through the roughly 1000
pages of dense text to identify the pages corresponding to the desired date. For
systematic analysis across longer periods of time, this e ort is multiplied.</p>
        <p>To improve access, we want to identify the date and start and end point of
meetings, to label resolutions with those dates, so users can search and select
resolutions by date or period. This also allows scaling the analysis of resolutions.
E.g. how many resolutions were discussed on each day? How many were
accepted, rejected or postponed? When or on how many dates were certain topics
discussed? There are no standard NLP tools to help with this.</p>
        <p>To enable the various types of research methods and questions, we extract
and operationalise the following elements of resolutions as metadata layers:
Meetings and meeting dates The speci c date on which a proposition
discussed and a decision reached.</p>
        <p>Attendance lists and president The persons who were present and involved
in the decision making process of each resolution.</p>
        <p>Resolutions The type of proposition that was submitted, e.g. a request, report
or missive, who submitted the proposition, whether it was accepted, rejected
or postponed for later discussion, and what action was decided on.
Insertions Extracts of earlier resolutions or of resolutions by one of the
Provincial States, or of memorandums, letters or requests submitted to the SG.
Named entities Persons who submitted propositions and persons, committees
selected to investigate and report on issues that were discussed, and other
named entities such as organisations, geographic locations and ship names.
Topics The topic of individual resolutions. This is partly provided by the
contemporary indices, lists of index terms and marginalia. But some form of key
phrase extraction or topic modelling could provide alternative (and di
erently biased) topical perspectives.</p>
        <p>Before we describe how we use this model for extract information (Section 4),
it is necessary to discuss the numerous text recognition challenges this resource
puts forward, because the quality of text recognition has a big impact on the of
information extraction process that follows it.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Text Recognition Challenges</title>
      <p>The physical state of the records is heterogeneous. The corpus has a mix of
formats, with an ongoing record of handwritten resolutions, that from the 18th
century were also published as printed volumes. The collection spans 220 years
and has a large variety of handwriting, caused by di erences between the
successive clerks that were employed by the SG as well as by changes in modes of
handwriting in general, that changed from 16th-century Gothic to 18th-century
roman script. Some printed volumes have single column pages, but others have
double column pages and there are more complex column splits, insertions of
letters and extracts, marginalia, tables (including multi-column and multi-page
tables) and indices organised by main terms (referred to as respecten in the
corpus).</p>
      <p>The automatic recognition of the texts of the resolutions is performed by
OCR (Optical Character Recognition) and HTR (Handwritten Text
Recognition). We use a typical pipeline consisting of Layout Analysis and detection of
baselines of text in the images. During the project we continuously update the
OCR and HTR models using ground truth data sets and feedback from the
information extraction process.</p>
      <p>
        OCR is standard technology, but requires training because of diverging
printing formats and especially orthographic di culties. The current model achieves
a Character Error Rate (CER) of 8, that is, 8 out every 100 character are
incorrectly recognised. HTR is more di cult because of the irregularities of
handwriting. Recognising the 425,000 pages of handwritten resolutions requires several
steps in a pipe line. We used the P2PaLA tool [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] on the scans without ground
truth for layout analysis of text regions and text (base)lines. Next, we created
a manual transcription of ground truth data set of 1,000 pages. Through
iterative recognition and correction of batches of pages in Transkribus, we currently
achieve a CER of 2.99 on a 100 page evaluation set in which the identi ed text
regions and baselines were manually corrected. The model is ne-tuned by
corrections made by volunteers in the Vele Handen crowd-sourcing platform that
uses the web version of Transkribus.
      </p>
      <p>On handwritten pages with no layout corrections, we observed that
misidenti ed regions result in much more variation in the number of errors per page
and a higher CER. We continue to expand the ground truth set and plan to
train speci c models for pages diverging from a `standard' resolution page (e.g.
indices and tables) and for handwriting on which the HTR+-model performs
worse than the average.</p>
      <p>In the next section we discuss the results of the information extraction process
on the OCR output of the printed resolutions.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Information Extraction Challenges</title>
      <p>
        A typical step in extracting information from historical texts is to use general
approaches like Named Entity Recognition (NER), part-of-speech tagging, and
lemmatisation to identify entities and topical words and phrases [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This step
is thwarted by both text recognition errors and the lack of good NLP-resources
for historical spelling and vocabulary in early modern corpora [
        <xref ref-type="bibr" rid="ref10 ref13 ref4">4,13,10</xref>
        ].
      </p>
      <p>
        On English texts these generic approaches work to some extent. English
orthography has not changed much since the 18th century, therefore resources
for modern English can be e ective [
        <xref ref-type="bibr" rid="ref15 ref6">6,15</xref>
        ]. In Dutch, changes are larger, making
generic approaches less useful. We annotated named entities in 200 pages of
manually transcribed resolutions and retrained the Spacy NER tagger4 with
90% of the pages and tested on the remaining 10%. This led to a precision of
0.49 and recall of 0.19. Although this is likely to improve by annotating more
pages, there are two hurdles. First, the upper bound for precision and recall
remains low because in the resolution texts, many nouns have uppercase initials,
which makes it hard to algorithmically distinguish them from named entities.
Second, precision and recall will be signi cantly lower on the vast majority of
pages that are not manually but automatically transcribed.
      </p>
      <p>
        Moreover, such techniques do not alleviate the problem of identifying the
start and end of daily sessions in the text, the precise date of each session, the
attendance lists, the start and end of individual resolutions and the type of
decision reached. Generic approaches of layout analysis can detect standard
structures like tables, gures, footnotes, headers and tables of content with varying
levels of success [
        <xref ref-type="bibr" rid="ref1 ref3">3,1</xref>
        ], but cannot interpret speci c semantics such as temporal
orderings of meeting dates and the geographical ordering in the attendants lists.
      </p>
      <p>
        We decided on an alternative approach that is based on a combination of 1)
exploiting repetitive structural elements such as the layout and ordering used for
indices and attendants lists, similar to Colavizza et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], 2) explicitly modelling
domain knowledge in lists of formulaic textual phrases, and 3) approximate string
      </p>
      <sec id="sec-4-1">
        <title>4 https://spacy.io/models/nl</title>
        <p>Text string found</p>
        <p>Frequency
searching and matching. First, the sessions have a xed structure and layout,
with the opening of a next session and the attendance list represented in a
di erent font and text alignment than the resolution summaries. Second, the
resolution texts are extremely repetitive, using the same phrasings with little
variation across decades of resolution. We use phrase models that contain lists
of frequently occurring phrases, such as the standard formulas used to announce
the next resolution or its decision, or the start of a new meeting (see Figure 2).
Formulaic phrases can be short or long, e.g. a single word or an entire sentence
including punctuation. Each phrase has a metadata category and label so that
an approximate match in the text can be tied to a metadata layer. Phrases can
also have known variants, e.g. alternative phrasings that we have encountered.
The phrase models represent knowledge of the domain and the corpus, what
information we expect to nd, where and in what order.</p>
        <p>Third, our approach exploits the fact that, even with a relatively high CER,
the majority of characters in frequently occurring phrases and names are correct
and in the right order, such that the string distance between the recognised text
and its corresponding phrase in our model is small. We have developed a fuzzy
searching algorithm that accepts one or more phrase models to nd approximate
matches, and uses con gurable string distance thresholds to control how much
textual variation is accepted.5</p>
        <p>The phrase model for opening formulas contains eight di erent formulas,
each with a list of variant phrasings and a label for the type of proposition.
For instance, the opening formula `Ontvangen een Missive van' (EN: Received
a Missive of) indicates the proposition is a missive. For the resolutions of year
1705, 950 matches are found with 315 di erent OCR strings (Table 1). The most
frequent OCR string occurs 286 times, so identi es only 30% of the resolutions
for missives. Other proposition types include reports, requests, memories and
(previous) resolutions. The evaluation of this approach is described below.</p>
        <p>We exploit domain knowledge di erently across multiple iterations of
information extraction. In the rst iteration we focus on extracting information with
very high precision, by using high thresholds for approximate string searching, to</p>
      </sec>
      <sec id="sec-4-2">
        <title>5 See https://github.com/marijnkoolen/fuzzy-search.</title>
        <p>build lists of e.g. the starting point of meeting sessions and names of attendants.
In later iterations, we use additional domain knowledge. For instance, if we have
found the starting points of the meeting sessions for 12 and 15 January 1725
in the rst iteration, we exploit our knowledge that sessions are chronologically
ordered, therefore know that the sessions for 13 and 14 January should be in
between these starting points. In the second iteration, we can search with lower
thresholds in a much smaller amount of text and much higher chances of success.
For the names of attendants, we exploit our knowledge that the president, being a
representative of one of the provinces, rotated every week between the provinces,
and that these persons were regular attendants during the other weeks. Once
we know some of the names of presidents, we use approximate searching to nd
them in the lists as either president but with recognition errors, or as regular
attendants, reducing the number of unknown and uncertain names in the list.
For the individual resolutions, we use the fact that they use extremely
repetitive language, with typical textual formulas for the opening proposition (see
Figure 2), e.g. 'Ontvangen een missive van ...', (English: 'Received the missive
of ...') and the decision, e.g. `WAAR op gedelibereert zynde, is goedgevonden
ende verstaan ...' (English: 'WHICH after deliberation has been accepted ...').
Through multiple iterations of extracting resolutions, we gradually build lists
of standard formulas, which we subsequently use in approximate searching to
nd variations of these formulas that may contain both text recognition errors
and slightly di erent spellings. For all elements we extract, we store them with
the fuzzy matches of phrases as evidence to explain how our metadata was
created. Together with the explicit phrase models that we publish in our GitHub
repository,6 this makes the decision process transparent and repeatable.
4.1</p>
        <sec id="sec-4-2-1">
          <title>Evaluation</title>
          <p>We evaluated our approach with various ground truth data sets:
Page type identi cation Identifying whether a page contains resolutions,
index entries or respect, and whether a page is the title page of a section (and
therefore the start of a section). Our model uses a combination of layout
and textual evidence and has an accuracy of &gt; 0:99 on a test set of 3376
manually annotated pages. The printed volumes contain 91,302 pages with
resolutions, 10,698 index pages and 828 pages with term lists.</p>
          <p>Meeting date identi cation Identifying the start of a meeting and the date
of that meeting. We created ground truth data for 500 randomly selected
meeting dates between 1703 and 1796, and annotated the starting point in
the text as well as the exact date and day of the week. We evaluated and
updated the phrase model in two iterations, using a batch of 100 meeting
dates per iteration. Our current phrase model, after a third iteration of
updating the phrase model, leads to a precision of 0.99 and recall of 0.93
on the test set consisting of the remaining 300 meetings. The extraction</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>6 See https://github.com/HuygensING/republic-project</title>
        <p>algorithm detects the correct start for 100% of the extracted meetings, but
in three out of 300 cases the identi ed date is incorrect. For several other
dates in the test set, no meeting start is found. The 91,302 pages of printed
resolutions are thereby transformed into a new layer consisting of 23,605
meetings. Of the meeting dates that were not correctly identi ed (7% of the
total), the algorithm signals for 65% of them that they are appended to the
previous meeting, so we know where to focus manual e ort to correct them.
Resolution identi cation Identifying the individual resolutions, including their
opening proposition, the decision reached and the closing summary. Based
on the annotated ground-truth, there are an estimated 311,586 resolutions
in the printed volumes (13.2 per meeting). On the ground truth test set of
198 resolutions, our phrase model currently achieves a precision of 0.94 and
recall of 0.76 in identifying the opening phrases. The relatively low recall
signals that our phrase model is incomplete. In future iterations, we will
create additional ground truth for testing, so as not to over t our model on
the initial ground truth data.</p>
        <p>Attendants identi cation Identify the names of the attendants and link
recurring names to the correct entities. We are currently developing ground
truth data for this.</p>
        <sec id="sec-4-3-1">
          <title>Index entry and reference identi cation Identify the lemma and page ref</title>
          <p>erence of an entry, and link the lemma to the correct resolution. We have a
rst version of a model, but have not nished the ground truth data yet.</p>
          <p>Large historical resources all have their own textual characteristics and
structural features, which require the modelling of expert knowledge of these resources
and incorporating these into generic NLP techniques. To know if this phrase
model and fuzzy search approach generalises to other collections, we have
experimented successfully with extraction of the dates and nding locations of
medieval charters, such that over 17,000 extracted mentions of place names can
be treated as historically dated attestations. Our goal is to continue to develop
this open and reusable toolkit as an approach for structure-driven information
extraction of digitised resources for historical research.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Linking Challenges</title>
      <p>Information extraction gives us a way to navigate, select and order the the
meetings and resolutions in individual volumes through the di erent layers of
metadata. The next step is to connect those layers to each other and to connect
them across the hundreds of volumes. Connecting the metadata across di erent
layers enables queries like a) when was nancing of the military discussed, b)
who were involved in the decisions made around this topic, and c) what kind of
decisions were reached. Here, connections between the topics from the indices
and the resolutions they refer to, as well as to the correct dates of the resolutions
are indispensable. The other dimension is connecting metadata elements within
a single layer but across di erent years and meetings found in di erent volumes.</p>
      <p>For instance, to enable a good layer of topical metadata, we need to connect the
indices and marginalia across the entire period, so that all resolutions regarding
nancing of the military can be retrieved for all 220 years.</p>
      <p>The marginalia and contemporary indices were made by di erent people at
di erent times, resulting in an incoherent system over the years. Marginalia di er
in level of extensiveness, indices in level of completeness, and their individual
terms in level of scope and interpretation. This makes connecting them into a
coherent layer of information to access the entire archive a challenge.</p>
      <p>We manually transcribed and merged the lists of index entries for seven
volumes of resolutions between 1742 and 1785. These turned out to contain
3934 distinct entries, with individual volumes containing between 673 and 994
entries. Geographic locations and organisations have a higher overlap between
years and are more stable across time than persons. Subjects seem more stable
as well, but it is challenging to establish whether, for instance, the index entry
on `declaration' refers to the same thing over the years. The overlap of entries
between subsequent years is around 33%, but drops to 10-20% for indices that
are decades apart. This is mainly caused by the individual person entries, which
may recur in subsequent years but logically disappear from the index over time.
Yet subject terms also change, which creates a challenge for longitudinal analysis
of topics such as long-term developments of economic policies or shifts in the
accessibility of the meetings and decisions for di erent classes in society.</p>
      <p>Our preliminary analysis of the overlap in index terms across volumes shows
there are two challenges in creating a single index across the whole series. One
is to link recurring terms that might have di erent spellings or OCR
representations, the other is to link orthographically di erent but semantically similar
terms. For the former, we use our fuzzy search strategy. For the latter, some
combination of manual categorisation and automatic clustering will be used.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Presentation Challenges</title>
      <p>The size and prominence of the corpus o ers research opportunities for many
years to come, but its potential can only be reached if it is presented to end users
with relevant and easy-to-use functionality. The Resolutions need a reliable yet
exible platform that stays in active development to publish information when it
becomes available. One problem of custom build publication environments is the
risk of stalling maintenance and further development when project nances dry
up. When development stops, stability and exibility decrease. We use a generic
publication environment we call Docere. Several other projects are currently
being developed on the platform. Development is mostly mutually bene cial,
a generic feature realised for one project, will be instantly available to all the
others. On the one hand, this means the cost of development is split over projects.
On the other, individual projects have less in uence on development decisions,
as a decision for one project will directly a ect all other projects.</p>
      <p>To o er exibility per project some components of the user interface can be
custom built, for instance the page header and footer, which search facets are
available, or how search results are shown. Full customisability is available for
the rendering of the text. The OCR and HTR output are mapped to custom
components, which means all extracted text elements can have custom features
attached to them. A meeting or resolution is shown in relation to previous and
next meetings or resolutions. Related meetings or resolutions can be found via
topical metadata or using full-text search, giving the user extra possibilities to
navigate the set. Docere leverages a search index to create a faceted search. A
facet can represent metadata (date, president, attendee and others) or entity
(person, place, topic and others) from a meeting or resolution.</p>
      <p>In addition, it is possible to directly query the data set through an
Application Programming Interface (API). This is especially bene cial to the growing
number of researchers for whom the user interface is not su cient to answer
their research questions. The API will expose the raw and enriched data and
will be essential in adhering to the FAIR principles.7
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>One of the challenges in digitising large archives is to provide di erent dimensions
and levels of access, as their potential users come with many di erent questions,
di erent background knowledge and di erent and di erent needs to explore the
connections between the records. There are often several structural elements
in the analogue version that support these di erent information access needs,
but operationalising these structures and connecting them in the digital version
requires tackling a cascade of challenges.</p>
      <p>This paper describes our approach of modelling, recognising and extraction
these structural elements, dealing with problems of text recognition errors,
historical language variation and the heterogeneity in structure and content found
in long serial publications. Instead of using generic Natural Language Processing
in a one-shot information extraction pipeline, we developed an iterative approach
in which expert knowledge of the records can be incorporated in the extraction
process, and insights from analysing the output can be transparently modelled
and fed back into the process. Evaluation of our results shows that this not only
leads to highly accurate layers of structured text, annotations and metadata,
but also to interpretable models that can be used as provenance to explain how
each algorithm came to its decisions.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This research is funded by the The Netherlands Organisation for Scienti c
Research (NWO) through the NWO Groot project REPUBLIC (an acronym for
REsolutions PUBlished In a Computational Environment) 2019-2023.</p>
      <sec id="sec-8-1">
        <title>7 https://www.go-fair.org/fair-principles/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Clausner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antonacopoulos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pletschacher</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Icdar2019 competition on recognition of documents with complex layouts-rdcl2019</article-title>
          .
          <source>In: 2019 International Conference on Document Analysis and Recognition (ICDAR)</source>
          . pp.
          <volume>1521</volume>
          {
          <fpage>1526</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Colavizza</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ehrmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bortoluzzi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Index-driven digitization and indexation of historical archives</article-title>
          .
          <source>Frontiers in Digital Humanities</source>
          <volume>6</volume>
          ,
          <issue>4</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Doermann</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tombre</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , et al.:
          <article-title>Handbook of document image processing and recognition</article-title>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Eijnatten</surname>
            ,
            <given-names>J.v.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pieters</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verheul</surname>
          </string-name>
          , J.:
          <article-title>Big data for global history: The transformative promise of digital humanities</article-title>
          .
          <source>BMGN-Low Countries Historical Review</source>
          <volume>128</volume>
          (
          <issue>4</issue>
          ),
          <volume>55</volume>
          {
          <fpage>77</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Head</surname>
          </string-name>
          , R.:
          <article-title>Knowing like a state: the transformation of political knowledge in swiss archives</article-title>
          ,
          <volume>1450</volume>
          {
          <fpage>1770</fpage>
          .
          <source>The Journal of Modern History</source>
          <volume>75</volume>
          (
          <issue>4</issue>
          ),
          <volume>745</volume>
          {
          <fpage>782</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hengchen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Quantifying the impact of dirty ocr on historical text analysis: Eighteenth century collections online as a case study</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          <volume>34</volume>
          (
          <issue>4</issue>
          ),
          <volume>825</volume>
          {
          <fpage>843</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hoekstra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koolen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Data scopes for digital history research</article-title>
          .
          <source>Historical Methods: A Journal of Quantitative and Interdisciplinary History</source>
          <volume>52</volume>
          (
          <issue>2</issue>
          ),
          <volume>79</volume>
          {
          <fpage>94</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Japikse</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , et al:
          <article-title>Resolutien der Staten-Generaal 1576-1625</article-title>
          . Nijhof, Den Haag (
          <year>1915</year>
          -1994)
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jeurgens</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Schurende systemen: Seriearchieven in de digitale wereld</article-title>
          . In: Berende, H., van der Heiden,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Thomassen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Jeurgens</surname>
          </string-name>
          , C.,
          <string-name>
            <surname>van der Ven</surname>
          </string-name>
          , C.,
          <string-name>
            <surname>de Man</surname>
          </string-name>
          , H. (eds.)
          <article-title>Schetsboek digitale onderzoek-omgeving en dienstverlening: Van vraag naar experiment</article-title>
          , pp.
          <volume>54</volume>
          {
          <fpage>61</fpage>
          .
          <string-name>
            <surname>Stichting</surname>
            <given-names>Archiefpublicaties</given-names>
          </string-name>
          , 's-Gravenhage (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Leemans</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maks</surname>
          </string-name>
          , E., van der Zwaan, J.,
          <string-name>
            <surname>Kuijpers</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steenbergh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Mining embodied emotions: A comparative analysis of bodily emotion expressions in dutch theatre texts 1600-</article-title>
          <year>1800</year>
          '.
          <source>Digital Humanities Quarterly</source>
          <volume>11</volume>
          (
          <issue>4</issue>
          ) (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Meron</surname>
          </string-name>
          <article-title>~o-Pen~uela,</article-title>
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Ashkpour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Van Erp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Mandemakers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Breure</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Scharnhorst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Schlobach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Van Harmelen</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          :
          <article-title>Semantic technologies for historical research: A survey</article-title>
          .
          <source>Semantic Web</source>
          <volume>6</volume>
          (
          <issue>6</issue>
          ),
          <volume>539</volume>
          {
          <fpage>564</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Opitz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Born</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nastase</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Induction of a large-scale knowledge graph from the regesta imperii</article-title>
          .
          <source>In: Proceedings of the Second Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage</source>
          ,
          <source>Social Sciences, Humanities and Literature</source>
          . pp.
          <volume>159</volume>
          {
          <issue>168</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Piersma</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ribbens</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Digital historical research: Context, concepts and the need for re ection</article-title>
          .
          <source>BMGN-Low Countries Historical Review</source>
          <volume>128</volume>
          (
          <issue>4</issue>
          ),
          <volume>78</volume>
          {
          <fpage>102</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Quiros</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>P2pala: Page to page layout analysis toolkit</article-title>
          . https://github.com/ lquirosd/P2PaLA (
          <year>2017</year>
          ), gitHub repository
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. van Strien,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Beelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Ardanuy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.C.</given-names>
            ,
            <surname>Hosseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>McGillivray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Colavizza</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          :
          <article-title>Assessing the impact of ocr quality on downstream nlp tasks</article-title>
          .
          <source>In: ICAART (1)</source>
          . pp.
          <volume>484</volume>
          {
          <issue>496</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Upward</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reed</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliver</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Recordkeeping informatics for a networked age</article-title>
          . Monash University (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>