<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Sunken Ships Shan't Sail: Ontology Design for Reconstructing Events in the Dutch East India Company Archives</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stella Verkijk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Piek Vossen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Language, Literature and Communication, Vrije Universiteit Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Huygens Institute</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <fpage>320</fpage>
      <lpage>332</lpage>
      <abstract>
        <p>This short paper describes ongoing work on the design of an event ontology that supports state-ofthe-art event extraction in the archives of the Dutch East India Company (VOC). The ontology models Dynamic Events (actions or processes) and Static Events (states). By modelling the transition of a given to a new state as a logical implication that can be inferred automatically from the occurrence of a Dynamic Event, the ontology supports implied information extraction. It also considers implied sub-event detection and models event arguments as coreferential between event classes where possible. By doing so, it enables the extraction of much more information than is only explicitly stated in the archival texts with minimal annotation e昀ort. We de昀椀ne this complete event extraction task that adopts both Natural Language Processing techniques as well as reasoning components as Event Reconstruction. The Event Reconstruction module will be embedded in a search interface that facilitates historical research in the VOC archives.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Natural Language Processing</kwd>
        <kwd>event modelling</kwd>
        <kwd>computational history</kwd>
        <kwd>ontology design</kwd>
        <kwd>reasoning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>and detecting and classifying event arguments (who, what, where, etc.). The second step
involves the classi昀椀cation of latent events denoted by verbs, nominal expressions or nouns (e.g.,
friendship evokes a relationship event): a type of implicit information extraction. Finally, the
last step of ER deals with a deeper level of implied information extraction, where we extract
implied events that cannot be anchored directly to a mention in the text, but are the logical
consequence of a mentioned event. In our case this entails logical reasoning using an ontology to
extract implied states. Obvious consequences or assumptions are o昀琀en not reported explicitly
but need to be inferred to reconstruct a complete story (a mention of one geopolitical entity
losing a battle from another implies they were at war). Other information that is o昀琀en implied
in text that we also aim to make explicit is how multiple smaller events s(ubevents) sometimes
implicitly describe a relevant supra-event (e.g., when the same shipsleave a harbour and arrive
at a di昀erent harbour at the same time these subevents describe a voyage event).</p>
      <p>Figure 1 illustrates these three levels of information extraction applied on a sample sentence
of our corpus.1 It shows how the custom ontology enables the modeling of ship voyages and
their locations over large time frames. The aim of the ontology presented in this paper is thus
to allow inferencing of states orStatic Events (such as being at a location or being in con昀氀ict )
from extracted Dynamic Events stated explicitly or implicitly in the text (such asleaving a
location or attacking). This type of common sense reasoning can be classi昀椀ed as causal reasoning:
inferring logical implications of events.</p>
      <p>We provide a taxonomy of event classes de昀椀ned by specialist historians, covering a wide
range of researchable topics in the VOC archives. It serves as the basis for a more elaborate
event ontology that enables the detection of both explicit and implicit information needed for
event reconstruction (currently under development)2. To our knowledge, this paper describes
the 昀椀rst event ontology designed for a source dating as far back as the Early Modern period.
We describe a project that applies state-of-the-art automatic inferencing for the humanities,
presenting the 昀椀rst ontology made for the historical 昀椀eld that supports reasoning with implicit
information from unstructured text for comprehensive event reconstruction.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>A vast amount of research has been conducted on knowledge representation for NLP, resulting
in many ontologies and resources. However, none were made speci昀椀cally for our domain, and
little leverage commonsense reasoning over event causality. In this section, we discuss the
ontologies, lexical resources and theories most relevant to the current project. When assessing
re-usability, we consider what information is needed to perform event reconstruction on
(pre)colonial sources.</p>
      <p>
        Most historical event detection research has been conducted on datasets from the 20th
century [
        <xref ref-type="bibr" rid="ref20 ref3">20, 21, 6, 3</xref>
        ]. These frameworks are o昀琀en not directly applicable to themes that are highly
1The source text in this 昀椀gure is a comprehensive translation to English; for the original text and a literal translation
to English, see 8.1
2For a representation of the taxonomy in OWL seehttps://github.com/globalise-huygens/nlp-event-detection/tre
e/main/built_resources, for a WiKi describing each event in detail seehttps://github.com/globalise-huygens/nlp
-event-detection/wiki, and for our annotation guidelines seehttps://docs.google.com/document/d/1ZL9BbEmGky
1tJeTR-yvnU-bvqVfA14mVKOWlGgEJT0g/edit?usp=sharing
nd the process structure
ys
*)+,Shortcut
$#
$'#
&amp;%
(%&amp;
createaconnector
duplicateanobject
08
/67
9
#
(
4
/13
#
=&lt;
=&lt;
=&lt;
reasons, the individual ship names are not stated in the squares representing extracted information.
relevant for our source (ship movement; trade; geopolitical/social relationships; power
imbalances). More importantly, these historical event models do not model implied states (event
implications), which is necessary for detecting information not explicitly mentioned.
      </p>
      <p>
        ER deals with storyline plausibility. For a system to be able to understand stories (chains
of events) it should have knowledge of causal reasoning 1[
        <xref ref-type="bibr" rid="ref11">0, 11</xref>
        ]. In order to make correct
judgments on storyline plausibility a system must consider pre- and postconditions of a given
event (e.g., a ship that sinks at sea will never arrive at its destination). With the results that
state-of-the-art language models (LMs) are showing at several NLP tasks, it could be
questioned whether the integration of an extra module for common sense knowledge integration
that makes certain event causalities explicit is needed for implied event detection. However,
Qasemi et al. (2021) show that several state-of-the-art language models fail to reason with
preconditions of events [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Their results show a 10-30% gap between LMs’ and human
performance on three tasks evaluating the ability to understand situational preconditions. Storks
et al. (2021) conclude that LMs struggle to support predictions on storyline plausability with
valid supporting evidence [
        <xref ref-type="bibr" rid="ref21">22</xref>
        ]. Many more studies show that LMs struggle to manage logical
and causal reasoning [
        <xref ref-type="bibr" rid="ref24 ref8">7, 8, 25</xref>
        ]. For our project, we expect these gaps to be even bigger due to
the challenging nature of the historical text (see Section3). Therefore, we want to rely on an
ontology to 昀椀ll these gaps. A hybrid approach to event implication detection, combining the
strong associations captured in an LM (to bridge a possible recall gap) and an ontology as a
checking and addition module (to bridge a precision gap), is also a possible line of research.
      </p>
      <p>
        Since the ontology supports an ongoing project, we aim to create a 昀氀exible and
comprehensive ontology to which event classes can be added in an iterative manner. Comprehensibility
is also important to guard annotation quality and modeling possibilities. General-purpose
ontologies like SUMO [13] and FrameNet [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] aim to model the whole world at a very high level
of granularity, making them less 昀氀exible and comprehensible. Also, SUMO distinguishes
between objects and processes but not between static and dynamic events. Therefore, although
SUMO allows for reasoning, it does not formalize logical implications of events or processes as
new states. FrameNet does not model causality or implications of events at all.
      </p>
      <p>
        The Rich Event Ontology aims to integrate lexical and ontological resources1[] and also
models causal relations between events [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Another aspect of the REO that is highly relevant
to ER is the di昀erentiation between hasSubclass and hasSubevent, the latter capturing how
subevents are temporally contained within their supra-class. We can draw on many principles
of the REO, however, they do not formally de昀椀ne how dynamic events logically imply new
states.
      </p>
      <p>
        Modeling causal relationships between events ontologically has been investigated
extensively by Segers [
        <xref ref-type="bibr" rid="ref16 ref17 ref18 ref19">19, 18, 16, 17</xref>
        ]. Segers o昀ers di昀erent versions of a detailed ontology (the
Circumstantial Event Ontology (CEO)) that models pre-, during- and post-conditions of calamity
events and the roles of the entities a昀ected by the event. By doing so, the ontology
represents implied causal relations between event classes. This framework served as a basis of our
ontology.
      </p>
      <p>A key aspect of event reconstruction is modelling the participants of an event, o昀琀en referred
to as semantic roles or event arguments. Li et al. (2022) show how event prediction can be
informed by implied information extraction9[], taking semantic roles into account. They utilize
graph schema induction as a means of predicting new nodes that represent future events. They
employ a copy mechanism to generate coreferential arguments (for example, theDetainee
argument is the Attacker of a previous Attack event).</p>
      <p>
        We adopt this idea and rely on PropBank 1[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for the linguistic argumentation behind this
type of participant modelling. PropBank, di昀erent from the semantics-central approach of
FrameNet, is based on syntax. The roles in PropBank are recyclable in the sense that they would
椀昀t in any sentence with any mentioned event. For almost any event, a Patient can be de昀椀ned
(similar to ARG1 in PropBank) and in most events an Agent can be de昀椀ned (similar to ARG0).
The downsides of PropBank for our purposes are threefold: i) it is a sentence-based approach,
which is not desirable for our project because of how information is presented in the archival
material (Section 3), ii) it is lexical and does not generalize over variants or synonyms that
express the same event and, iii) the PropBank roles are not consistent enough to recycle among
event classes.3 While PropBank arguments rely on the syntactic structure of a sentence, which
is linked to the transitivity of a speci昀椀c verb, we want to rely on the semantic implications for
an argument, the role it plays in an action and how that role can be translated to the role it
plays in an inferred situation.
3For example, the location of an event sometimes takes up the ARG1 slot, sometimes ARG2, etcetera, depending on
the syntactic structure a predicate (event) dictates in a sentence
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <p>The project in which our ER module is embedded focuses on a subset of the VOC archives,
namely the Overgekomen Brieven en Papieren (‘Received Letters and Papers’, OBP) and within
that the Generale Missiven (‘General Missives’). These are documents written by the governor
general and council of Batavia summarizing the status of, for example, trade and con昀氀ict in Asia.
They are written in Early Modern Dutch, which di昀ers from contemporary Dutch with respect
to lexical variation, spelling, style, and grammar. Most noteworthy is the frequent occurrence
of very long-range dependencies, with pronouns sometimes referring to an entity several
sentences or even pages back. Imperfect automatic transcriptions of handwritten material result
in even more variation. Given the nature of the historical data and the lack of training data, an
ontology can have a strong added value to guide the interpretation of any models that need to
process these texts and reconstruct events. See8.2 for some corpus characteristics.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Ontology Description</title>
      <sec id="sec-4-1">
        <title>4.1. Ontology Requirements</title>
        <p>The design of the ontology is based on six requirements following from the interdisciplinary
collaboration between NLP scientists and historians and the needs for an ER system.</p>
        <p>R1 Domain-speci昀椀c event classes</p>
        <p>The ontology should contain event classes that historians have indicated to be relevant
in this speci昀椀c corpus, and the classes should be de昀椀ned by historians.</p>
        <p>R2 Scalar and binary change</p>
        <p>The ontology must accommodate for scalar change as well as binary change. This
enables, for example, the modeling of price changes or the conditions of ships over time
through ER.</p>
        <p>R3 Implied event detection and reasoning</p>
        <p>The ontology must accommodate automatic reasoning over events in a way that enables
it to extract implied states and events.</p>
        <p>R4 Flexibility</p>
        <p>The ontology must have a (taxonomic) structure that allows for the addition of events in
the future through an iterative process.</p>
        <p>R5 Comprehensibility</p>
        <p>The ontology must remain as comprehensible and transparent as possible in order to
guard both annotation quality and modeling possibilities, taking into account the
complexity and noisiness of the source data. The ontology should serve to reduce annotation
e昀ort where possible.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Event classes and ontology structure</title>
        <p>A昀琀er considering several linguistic frameworks and ontologies, we opted to take building
blocks from di昀erent resources to create an ontology that is as robust as possible for our dataset
and research purposes. The CEO served as a basis for the de昀椀nition of event classes because it
focuses less on how events are semantically manifested in text (like FrameNet), and more on the
practical notion that dynamic events o昀琀en imply a static event. However, many event classes
had to be adapted, deleted or added, since the CEO was built for processing modern new4s.
We copied the notion that tokens can be annotated with semantic roles cross-sententially from
FrameNet (as opposed to PropBank where the whole syntactic structure of one sentence should
be represented in the annotation). We took the notion that event arguments do not have to be
re-de昀椀ned for each event class speci昀椀cally from PropBank. We de昀椀ned our own limited set of
argument types in order to guard annotation quality and modeling purposes.</p>
        <p>The present ontology contains events that can be categorised under three themes deemed
relevant by historians: ship movement, trade, and (geo)political/social relations. The event
classes are de昀椀ned so that they mainly represent observable events, steering clear of concepts
that have an inherently subjective character. For example, whether an action is legal or illegal
depends on the context, which should be studied 昀椀rstly by a historian. Which events can
be classi昀椀ed as observable and which as subjective is an open and ongoing interdisciplinary
discussion. Since there also exist causal relations that are debatable or subjective2[4], we only
aim to model logical implications of events (which means only modelling necessary and not
possible consequences of dynamic events).</p>
        <p>For an overview of the current taxonomy of dynamic events, see Figure 2. Most dynamic
events automatically imply the transition from a given to a new static event. The static events
are modeled as event classes with their own arguments. This allows us to annotate and extract
them at the same complexity level as dynamic events. For a sample of the graph indicating
how event classes relate, see Figure 3. Our Wiki describes all event classes5.</p>
        <p>
          Previous research [
          <xref ref-type="bibr" rid="ref22">23</xref>
          ] shows that events are o昀琀en described as relative changes on a scale
rather than in absolute values. A scalar model allows for capturing and reasoning over such
relative and under-speci昀椀ed reports. We solve the inclusion of scalar changes without
having to annotate and model the absolute value of change by including events likeIncreasing,
Decreasing, Repairing and Damaging and linking them to states likeHavingInternalState+ and
HavingInternalState-.6
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Event participant modelling</title>
        <p>FrameNet and the CEO de昀椀ne participants speci昀椀c to each event class. In order to be able to
model actors of implied changes of state, semantic roles have to be recyclable from one event
class to another (if necessary with minimal conversion rules). That is to say, we want to infer
that the Agent in an Attacking event is a Patient in a BeingInCon昀氀ict state. FrameNet shi昀琀s
the semantics from the event class to the role, which limits the generalization across roles.
The CEO o昀ers a solution to this with the incorporation of assertion rules. The goal for our
ontology is to have even more intuitive and general roles than in the CEO because the textual
data we are working with is extremely noisy and information about one event is stretched
out over large portions of text (see Section3). By adopting more intuitive and general roles,
4The 昀椀nal ontology will link classes to the CEO, SUMO, FrameNet and other relevant ontologies with SKOS relations
5https://github.com/globalise-huygens/nlp-event-detection/wiki
6A ship being repaired impliesHavignInternalState+ and a deteriorated city
impliesHavingInternalStatewe simplify the annotation process as much as possible, which will lead to more consistent
annotation. Also, we suspect that the de昀椀nition of more general semantic roles might be more
informative for a machine learning system that needs to generalize on limited training data.</p>
        <p>Table 1 gives an overview of our arguments. Participants can be animate and non-animate.
The AgentPatient role is speci昀椀cally de昀椀ned for events in which a causer is simultaneously
undergoing the event SocialInteraction, BeginningARelationship, EndingAContractualAgreement.
For now, Time can be anything from a date to an hour to a day in the week. Source and Target
are de昀椀ned for Translocation events that may indicate both the starting and ending point of a
movement. Roles are de昀椀ned per event in our Wiki. 7. In our ontology, event participants are
represented as properties: i.e. event classes have properties likehasAgent, hasPatient.</p>
        <p>The roles can sometimes be recycled directly from a dynamic event class to their supra-class
as well as the static class they change, and sometimes they need conversion rules. See Table
2 for an example of how roles are translated between classes. It shows how the Bene昀椀ciary of
a Selling event (the person to whom is sold) is an Agent in the static eventHavingInPossession,
and it can be deducted that this Agent has the Patient of theSelling event (the thing sold) in
possession from the moment the event happened onward.
7https://github.com/globalise-huygens/nlp-event-detection/wiki</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Future steps: Operationalization</title>
      <p>An annotation pilot has been conducted and the ontology and event guidelines have been
updated accordingly. We plan to enhance the ontology in this iterative manner over a few
annotation rounds. When the ontology and guidelines are stabilized, we will run consecutive
annotations through the pipeline presented in Figure 4. The resulting Knowledge Graph (KG)
is to be integrated in the complete KG supporting the online infrastructure for historical
research.</p>
      <p>When representing information extracted from text it is important to be transparent about
provenance [4]. For each event triple represented in the KG, we will show whether this triple
was retrieved through human annotation or through the NLP pipeline. The GRaSP model5[]
allows for provenance tracing so that every triple is part of a claim by a source.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This paper has o昀ered an overview of considerations that arise when designing an event
ontology for a speci昀椀c domain and a speci昀椀c purpose within an interdisciplinary e昀ort, relying,
where possible, on existing resources. We make clear that event classes should be de昀椀ned
AP
AP
A
P
A
P
A
P
A
B
P
A
B
P
A
B</p>
      <p>P</p>
      <sec id="sec-6-1">
        <title>BeingInConflict</title>
      </sec>
      <sec id="sec-6-2">
        <title>BeingInConflict</title>
      </sec>
      <sec id="sec-6-3">
        <title>Unrest</title>
      </sec>
      <sec id="sec-6-4">
        <title>Unrest</title>
      </sec>
      <sec id="sec-6-5">
        <title>HavingInPossession</title>
      </sec>
      <sec id="sec-6-6">
        <title>HavingInPossession</title>
      </sec>
      <sec id="sec-6-7">
        <title>HavingInPossession</title>
        <p>Ontological 
Reasoning</p>
        <p>Event enrichment
module: inferring
(Static) Events
Provenance tracking
P1
P2
P1
P2
P1
P2
P1
P2
A1
A2
P
A1
A2
P
A2
A1
P</p>
      </sec>
      <sec id="sec-6-8">
        <title>P1 inConflict with P2</title>
      </sec>
      <sec id="sec-6-9">
        <title>P2 inConflict with P1</title>
      </sec>
      <sec id="sec-6-10">
        <title>P1 inConflict with P2</title>
      </sec>
      <sec id="sec-6-11">
        <title>P2 inConflict with P1</title>
      </sec>
      <sec id="sec-6-12">
        <title>Unrest between P1 and P2</title>
      </sec>
      <sec id="sec-6-13">
        <title>Unrest between P1 and P2</title>
      </sec>
      <sec id="sec-6-14">
        <title>Unrest between P1 and P2</title>
      </sec>
      <sec id="sec-6-15">
        <title>Unrest between P1 and P2</title>
      </sec>
      <sec id="sec-6-16">
        <title>A1 notHasInPossession P</title>
      </sec>
      <sec id="sec-6-17">
        <title>A2 hasInPossession P</title>
      </sec>
      <sec id="sec-6-18">
        <title>A1 notHasInPossession P</title>
      </sec>
      <sec id="sec-6-19">
        <title>A2 hasInPossession P</title>
      </sec>
      <sec id="sec-6-20">
        <title>A2 hasInPossession P A1 notHasInPossession P</title>
        <p>according to the characteristics of the dataset and needs of the 昀椀nal product: in our case
de昀椀ning events speci昀椀c enough to the corpus but general enough to leave space for historical and
contextual interpretation. We distinguish between event extraction and event reconstruction,
de昀椀ning the latter as a combination of several NLP-tasks and reasoning. We propose an
ontological method of extracting implicit information and incorporating it in an event
reconstruction pipeline.</p>
        <p>By focusing on this case-study of event reconstruction in Early Modern Dutch documents we
expect to come to new insights about how to push the state-of-the-art in event extraction (using
ontologies) forward. Our data and ontology can serve as a robust use-case for investigating to
what extend LMs need external common sense knowledge integration to successfully perform
certain NLP tasks. Finally, we aim to facilitate a new way of doing historical research with the
so昀琀ware supported by the presented ontology.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Acknowledgements</title>
      <p>This research is part of the GLOBALISE project, funded by the Dutch Research Council (NWO)
under project number 175.2019.003. We thank Kay Pepping, dr. Lodewijk Petram, dr. Pia
Sommerauer and dr. Micky Cornelissen for their help revising the paper, as well dr. Manjusha
Kuruppath, Maartje Hids, Henrike Vellinga and Brecht Nijman for their involvement in the
annotation pilot.
[21] R. Sprugnoli and S. Tonelli. “One, no one and one hundred thousand events: De昀椀ning and
processing events in an inter-disciplinary perspective”. In:Natural language engineering
23.4 (2017), pp. 485–506. doi: https://doi.org/10.1017/S1351324916000292.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Appendices</title>
      <sec id="sec-8-1">
        <title>8.1. Additional examples</title>
        <p>• Original source
‘Tsedert onsen Jongsten (hiernevens in copie gaande) met de schepen Frederick Hendrik,
Hollandia, ’t Wapen van Del昀, ’s landts Hollandia, ende de Galias den 12en November
passato in compagnie van Bataviagescheyden ende den 15en ditto door de Strate Sunda
geraeckt, sijn hier, Godtlo昀, den 17en ende den 21en van Teyouhan ende Jappan
successive wel aengecomen ’t jacht Erasmus ende ’t schip de Vreede, t’samen geladen 810 picol
rouwe Chineesche syde ende 332 picol coper8.’
• Literal translation</p>
        <p>Since our youngest (herewith in copy) with the ships Frederick Hendrik Hollandia, ’t
Wapen van Del昀, ’s landts Hollandia, and de Galias on the 12th of November passato in
company from Bataviaseperated and the 15th of the same have made its way through the
Sunda Strait, did here, God bless, on the 17th and the 21st from Teyouhan and Jappan
successively arrive ’t yacht Erasmus and the ship de Vreede, together loaded 810 picul
raw Chinese silk and 332 picul copper.</p>
      </sec>
      <sec id="sec-8-2">
        <title>8.2. Corpus characteristics</title>
        <p>8‘General Missive of 6-1-1628’, National Archive, The Hague, The Netherlands, 1.04.02 (Archive of the VOC),
inventory no. 1092, folio 1, r.</p>
        <sec id="sec-8-2-1">
          <title>Name Period</title>
        </sec>
        <sec id="sec-8-2-2">
          <title>General Missives 1618-1792 OBP 1610-1796</title>
        </sec>
        <sec id="sec-8-2-3">
          <title>Description</title>
        </sec>
        <sec id="sec-8-2-4">
          <title>Narrative reports</title>
          <p>from Council of India (Batavia)
to Gentlemen Seventeen</p>
          <p>(Dutch Republic).</p>
        </sec>
        <sec id="sec-8-2-5">
          <title>O昀琀en ordered by region;</title>
          <p>small summaries in margins</p>
        </sec>
        <sec id="sec-8-2-6">
          <title>Collection of General Missives and varied documents on which these missives are based.</title>
        </sec>
        <sec id="sec-8-2-7">
          <title>Number &amp; size</title>
          <p>Avg doc length
923 documents;
191.725 handwritten
pages
c. 250,000 documents;
c. 7 million
handwritten pages
c. 28 pages
(rough estimate)</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bonial</surname>
          </string-name>
          , S. Brown, M. Palmer, and
          <string-name>
            <surname>G. Kazeminejad.</surname>
          </string-name>
          “The Rich Event Ontology:
          <article-title>Ontological Hub for Event Representations”</article-title>
          .
          <source>In:Computational Analysis of Storylines: Making Sense of Events</source>
          (
          <year>2021</year>
          ), pp.
          <fpage>47</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. W.</given-names>
            <surname>Brown</surname>
          </string-name>
          , C. Bonial,
          <string-name>
            <given-names>L.</given-names>
            <surname>Obrst</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmer</surname>
          </string-name>
          . “
          <article-title>The rich event ontology”</article-title>
          .
          <source>InP:roceedings of the Events and Stories in the News Workshop</source>
          . Vancouver, Canada,
          <year>2017</year>
          , pp.
          <fpage>87</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[3] [4] [5]</source>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Cybulska</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          . “
          <article-title>Historical event extraction from text”</article-title>
          .
          <source>In:Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities</source>
          . Portland,
          <string-name>
            <surname>OR</surname>
          </string-name>
          , USA,
          <year>2011</year>
          , pp.
          <fpage>39</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Fokkens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Ter</given-names>
            <surname>Braake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ockeloen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Legêne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          , et al. “
          <article-title>BiographyNet: Methodological Issues when NLP supports historical research”</article-title>
          .
          <source>In:Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC'14)</source>
          . Reykjavik, Iceland,
          <year>2014</year>
          , pp.
          <fpage>3728</fpage>
          -
          <lpage>3735</lpage>
          . doi: https://doi.org/10.1162/coli\_a\_
          <volume>00347</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Fokkens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rospocher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoekstra</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. R. van Hage</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F. B.</given-names>
            <surname>Kessler</surname>
          </string-name>
          . “Grasp:
          <article-title>Grounded representation and source perspective”</article-title>
          .
          <source>In:Proceedings of the Workshop Knowledge Resources for the Socio-Economic Sciences and Humanities associated with RANLP. Varna, Bulgaria</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>N.</given-names>
            <surname>Ide</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Woolner</surname>
          </string-name>
          . “
          <article-title>Historical ontologies”. InW:ords and Intelligence II: Essays in Honor of Yorick Wilks (</article-title>
          <year>2007</year>
          ), pp.
          <fpage>137</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>O.</given-names>
            <surname>Ignat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abzaliev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Biester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gunal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kazemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khalifa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Koh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          , D. June Min,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nwatu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Perez-Rosas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          . “
          <article-title>A PhD Student's Perspective on Research in NLP in the Era of Very Large Language Models”</article-title>
          .
          <source>In:arXiv preprint arXiv:2305.12544</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lalwani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Vaidhya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sachan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schölkopf</surname>
          </string-name>
          . “
          <article-title>Logical fallacy detection”</article-title>
          .
          <source>In:arXiv preprint arXiv:2202.13758</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ji</surname>
          </string-name>
          , J. Han, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Voss</surname>
          </string-name>
          . “
          <article-title>The future is not onedimensional: Complex event schema induction by graph modeling for event prediction”</article-title>
          .
          <source>In: arXiv preprint arXiv:2104.06344</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ilievski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Francis</surname>
          </string-name>
          , E. Nyberg,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Oltramari</surname>
          </string-name>
          . “
          <article-title>Coalescing Global and Local Information for Procedural Text Understanding”</article-title>
          .
          <source>In:arXiv preprint arXiv:2208.12848</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Mirza</surname>
          </string-name>
          . “
          <article-title>Event Causality”</article-title>
          .
          <source>In: Computational Analysis of Storylines: Making Sense of Events</source>
          <volume>106</volume>
          (
          <year>2021</year>
          ), pp.
          <fpage>106</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12] [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gildea</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Kingsbury</surname>
          </string-name>
          . “
          <article-title>The proposition bank: An annotated corpus of semantic roles”</article-title>
          .
          <source>In: Computational linguistics 31.1</source>
          (
          <issue>2005</issue>
          ), pp.
          <fpage>71</fpage>
          -
          <lpage>106</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Pease</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Niles</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          . “
          <article-title>The suggested upper merged ontology: A large ontology for the semantic web and its applications”</article-title>
          .
          <source>In:Working notes of the AAAI-2002 workshop on ontologies and the semantic web</source>
          . Vol.
          <volume>28</volume>
          .
          <year>2002</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E.</given-names>
            <surname>Qasemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ilievski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Szekely</surname>
          </string-name>
          . “Paco:
          <article-title>Preconditions attributed to commonsense knowledge”</article-title>
          .
          <source>In:arXiv preprint arXiv:2104.08712</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ruppenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ellsworth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schwarzer-Petruck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Sche昀czyk. FrameNet</surname>
          </string-name>
          <string-name>
            <surname>II</surname>
          </string-name>
          :
          <article-title>Extended theory and practice</article-title>
          .
          <source>Tech. rep. International Computer Science Institute</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Segers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Caselli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          . “
          <article-title>The circumstantial event ontology (CEO)”</article-title>
          .
          <source>In:Proceedings of the Events and Stories in the News Workshop</source>
          . Vancouver, Canada,
          <year>2017</year>
          , pp.
          <fpage>37</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>R.</given-names>
            <surname>Segers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Caselli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          . “
          <article-title>The Circumstantial Event Ontology (CEO) and ECB+/CEO; an Ontology and Corpus for Implicit Causal Relations between Events”</article-title>
          .
          <source>In: Proceedings of the 11th International Conference on Language Resources and Evaluation (LREC'18)</source>
          . Miyazaki, Japan,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>R.</given-names>
            <surname>Segers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rospocher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          , E. Laparra,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Rigau, and</article-title>
          <string-name>
            <given-names>A.-L.</given-names>
            <surname>Minard</surname>
          </string-name>
          . “
          <article-title>The event and implied situation ontology (eso): Application and evaluation”</article-title>
          .
          <source>InP:roceedings of the 10th International Conference on Language Resources and Evaluation (LREC'16)</source>
          .
          <year>2016</year>
          , pp.
          <fpage>1463</fpage>
          -
          <lpage>1470</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R.</given-names>
            <surname>Segers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          , E. Laparra, G. Rigau,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rospocher</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.-L.</given-names>
            <surname>Minard</surname>
          </string-name>
          . “
          <article-title>The Event and Implied Situation Ontology (ESO)”</article-title>
          .
          <source>In:Clin26</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Tonelli</surname>
          </string-name>
          . “
          <article-title>Novel event detection and classi昀椀cation for historical texts”</article-title>
          .
          <source>In: Computational Linguistics 45.2</source>
          (
          <issue>2019</issue>
          ), pp.
          <fpage>229</fpage>
          -
          <lpage>265</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Storks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Chai</surname>
          </string-name>
          . “
          <article-title>Tiered reasoning for intuitive physics: Toward veri昀椀able commonsense language understanding”</article-title>
          .
          <source>In: arXiv preprint arXiv:2109.04947</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Agerri</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Aldabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cybulska</surname>
          </string-name>
          , M. van
          <string-name>
            <surname>Erp</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Fokkens</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Laparra</surname>
            ,
            <given-names>A.-L.</given-names>
          </string-name>
          <string-name>
            <surname>Minard</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          <string-name>
            <surname>Aprosio</surname>
          </string-name>
          , and G. Riga. “
          <article-title>NewsReader: using knowledge resources in a crosslingual reading machine to generate more knowledge from massive streams of news”</article-title>
          .
          <source>In: Knowledge-Based Systems</source>
          <volume>110</volume>
          (
          <year>2016</year>
          ), pp.
          <fpage>60</fpage>
          -
          <lpage>85</lpage>
          . doi: https://doi.org/10.1016/j.
          <source>knosys.20 16.07</source>
          .013.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>P.</given-names>
            <surname>Vossen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Caselli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Segers</surname>
          </string-name>
          .
          <article-title>“A Narratology-Based Framework for Storyline Extraction”</article-title>
          .
          <source>In: Computational Analysis of Storylines: Making Sense of Events</source>
          <volume>125</volume>
          (
          <year>2021</year>
          ), pp.
          <fpage>125</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>M.</given-names>
            <surname>Willig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>ZEČEVIĆ</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Dhami</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Kersting. “Causal Parrots: Large Language Models May Talk Causality But Are Not Causal</surname>
          </string-name>
          <article-title>”</article-title>
          . In:preprint (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>