<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Participation in the KONVENS 2021 Shared Task on Scene Segmentation Using Temporal, Spatial and Entity Feature Vectors</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Florian Barth, Tillmann D o ̈nicke G o ̈ttingen Centre for Digital Humanities University of G o ̈ttingen</institution>
        </aff>
      </contrib-group>
      <fpage>35</fpage>
      <lpage>41</lpage>
      <abstract>
        <p>This paper describes the team's efforts in solving the KONVENS 2021 Shared Task on Scene Segmentation. It presents a statistical approach and puts a focus on the design of feature vectors that cover the key criteria for scene boundaries, namely the change of time, space, and/or entities between two scenes. Combining our feature set with a random forest classifier achieves micro-averaged F1's of 0.07 (in-domain) and 0.12 (off-domain), which puts our system in third place (out of five) in the shared task but does not improve over the performance of the neural model previously published by the organisers (Zehe et al., 2021a). Nevertheless, we think that handcrafted features can, in combination with distributional embeddings, improve the task of scene segmentation and this paper might inspire future work in this direction.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The analysis of narratological phenomena is a
major field of research within computational
literary studies and focuses on aspects like time
        <xref ref-type="bibr" rid="ref6">(Kearns, 2020)</xref>
        , space
        <xref ref-type="bibr" rid="ref10">(Pustejovsky et al., 2015)</xref>
        ,
narrative levels
        <xref ref-type="bibr" rid="ref11">(Reiter et al., 2019)</xref>
        as well
as perspective or involvement of the narrator
        <xref ref-type="bibr" rid="ref3">(Eisenberg and Finlayson, 2016)</xref>
        . These tasks
typically involve a critical reflection on the
theoretical background in conjunction with an
extended refinement of annotation guidelines
        <xref ref-type="bibr" rid="ref1">(cf.
B o¨gel et al. 2015)</xref>
        , which leads to complex
categories for which annotators mostly achieve
moderate to substantial agreement.
      </p>
      <p>
        The current task focuses on the detection of
narrative scenes
        <xref ref-type="bibr" rid="ref13 ref14">(Zehe et al., 2021b)</xref>
        . The concept
introduced by
        <xref ref-type="bibr" rid="ref4">Genette (1983</xref>
        ) describes a strong
relationship or even equality of narrated time (time
of discours) and story time (time of histoire) as an
exclusive criterion for the definition of a scene.
trial
train
eval 1
eval 2
texts
sents
      </p>
      <p>Zehe et al. (2021a) expand the definition to 4
criteria that constitute scenes: the equality or
continuity of 1) time and 2) space, 3) the centrality
of a specific action, and 4) a constant character
constellation. In this paper, we present a statistical
classifier that is based on a feature design including
each of these criteria for scenes. This allows
an analysis of feature importance and a better
understanding of how the defining criteria are
processed by the learning algorithm.</p>
      <p>The observations on scene detection in fictional
texts can furthermore contribute to similar tasks in
other textual domains such as news stories, or it
might serve as a basis for an in-depth understanding
of the plot structure within a narration.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>The shared task data consists of 27 novels from
which 20 serve as train set, 1 was given as trial
data, and 6 as test set. The latter is split into
two evaluation tracks: the first consists of 4
dime novels and the second contains 2 novels
of contemporary high literature. All training
and trial texts are dime novels. On average,
each text of the training data consists of 35,801
tokens (standard deviation: 10,425) and 2,823
sentences (standard deviation: 550). The target
classes of the task are the boundaries between
scenes and/or nonscenes. Therefore, 3 different
classes exist: Scene-to-Scene, Scene-to-Nonscene,
and Nonscene-to-Scene (Nonscene-to-Nonscene is
excluded by definition). The proportion of scenes
within the training data is considerably higher,
which is why the 2 classes involving nonscenes
are rather underrepresented (cf. Table 1).
3</p>
    </sec>
    <sec id="sec-3">
      <title>Preprocessing</title>
      <p>
        We preprocessed the already sentencised texts
with spaCy1. We used its default tokenizer and
lemmatizer for German and added several custom
preprocessing components. First, we added the
Universal Dependency parser, morphological
analyzer, clausizer and tense–mood–voice–
modality tagger from Do¨nicke (2020). Second,
we added a direct speech tagger that recognises
text between opening and closing quotation
marks. Third, we added a coreference resolver
based on the algorithm from Krug et al. (2015),
which we extended to create coreference clusters
for all noun phrases in a text and not only
character mentions. Fourth, we added a temponym
tagger that recognises and normalises temporal
expressions using regular expressions. For this, we
used the German resource files from HeidelTime
        <xref ref-type="bibr" rid="ref12">(Stro¨tgen and Gertz, 2010)</xref>
        2. And fifth, we added
a verb tagger that assigns Levin (1995)’s verb
categories from GermaNet
        <xref ref-type="bibr" rid="ref5">(Hamp and Feldweg,
1997)</xref>
        to the verb of the matrix clause, based on a
disambiguation with respect to synset distances of
verb–subject and verb–object(s).
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Features</title>
      <p>We extract features from different syntactic units in
a sentence: from the sentence itself, from clauses,
from noun phrases, and from tokens. When
vectorising a document D = (s1, . . . , sn), we
get sentence vectors (~s1, . . . , ~sN ), which we then
concatenate to context-sensitive vectors XD =
(~x1, . . . , ~xn) using a window of 5 sentences: ~xi :=
~si−2 ◦ . . . ◦ ~si+2. The following subsection briefly
describes the features which we extract from each
sentence.
4.1</p>
      <sec id="sec-4-1">
        <title>General Features</title>
        <p>General features should catch structural markers for
scene boundaries, e.g. (changes in) grammatical
1https://spacy.io/
2https://github.com/HeidelTime/
heideltime/tree/master/resources/german
features such as tense and aspect, direct speech,
presence of punctuation, and discourse connectives
(usually the first word of a sentence).</p>
        <p>From the matrix clause, we extract: tense (fut/
past/pres), aspect (imperf/perf), mood (imp/ind/
subj:past/subj:pres), voice (active/pass:dynamic/
pass:static), and the lemmas of modal verbs
(ko¨nnen/mu¨ssen/wollen/...); whether it is inside
direct speech (no/yes); and for all verbs, the
part of speech (AUX/VERB) and Levin category
(Communication/Cognition/...).</p>
        <p>From subordinate clauses, we extract: the
root dependency relation (acl/ccomp/csubj/...);
whether it is inside direct speech; and for all verbs,
the Levin category.</p>
        <p>From the sentence, we extract: the lemmas of
punctuation tokens (,/;/*/...); the lemma of the
last token if it is a punctuation token; whether
the first, last or any other token is a space token
(e.g. an empty line)3; the part of speech (PRON/
SCONJ/...) and dependency relation (mark/nsubj/
...) of the first non-punctuation, non-space token;
and the number of tokens and the number of
clauses.</p>
        <p>Because we had the impression that nonscenes
increasingly occur in the beginnings and ends
of texts, we also added the sentence’s index as
feature, once counted from the start and once
counted from the end. To prevent the system from
simply memorising the tags for all indices, we set
the feature to 10 for indices greater or equal to 10.4
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Temporal Features</title>
        <p>Temporal expressions are recognised and
normalised by the temponym tagger. We split
the norm value of a temporal expression at
dashes and camel-cased letters into substrings that
we use as substring features. For example, the
expression [am] na¨chsten Tag ‘[the] next day’ is
normalised as UNDEF-next-day which is split into
{UNDEF, next, day}.</p>
        <p>For all temporal expressions, the temponym
tagger further returns a type (date/duration/
interval/set/time), and optionally a modifier (END/
3We included this feature because headings, which are
surrounded by empty lines, are usually nonscenes, and spaCy
inserts special space tokens for such empty lines. However,
the organisers of the shared task seem to have replaced all
whitespace with single spaces beforehand, so there are no
space tokens.</p>
        <p>49/21 (43%) of the trial and training texts have a nonscene
boundary among the first 10 sentences; 2/21 (10%) have a
nonscene boundary among the last 10 sentences.</p>
        <p>MID/START), a quantifier (EVERY), and a
frequency (1M/1S/1W), which we also add as
features.</p>
        <p>We ignore temporal expressions within direct
speech for the same reason as we ignore spatial
and other mentions within direct speech (see next
subsection).
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Mention Features</title>
        <p>Mentions are all noun phrases, including pronouns,
in a document that are part of a coreference cluster.
(These are all noun phrases with a few exceptions,
e.g. expletive and interrogative pronouns.) Hereby,
all mentions of a coreference cluster should denote
the same entity. For the shared task, we ignore
all mentions within direct speech because we only
want to consider entities that are present in a scene,
and direct speech frequently contains mentions of
absent entities.</p>
        <p>We differentiate between spatial entities
(toponyms, nouns with inherent spatiality, e.g.
buildings, inner rooms, landscapes, etc.) and other
entities (characters, objects, concepts, etc.). For
the distinction, we extracted a list of 18,345 spatial
nouns from GermaNet based on their affiliation to
a certain upper-level synset and define a mention
to denote a spatial entity if its head noun is in the
list. We consider times, spaces, and characters
to be most relevant for the shared task but do
not further sub-categorise other (i.e. non-spatial)
entities into characters and non-character entities.
We think that this is not necessary since we
assume clusters of characters to stand out through
a high rate of proper-noun mentions and include
part-of-speech-based features, as we will describe
below.</p>
        <p>We extract an identical set of features for
mentions of spatial entities and mentions of
other entities. In the following, we describe the
procedure for either.</p>
        <p>First, we determine the sentence distance of
each mention, i.e. how many sentences ago the
corresponding entity was last mentioned. If an
entity is mentioned for the first time, we set the
distance to -1. Then, we take the mention with the
lowest distance and the mention with the highest
distance for feature selection and discard all other
mentions.5</p>
        <p>5A sentence contains an unfixed number of mentions but
a feature vector has a fixed number of dimensions, leaving
two options: 1) One can extract features from every mention
but makes it impossible to reallocate the feature values to</p>
        <p>From the two mentions, we extract: the
sentence distance; whether the mention’s head
is a pronoun (no/yes), proper noun (no/yes),
or common noun (no/yes); the root dependency
relation / grammatical role (iobj/nmod/nsubj/obj/
obl); case (acc/dat/gen/nom), person (1per/2per/
3per), number (plu/sing), and gender (fem/masc/
neut); whether the mention contains a determiner
(no/yes) or numeral (no/yes); and the lemmas of
determiners (der/dies/...). Furthermore, we add
a feature that indicates whether the two mentions
are the same (if there is only one mention in the
sentence).</p>
        <p>We are aware that our coreference algorithm is
not perfect and sometimes returns more than one
cluster for the mentions of a single entity or, even
worse, merges the clusters of two or more entities.
We therefore add some features that should give a
hint on the trustworthyness of a coreference cluster.
From the clusters of the mentions, we extract:
number of mentions; percentage of pronoun
mentions, percentage of proper-noun mentions,
and percentage of common-noun mentions; and
its lemma-type ratio, which we define as the
fraction
r(C) =
max#2`|{m ∈ Cnouns : L(H (m)) = `}|
max#1`|{m ∈ Cnouns : L(H (m)) = `}|
with Cnouns = {m ∈ C : is noun(H (m))}, H (m)
being the head of m, L(w) being the lemma of w,
and max#nS being the n-th highest element in S.
Thus, r(C) divides the frequency of the
secondmost frequent lemma in Cnouns by the frequency of
the most frequent lemma in Cnouns.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Statistics</title>
        <p>To counteract overfitting, we exclude features that
occur less than 5 times in the training data, reducing
the number of features from 2,820 to 1,744.
Categorical features are then binarised so that each
feature–value combination becomes a Boolean
feature. This results in 2,104 features altogether.
Table 2 shows the top-45 features scored and
ranked by ANOVA F-value for classifying the
individual mentions, or 2) one can create feature groups for
individual mentions but only extracts features for a fixed
number of mentions. We chose to go with the second option
because we think that the combination of features is very
important, e.g. knowing that a (single) mention has case
= nominative and number = singular is presumably much
more important than knowing that some mention has case =
nominative and some (possibly different) mention has number
= singular.
sentences of the training set into four classes:
None (no boundary), Nonscene-to-Scene,
Sceneto-Nonscene, and Scene-to-Scene.</p>
        <p>The attributes max and min in e.g. max space
mention indicate whether it is the mention with the
highest or lowest sentence distance, respectively.
We can see that some minimal pairs of features
receive (almost) equal F-values, e.g. max space
mention case = nom and min space mention case =
nom. This is due to the fact that most sentences only
mention up to one spatial entity. Among the top-45
features, there are 16 features addressing mentions
of spaces and 8 features addressing temporal
expressions. Further 8 features cover the first token
in a sentence, mainly its part of speech, 5 features
cover the sentence-final punctuation in a sentence,
and 4 features cover the punctuation anywhere
in a sentence. The remaining features address
other mentions (2×) and grammatical features of
subordinate clause (1×) and matrix clause (1×). 34
features cover the current sentence; the remaining
features cover the second last (3×), last (3×), and
next (5×) sentence.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Classifier</title>
      <p>A sentence can either be the start of a scene, the
start of a nonscene, or none of both. Thus, a 3-class
classification (Scene, Nonscene, None) would be
sufficient for training a classifier and constructing
spans for scenes and nonscenes. A span tagged
with X followed by a span tagged with Y then
produces the boundary class X-to-Y for the shared
task’s evaluation. However, to sensitise our model
with respect to types of scene boundaries, we used 4
classes during training (None, Nonscene-to-Scene,
Scene-to-Nonscene, Scene-to-Scene), where the
class Scene is divided into Nonscene-to-Scene and
Scene-to-Scene. In this classification, the first
sentence of a document was treated as
Scene-toScene or Scene-to-Nonscene boundary, depending
on whether the document starts with a scene or a
nonscene. A span tagged with A-to-X followed
by a span tagged with B-to-Y then produces the
boundary class X-to-Y in the evaluation, ignoring
A and B.</p>
      <p>We trained a random forest classifier with
100 decision trees, entropy as split criterion, a
maximum tree depth of 11, and at least 3 samples
per leaf. The class weights were balanced for
all 4 classes. These parameters showed the best
results in a cross-validation (see 6.1). We also
Scene-to-Scene
micro-avg
macro-avg
mean
tested other classification methods, including Naive
Bayes, SVM and k-NN, but decision trees yielded
the best results, presumably because they are able
to learn dependencies between features.
6
6.1</p>
    </sec>
    <sec id="sec-6">
      <title>Results</title>
      <sec id="sec-6-1">
        <title>Cross-Validation</title>
        <p>Due to the sparseness of nonscenes within the
corpus, we also include the trial data for evaluating
our hyperparameters. We perform a 21-fold
crossvalidation where each text represents one fold
(leave-one-out evaluation). For the micro-averaged
F1 including all classes, we achieve a mean over all
folds of 0.070 with a standard deviation of 0.062
(see Table 3). The best result for an individual
text/fold in the micro-averaged evaluation of all
classes is 0.202. Our classifier does not detect
any nonscene correctly, hence the macro-averaged
F1 for the cross-validation is considerably lower.
The F1 for the most-frequent class Scene-to-Scene
is slightly higher than the micro-averaged F1,
which reflects that our classifier is optimised on
the detection of this boundary type and the
microaveraged F1 essentially depends on this class.</p>
        <p>Overall, the number of predictions by the
classifier is similar to the number of gold
boundaries, which creates a balance of precision
and recall (see Table 4).
6.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Final Evaluation</title>
        <p>In the final evaluation, the classifier achieves
slightly better results than in the cross-validation.
Table 5 shows the micro-averaged F1 on both
evaluation sets. Evaluation on the dime novels
yields an average F1 of 0.07 whereas evaluation on
high literature yields an average F1 of 0.12. This
is surprising insofar that the training set consists
only of dime novels which means that the
offdomain performance is higher. Note, however, that
the evaluation sets consist only of 4 and 2 texts,
respectively, and the difference in performance lies
in the range of variation that we observed in the
cross-validation.</p>
        <p>As in the cross-validation, the classifier does not
make correct predictions for Scene-to-Nonscene
or Nonscene-to-Scene boundaries. Thus, the
performance solely relies on the majority class
Scene-to-Scene.</p>
        <p>
          Table 6 additionally shows the γ metric
          <xref ref-type="bibr" rid="ref9">(Mathet
et al., 2015)</xref>
          , which computes the agreement of
gold spans and predicted spans. The value range
is [−∞; 1], where a value of 1 indicates exact
agreement, a value of 0 indicates agreement by
chance, and values &lt; 0 indicate worse-than-chance
agreement. We achieve an agreement slightly better
than chance on both evaluation sets.
7
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Discussion</title>
      <p>With respect to the narratological foundation
of the task, we present a feature design that
reflects the defining criteria for scenes like the
temponym extraction for the temporal dimension,
spatial nouns to describe the current setting,
Levin verb categories for specific actions and the
mention features as representation of the character
constellation. We supplement these features with a
variety of linguistic features such as tense, mood,
voice, modal verbs, and direct speech.</p>
      <p>The use of a random forest algorithm intends
a transparency of the classification but besides
the limited success rate, a manual inspection of
individual trees gave no further insights about
the connection between the learned decisions
and the theoretical foundation of the feature
design (e.g. the assumption that changes of
characters, places or time trigger a new scene). We
optimised the classifier in a way that the amount
of predictions approximately matches the number
of each boundary class within the gold data, which
yields a balanced relation of precision and recall.
Despite of this ratio, the algorithm never detects
nonscenes properly and only a few predictions
for the boundary type Scene-to-Scene match the
correct sentence position.</p>
      <p>Apparently, the only decisive rule which the
classifier has learnt about nonscenes is that
they occur in the beginning of texts. The
class Nonscene-to-Scene was predicted as first
boundary for every text in the cross-validation and
every but one text in the final evaluation. The
class Scene-to-Nonscene followed by
Nonsceneto-Scene was predicted only 3 times in the middle
of a text in the cross-validation and never in the
final evaluation. Neither Scene-to-Nonscene nor
Nonscene-to-Scene boundaries were predicted as
last boundary for any text. This behaviour is caused
by the sentence’s index feature—when we removed
it, the text-initial nonscenes were not predicted
anymore. We kept the feature because a lot of
texts do indeed start with a nonscene6, although
the prediction for the first sentence is not relevant
for the shared task (the classification of the first
sentence only produces a correct boundary when
position and type of the next scene are detected
properly).</p>
      <p>Considering the complexity of the task, the
context window of 2 sentences for each feature
seems rather small, nevertheless, we could not
achieve a better performance with a broader
window. As we have seen in Section 4.4, features
of context sentences are underrepresented in the
feature top-list, which is probably caused by their
sparsity. We hypothesise that a larger context for
individual features might improve the classification
as well as the additional usage of contextual
embeddings. In general, training a classifier with
a large number of (sparsely attested) features on
a small data set is prone to over-/underfitting. We
expect that a larger training set would reduce
67/21 (33%) of the trial and training texts start with a
nonscene.
the sparsity of features and possibly increase the
performance of our approach.</p>
      <p>So far, our mention features include places and
characters (and other entities) but we do not verify
if a character actually appears or if a place is part
of a current action (if not, for example, places
are not relevant as scene features as New York in
the sentence: In a month, she plans to go to New
York with her brother). This possibly causes the
prediction of incorrect boundaries. The approach
of capturing actions by Levin’s verb categories
might be improved by a direct syntactical relation
to a considered entity, still, it is limited to a
single sentence (or its clauses), and a more
contextsensitive approach for capturing actions would be
desirable. The lack of distinction between present
and absent entities influences both the correct
prediction of Scene-to-Scene boundaries and the
rare occurrences of nonscenes.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>This paper describes the participation in the
KONVENS 2021 Shared Task on Scene Detection.
The categories of scene and nonscene are based
on a narratological concept which is streamlined
to four concrete criteria: (change of) time, place,
character constellation, and/or type of action. The
training set consists of 20 dime novels, in which the
classes (boundary types of scenes/nonscenes) are
considerably unbalanced. We developed general,
linguistic features as well as specific ones with
respect to the defining criteria of scenes. Despite
the detailed feature design, our random forest
classifier is only able to detect some occurrences
of one boundary type (Scene-to-Scene) and never
finds nonscenes. Inspecting the best features and
the decision trees revealed that features for time and
space play an important role in the functionality
of the algorithm but within the given context they
cannot develop to their full potential. Given the
current feature design, a stronger focus on a broader
context as well as modelling actions in relation
to space and character entities might improve the
overall results. For that, a larger training set would
be desirable.</p>
      <p>The complete code and the presented model
can be found here: https://gitlab.gwdg.de/
florian.barth/stss.7</p>
      <p>7All results can be reproduced with the annotation data of
the task. Since the data is under copyright, we cannot publish
it within this repository.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Bo</surname>
          </string-name>
          ¨gel, Michael Gertz, Evelyn Gius, Janina Jacke, Jan Christoph Meister, Marco Petris, and Jannik Stro¨tgen.
          <year>2015</year>
          .
          <article-title>Collaborative text annotation meets machine learning: heurecle´a, a digital heuristic of narrative</article-title>
          .
          <source>DHCommons Journal</source>
          ,
          <volume>1</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Tillmann</given-names>
            <surname>Do</surname>
          </string-name>
          ¨nicke.
          <year>2020</year>
          .
          <article-title>Clause-level tense, mood, voice and modality tagging for German</article-title>
          .
          <source>In Proceedings of the 19th International Workshop on Treebanks and Linguistic Theories</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          , Du¨sseldorf, Germany. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Joshua</given-names>
            <surname>Eisenberg</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Finlayson</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Automatic identification of narrative diegesis and point of view</article-title>
          .
          <source>In Proceedings of the 2nd Workshop on Computing News Storylines (CNS</source>
          <year>2016</year>
          ), pages
          <fpage>36</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>Ge´rard Genette</source>
          .
          <year>1983</year>
          .
          <article-title>Narrative discourse: An essay in method</article-title>
          . Cornell University Press.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Birgit</given-names>
            <surname>Hamp</surname>
          </string-name>
          and
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Feldweg</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Germaneta lexical-semantic net for german. Automatic Information Extraction and Building of Lexical Semantic Resources for NLP Applications</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Edward</given-names>
            <surname>Kearns</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Annotating and quantifying narrative time disruptions in modernist and hypertext fiction</article-title>
          .
          <source>In Proceedings of the First Joint Workshop on Narrative Understanding, Storylines, and Events</source>
          , pages
          <fpage>72</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Markus</given-names>
            <surname>Krug</surname>
          </string-name>
          , Frank Puppe, Fotis Jannidis, Luisa Macharowsky, Isabella Reger, and
          <string-name>
            <given-names>Lukas</given-names>
            <surname>Weimar</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Rule-based coreference resolution in German historic novels</article-title>
          .
          <source>In Proceedings of the Fourth Workshop on Computational Linguistics for Literature</source>
          , pages
          <fpage>98</fpage>
          -
          <lpage>104</lpage>
          , Denver, Colorado, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Beth</given-names>
            <surname>Levin</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>English verb classes and alternations</article-title>
          .
          <source>A preliminary Investigation</source>
          ,
          <volume>1</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Yann</given-names>
            <surname>Mathet</surname>
          </string-name>
          , Antoine Widlo¨cher, and
          <string-name>
            <surname>Jean-Philippe Me</surname>
          </string-name>
          ´tivier.
          <year>2015</year>
          .
          <article-title>The Unified and Holistic Method Gamma (γ) for Inter-Annotator Agreement Measure and Alignment</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>41</volume>
          (
          <issue>3</issue>
          ):
          <fpage>437</fpage>
          -
          <lpage>479</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>James</given-names>
            <surname>Pustejovsky</surname>
          </string-name>
          , Parisa Kordjamshidi, MarieFrancine Moens, Aaron Levine, Seth Dworman, and
          <string-name>
            <given-names>Zachary</given-names>
            <surname>Yocum</surname>
          </string-name>
          .
          <year>2015</year>
          . SemEval
          <article-title>-2015 task 8: SpaceEval</article-title>
          .
          <source>In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval</source>
          <year>2015</year>
          ), pages
          <fpage>884</fpage>
          -
          <lpage>894</lpage>
          , Denver, Colorado. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Nils</given-names>
            <surname>Reiter</surname>
          </string-name>
          , Marcus Willand, and
          <string-name>
            <given-names>Evelyn</given-names>
            <surname>Gius</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A shared task for the digital humanities chapter 1: Introduction to annotation, narrative levels and shared tasks</article-title>
          .
          <source>Journal of Cultural Analytics</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Jannik</given-names>
            <surname>Stro</surname>
          </string-name>
          <article-title>¨tgen</article-title>
          and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Gertz</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>HeidelTime: High quality rule-based extraction and normalization of temporal expressions</article-title>
          .
          <source>In Proceedings of the 5th International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>321</fpage>
          -
          <lpage>324</lpage>
          , Uppsala, Sweden. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Albin</given-names>
            <surname>Zehe</surname>
          </string-name>
          , Leonard Konle, Lea Katharina Du¨mpelmann, Evelyn Gius, Andreas Hotho, Fotis Jannidis, Lucas Kaufmann, Markus Krug, Frank Puppe,
          <string-name>
            <given-names>Nils</given-names>
            <surname>Reiter</surname>
          </string-name>
          , et al. 2021a.
          <article-title>Detecting scenes in fiction: A new segmentation task</article-title>
          .
          <source>In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume</source>
          , pages
          <fpage>3167</fpage>
          -
          <lpage>3177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Albin</given-names>
            <surname>Zehe</surname>
          </string-name>
          , Leonard Konle, Svenja Guhr, Lea Katharina Du¨mpelmann, Evelyn Gius, Andreas Hotho, Fotis Jannidis, Lucas Kaufmann, Markus Krug, Frank Puppe, Nils Reiter, and
          <string-name>
            <given-names>Annekea</given-names>
            <surname>Schreiber</surname>
          </string-name>
          . 2021b.
          <article-title>Shared task on scene segmentation@konvens2021</article-title>
          .
          <source>In Shared Task on Scene Segmentation.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>