<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Twin BERT Contextualized Sentence Embedding Space Learning and Gradient-Boosted Decision Tree Ensembles for Scene Segmentation in German Literature</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sebastian Gombert</string-name>
          <email>gombert@dipf.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Center for Education DIPF: Leibniz Institute for Research and Information in Education Frankfurt am Main</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>42</fpage>
      <lpage>48</lpage>
      <abstract>
        <p>This paper documents a submission to the shared task on scene segmentation hosted at KONVENS 2021 (Zehe et al., 2021b). The aim of this shared task was to find methods for segmenting narrative texts into different scenes - segments of text where location, time and the constellation of characters stay more or less coherent. This task is formulated as a sentence classification task where sentences bordering the scenes have to be distinguished from in-scene sentences. The approach presented in this paper is based on two steps. In the first one, a twin BERT training setup is used to learn a sentence embedding space in which sentences functioning as scene borders are well-separated from ones that are in-scene. In the second one, the sentence embeddings generated by this model are used as feature vectors to feed a gradient-boosted decision tree ensemble which conducts final predictions. In the shared task leaderboard, the system ranked second in track 1 and first in track 2.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Scene segmentation in narrative texts is a novel task
in natural language processing introduced by Zehe
et al. (2021a). The aim of this task is to segment
pieces of literature into scenes – sections of text
where the relation of story time and discourse time,
the location and character constellations stay more
or less the same. From a formal point of view,
this problem can be interpreted as a sentence in
context classification task where sentences
separating scenes have to be distinguished from in-scene
ones. This is needed as the typical length of longer
narrative texts such as novels prevents techniques
such as co-reference resolution useful for
proceeding steps of analysis from functioning well
        <xref ref-type="bibr" rid="ref15 ref16">(Zehe
et al., 2021a)</xref>
        . With a text being segmented into
coherent scenes, each scene can be processed
separately improving the performance for such follow
up processing.
      </p>
      <p>
        This paper presents an a participating system
at the KONVENS 2021 shared task on scene
segmentation
        <xref ref-type="bibr" rid="ref15 ref16">(Zehe et al., 2021b)</xref>
        and relies on two
steps. For the first one, a BERT-based
        <xref ref-type="bibr" rid="ref2">(Devlin
et al., 2019)</xref>
        neural network trained in a twin
network setup is used to predict embeddings for
respective input sentences
        <xref ref-type="bibr" rid="ref12 ref6">(Reimers and Gurevych,
2019)</xref>
        . This network was trained to provide an
embedding space in which sentences bordering scenes
are well-separated from in-scene ones. For the
second step, gradient-boosted decision tree ensembles
        <xref ref-type="bibr" rid="ref7">(Mason et al., 1999)</xref>
        are then fed these sentence
embeddings as feature vectors to carry out final
predictions.
      </p>
      <p>For shared task evaluations, this system was
trained on a data set consisting of various
German dime novels where scene borders had been
previously annotated. Participating systems were
evaluated in two tracks using F1 scores. In the
first track, the models were evaluated using a test
set consisting of additional dime novels. In this
track, the system presented in this paper achieved
the second place with an F1 of 0.16. In the second
track, domain-adaptability was probed by
evaluating the systems on a set of German contemporary
highbrow literature. Here, the system presented
performed better and was ranked first with an F1
of 0.26.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <sec id="sec-2-1">
        <title>Task Description</title>
        <p>
          In Zehe et al. (2021a), the authors interpreted the
task of scene segmentation as a sentence
classification task. They defined four different classes
of sentences: no border, scene-to-scene,
scene-tononscene and nonscene-to-scene. The three latter
of these are used to mark the different kinds of
textual borders among the sentences. They trained a
BERT-based
          <xref ref-type="bibr" rid="ref2">(Devlin et al., 2019)</xref>
          classifier
utilising a sliding windows over multiple sentences for
context encoding to carry out sentence
classification.
        </p>
        <p>
          This approach was evaluated against the
unsupervised TextTiling
          <xref ref-type="bibr" rid="ref3">(Hearst, 1997)</xref>
          and TopicTiling
          <xref ref-type="bibr" rid="ref13">(Riedl and Biemann, 2012)</xref>
          methods on a corpus
consisting of 15 German dime novels using cross
validation. While the supervised BERT model
achieved superior results (γ 0.15) compared to the
unsupervised methods (γ 0.01; γ 0.02), the
overall results turned out subpar which led the authors
conclude that scene segmentation can be regarded
as an inherently hard task.
        </p>
        <p>For the KONVENS 2021 shared task, the
organizers provided an expanded version of the data
set presented by Zehe et al. (2021a). This data set
is composed of various German dime novels. The
authors chose this genre as they deemed it easier
for potential models to deal with.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Related Work</title>
        <p>While segmenting text into smaller units such as
tokens, sentences or spans is one of the oldest and
most researched topics in natural language
processing, the task of semantically segmenting narrative
texts into scenes is a new one. In this form, scene
segmentation was first introduced by Zehe et al.
(2021a). From a problem-centric point of view,
Zehe et al. (2021a) relate scene segmentation to the
task of topic segmentation, the task of segmenting
a text by topic changes, as changes of time, place
and character constellation can be interpreted as a
special cases of topic changes.</p>
        <p>
          Most of the more recent work in this area
          <xref ref-type="bibr" rid="ref13 ref8">(Riedl
and Biemann, 2012; Misra et al., 2011)</xref>
          is built
upon latent Dirichlet allocation
          <xref ref-type="bibr" rid="ref1">(Blei et al., 2003)</xref>
          .
This method discovers fields of words consistently
co-occuring in the same contexts. By monitoring
changes in their distribution throughout a text, one
can define topic-wise section borders. Another
related topic according to Zehe et al. (2021a) is
discourse coherence. Recent approaches in this area
rely on neural networks to detect textual coherence
in various setups and use cases
          <xref ref-type="bibr" rid="ref10 ref5">(Li and Jurafsky,
2017; Pichotta and Mooney, 2016)</xref>
          . Changes in
these coherence scores can be used for detecting
borders within texts, as well.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>System Description</title>
      <p>My code can be found under 1.
3.1</p>
      <sec id="sec-3-1">
        <title>Adjustments to the Tag Set</title>
        <p>While Zehe et al. (2021a) used a quaternary tag
set which distinguished scene to scene- and
nonscene to scene borders which is also used for
official shared task evaluations, my system internally
relies on a tertiary tag set consisting of the tags
O, SCENE and NONSCENE. The latter two refer
to the first sentence of an according section. The
reason for this adjustment is that the number of
border sentences is low compared to the number
of non-border sentences. My tertiary tag set is
the smallest classification setup which can be used
to distinguish scenes and non-scenes. Using this
tertiary tagset results in all scene to scene- and
nonscene to scene sentences being grouped under the
SCENE task, and all scene-to-nonscene ones under
the NONSCENE.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Twin BERT Embedding Space Learning</title>
        <p>My system is built around the idea of neural
embedding space learning. Reimers and Gurevych
(2019) introduced the idea of using twin and triplet
network-based training setups for fine-tuning
transformer language models to map sentences into
meaningful semantic vector spaces under the name
Sentence Transformers. In their training setup,
two or three different sentences are fed into the
same transformer language model. These pairs
and triplets of sentences are assigned scores such
as cosine similarity or concrete training labels. A
prediction head which is fed the output of the
transformer language model for all two or three
sentences is trained to predict the assigned scores or
labels. After this training process, the transformer
language model can embed sentences into a vector
space where they are well-separated according to
the respective training objective.</p>
        <p>The idea behind the system presented in this
paper is to combine this approach of twin network
embedding space learning with the sliding
windowbased approach from Zehe et al. (2021a). More
precisely, my approach is to utilise a twin
networkbased training setup to learn an embedding space
encoding information about a sentence as well as
the sentences surrounding it. The goal here is that,</p>
        <sec id="sec-3-2-1">
          <title>1https://github.com/SGombert/</title>
          <p>ssts-2021-sego
within this vector space, the embeddings of
sentences bordering scenes are well-separated from
them of in-scene ones.</p>
          <p>
            Instead of a single BERT model as
            <xref ref-type="bibr" rid="ref12 ref6">(Reimers
and Gurevych, 2019)</xref>
            , it uses two of them with one
functioning as sentence encoder and one as context
encoder. In both cases, the regular pooling layer
output of these networks is used to encode given
input sentences. While the sentence encoder is only
used to predict a sentence embedding for a given
target sentence, the context encoder also predicts
sentence embeddings for a context window of n
sentences to the left and to the right around this
target sentence. The output of both encoders is
concatenated to acquire the final embeddings for
embedding a sentence and its context into vector
space.
          </p>
          <p>m(st) = esent(st) ⊕ econt(st)</p>
          <p>esent(st) = B1(st)
econt(st) = cleft(st) ⊕ B2(st) ⊕ cright(st) (3)
cleft(st) = B2(st−n) ⊕ · · · ⊕ B2(st−1)
cright(st) = B2(st+1) ⊕ · · · ⊕ B2(st+n)
In these equations, st is a given sentence at time
step (position in text) t. m(s) refers to the
function used for predicting embeddings. esent(s) and
econt(s) are the two different encoder networks. B1
and B2 refer to the two underlying BERT networks,
and cleft(st) and cright(st) are the functions used
for acquiring the context of a given sentence st. n
determines the size of this context.</p>
          <p>For training such a sentence embedding model, I
randomly sampled 15000 pairs of sentences which
(1)
(2)
(4)
(5)
were both either scene- or non-scene borders and
15000 pairs where both sentences were from
different categories, the majority of them being pairs of
scene border and in-scene sentences, from the
training set. While the prior set of pairs is assigned a
score of 1, the pairs from the latter set are assigned
a score of -1.</p>
          <p>mconcat(p) = m(s1(p)) ⊕ m(s2(p))
f (p) = L(mconcat(p))
(6)
(7)</p>
          <p>In these equations, p refers to a triple of two
sentences from the training set and an according score
(-1 or 1, depending on class equality), s1(p) and
s2(p) are functions retrieving the first respectively
second sentence from a given training input triple.
f (p) refers to the final output score calculated by
the network during training and L to a linear
feedforward layer. During training both sentences of a
triple and their according local context sentences
are propagated through both the sentence
respectively the context encoders. Their pooling layer
outputs for both sentences are concatenated and
propagated into a linear layer whose single output
neuron is trained to predict the according score
using hinge embedding loss:
f (x, y) =
(x if y = 1
max(0, δ − x) if y = -1
(8)</p>
          <p>Within this function, x is a predicted score, y
a gold standard one and δ the so-called margin, a
hyper parameter which can be used to control the
distances between the vectors a given model learns.
This function is used to learn a maximum
marginlike embedding space which separates scene
borders from in-scene sentences.</p>
          <p>
            The GermanBERT variant provided by
Huggingface Transformers
            <xref ref-type="bibr" rid="ref14">(Wolf et al., 2020)</xref>
            under the
id bert-base-german-dbmdz-uncased2 is used as
a base for both sentence encoder and context
encoder. The reason for choosing this model was
that the data it was pre-trained on includes
narrative texts which makes it an appropriate basis
for a model dealing with literary data. The model
was trained using AdamW
            <xref ref-type="bibr" rid="ref12 ref4 ref6">(Kingma and Ba, 2015;
Loshchilov and Hutter, 2019)</xref>
            with the learning rate
          </p>
        </sec>
        <sec id="sec-3-2-2">
          <title>2https://huggingface.co/</title>
          <p>bert-base-german-dbmdz-uncased
set to 0.000001 and weight decay to 0.0001. The
embedding model was trained for one epoch using
a constant warm up schedule with a constantly
increasing learning rate for the first 1000 iterations.
No batch processing was used during training.</p>
          <p>As visible in figure 3, the model indeed learned
to embed sentences into a vector space in which
they were well-separated into two distinct clusters.
However, it does not seem that the model
generalized the idea of what exactly is a scene border
well from the training data. While for ’Der kleine
Chinesengott’, the German dime novel provided
as trial corpus, the majority of scene borders is
located in the smaller of the two clusters, there are
also borders located in the larger cluster, and,
moreover, many in-scene sentences are also sorted into
the smaller cluster. This phenomenon was visible
after multiple training runs with different sampled
pairs of sentences which implies that drawing clear
distinctions between scene borders and in-scene
sentences is hard for solely BERT-based models.
3.3</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Gradient Boosted Decision Tree</title>
      </sec>
      <sec id="sec-3-4">
        <title>Ensembles</title>
        <p>
          As the embedding model did seemingly not learn
a precise enough distinction between scene
borders and in-scene sentences, using maximum
margin classification with the resulting embeddings
as feature vectors was no option. Instead, I chose
gradient boosted decision tree ensembles
          <xref ref-type="bibr" rid="ref7">(Mason
et al., 1999)</xref>
          as classification algorithm because of
its ability to select distinctive features and ignore
less distinctive ones.
        </p>
        <p>During training, this algorithm creates an
ensemble of weak regression trees trained to predict the
logits within a specialized logistic regression setup.
Combining enough of such trees results in a strong
learner. This is conducted by means of gradient
descent and decision tree learning. Each subsequent
tree is trained to correct erroneous predictions of
the previous ones. As each of them is limited to use
only a small subset of the input features provided
in given input feature vectors, the trained ensemble
can automatically isolate features which globally
distinguish scene borders from in-scene sentences
the best within the training set.</p>
        <p>
          For implementing this part of the system, I used
Catboost
          <xref ref-type="bibr" rid="ref11">(Prokhorenkova et al., 2018)</xref>
          as
framework. The model is based upon its multi class
classification mode. The tree growth policy is set
to lossguide and class weights are used. The
following formula is used for calculating them:
wc = 1 −
        </p>
        <p>num(c)
PcC0 num(c0)
(9)
wc is a respective class weight, c a class, C the
set of all classes, c and c0 classes and num(c) a
function which returns the number of training
examples for a given class. Additionally, I used early
stopping to prevent overfitting. For this, I set the
number of training iterations to 5000, let the
framework choose a learning rate automatically, and then
used the checkpoint of the model which performed
best on the trial dime novel.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <sec id="sec-4-1">
        <title>Results</title>
        <p>Shared task evaluations were carried out on two
different corpora resulting in two different
evaluation tracks. The first of these corpora consisted
of 5 more dime novels similar to the ones systems
were trained on to address in-domain transfer
capabilities of the participating systems. The corpus
used for the second track consisted of two pieces of
highbrow German literature. The aim of this track
was to evaluate out-of-domain transfer capabilities
of the participating systems. My system ranked
second out of four in the first track reaching a micro
F1 of 0.16 and first out of five in the second track
reaching a micro F1 score of 0.26. These results
confirm the difficulty of this task observed by Zehe
et al. (2021a).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2 Qualitative Error Analysis</title>
        <p>To further analyze the results of my system, I turned
to qualitative error analysis. For this purpose, I
collected the false negative and false positive scene
border sentences detected by my system for the trial
corpus and analyzed a selection of them with regard
to common structural patterns. 128 of the sentences
marked as scene borders within the trial corpus
were false positives. What became quickly visible
was that some false positives contained changes
of time, character constellations and/or location.
As these function as important signals for a scene
change, the model seems to have overgeneralized
such cases. The following utterances are examples
for a signified change in time from false positives:
Langsam verstrich die Zeit.</p>
        <p>Natu¨rlich kamen wir zu spa¨t.</p>
        <p>unendlich langsam verstrich die Zeit [...].</p>
        <p>Ich wartete also noch eine Weile, dann aber [...]</p>
        <p>Gerade in dem Moment vernahm ich [...]</p>
        <p>Examples for a change in character constellation
are the following:</p>
        <p>Bills Alarmruf hatte den Spitzbuben verscheucht.</p>
        <p>Der Verfolger war [...] untergetaucht.</p>
        <p>Da ho¨rte ich Tom plo¨tzlich aufstehen [...].</p>
        <p>Tom erhob sich jetzt und entschuldigte sich [...].</p>
        <p>Dem herbeieilenden Portier berichtete ich [...].</p>
        <p>Ich war wieder allein [...].</p>
        <p>Bill meldete in diesem Moment den Besuch Dr. Tu¨rks.</p>
        <p>Ich fand ihn ohnma¨chtig auf dem Fußboden liegen.</p>
        <p>The following utterances are examples for a
location change:</p>
        <p>Wir verließen unser H a¨uschen [...].</p>
        <p>”Schnell, zu Wertheim,” raunte Tom mir zu.</p>
        <p>Wir trafen uns erst wieder draußen in der Linienstraße.
Wir durchsuchten noch einmal das Arbeitszimmer [...].
Endlich erreichten wir den kleinen Antiquita¨tenladen.</p>
        <p>Ich fuhr zur Linienstraße.</p>
        <p>Dann aber schlich ich mich in den dunklen Hausflur.</p>
        <p>Most false positive sentences mention time,
characters or location without explicitly signifying a
change. This speaks for the assumption that the
model might have overgeneralized these signals:
In der Na¨he des schlesischen Bahnhofs.</p>
        <p>”Tom, was tust Du, mußte das sein!”</p>
        <p>Bill lag wieder still.</p>
        <p>Auch Tom lauschte und schien unschlu¨ssig zu sein.</p>
        <p>Isaak Kornblum besaß Telephon.</p>
        <p>Ich tat es.</p>
        <p>On the other hand, many of the false negatives
contain similar signals. This puts the assumption
that the model might have overgeneralized upon
such signals into question. Of course, one needs
to consider that the majority of dimensions of the
respective embeddings encode sentences from the
context of a particular target sentence. Given this
fact in combination that with the observation that
false positives and false negatives share similar
patterns, it seems very likely that these local context
sentences have played a major role for
classification. The following utterances are examples for
false negatives:</p>
        <p>Tom eilte jetzt die Treppe empor [...].</p>
        <p>Mein Weg ging u¨ber die Gartenmauer.</p>
        <p>Dann verschwand er lautlos durch die Vordiele.</p>
        <p>Wir [...] verließen schnell den Laden.</p>
        <p>dann stieg er die Leiter empor.</p>
        <p>Tom verschwand schnell durch die Verbindungstu¨r [...].
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion &amp; Outlook</title>
      <p>I presented my submission to the shared task on
scene segmentation at KONVENS 2021, a system
aimed at segmenting German narrativew texts into
distinct scenes, spans of text where character
constellations, discourse- and story time, and locations
stay more or less the same. For its implementation,
the task was interpreted as a sentence in context
classification task. For solving this task, I first
trained a neural model consisting of two
GermanBERT networks, the sentence encoder and context
encoder, which, in conjunction, predict
contextualized sentence embeddings. This was conducted in a
twin network setup where triplets of two sentences
and an according score were fed to a a linear layer
responsible for predicting such an according score.</p>
      <p>The goal behind this was to train a model which
would be able to embed sentences into a vector
space in which sentences functioning as scene
borders would be well-separated from in-scene ones
which could then be used as feature vectors in
regular classification. While the model indeed learned
a vector space in which sentences were more or
less sorted into two distinct clusters, these clusters
did not seem to capture a general understanding of
the concept of scene borders. This is shown by the
observation that gold standard scene borders from
the trial set were sorted into both clusters when
embedded by the model.</p>
      <p>For this reason, gradient boosting was chosen
as a subsequent classification algorithm for its
ability to isolate a subset of features which would still
be able to separate classes well. Early stopping
was used during training, meaning that the model
was trained for 5000 iterations on the shared task
training data and the iteration of the model which
achieved best results on the trial data set was
chosen as final. This achieved comparably poor results
with micro F1 scores of 0.16 for track 1
respectively 0.26 for track 2. Nonetheless, these results
were sufficient for ranks 2/4 respectively 1/5 in the
two tracks.</p>
      <p>It is an interesting observation that my system
performs better for highbrow literature in spite of
the fact that its training data consisted solely of
dime novels as it contradicts the assumption of the
authors that dime novels would be potentially
easier to deal with for participating systems compared
to highbrow literature. A possible explanation for
this could lie in the more formal nature of
highbrow literature which might result in more
regularities that are useful for successful classification.
However, without further inspection, this remains
speculation.</p>
      <p>Further work could be the optimization of the
architecture and training procedure of the
contextualized sentence embedding model presented in this
paper. This might lead to improved downstream
training results. Moreover, as gradient boosting
functions as feature-based learning algorithm, it
could be an option to combine contextualized
sentence embeddings with statistical and hand-crafted
features for representing sentences in context. In
general, it can be said that the problem is far from
solved as sugggested by the poor results.
However, the idea of learning contextualized sentence
embeddings and the optimization of the according
training procedure could be a useful option to for
future work on the topic.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>David M.</given-names>
            <surname>Blei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Andrew Y.</given-names>
            <surname>Ng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael I.</given-names>
            <surname>Jordan</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>3</volume>
          (
          <issue>4</issue>
          -5):
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          , Minneapolis, Minnesota. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Marti A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Text tiling: Segmenting text into multi-paragraph subtopic passages</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>23</volume>
          (
          <issue>1</issue>
          ):
          <fpage>33</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Diederik P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Ba</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>In 3rd International Conference on Learning Representations, ICLR</source>
          <year>2015</year>
          , San Diego, CA, USA, May 7-
          <issue>9</issue>
          ,
          <year>2015</year>
          , Conference Track Proceedings.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Jiwei</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Neural net models of open-domain discourse coherence</article-title>
          .
          <source>In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>198</fpage>
          -
          <lpage>209</lpage>
          , Copenhagen, Denmark. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Loshchilov</surname>
          </string-name>
          and
          <string-name>
            <given-names>Frank</given-names>
            <surname>Hutter</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Decoupled weight decay regularization</article-title>
          .
          <source>In 7th International Conference on Learning Representations, ICLR</source>
          <year>2019</year>
          ,
          <article-title>New Orleans</article-title>
          , LA, USA, May 6-
          <issue>9</issue>
          ,
          <year>2019</year>
          . OpenReview.net.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Llew</given-names>
            <surname>Mason</surname>
          </string-name>
          , Jonathan Baxter,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Bartlett</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Marcus</given-names>
            <surname>Frean</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Boosting algorithms as gradient descent</article-title>
          .
          <source>In Proceedings of the 12th International Conference on Neural Information Processing Systems</source>
          ,
          <source>NIPS'99, page 512-518</source>
          , Cambridge, MA, USA. MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Hemant</given-names>
            <surname>Misra</surname>
          </string-name>
          , Franc¸ois Yvon, Olivier Cappe´, and Joemon Jose.
          <year>2011</year>
          .
          <article-title>Text segmentation: A topic modeling perspective</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>47</volume>
          (
          <issue>4</issue>
          ):
          <fpage>528</fpage>
          -
          <lpage>544</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Karl</given-names>
            <surname>Pearson</surname>
          </string-name>
          .
          <year>1901</year>
          . LIII.
          <article-title>on lines and planes of closest fit to systems of points in space</article-title>
          .
          <source>The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science</source>
          ,
          <volume>2</volume>
          (
          <issue>11</issue>
          ):
          <fpage>559</fpage>
          -
          <lpage>572</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Karl</given-names>
            <surname>Pichotta</surname>
          </string-name>
          and
          <string-name>
            <given-names>Raymond J.</given-names>
            <surname>Mooney</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Learning statistical scripts with lstm recurrent neural networks</article-title>
          .
          <source>In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence</source>
          ,
          <source>AAAI'16</source>
          , page 2800-
          <fpage>2806</fpage>
          . AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Liudmila</given-names>
            <surname>Ostroumova</surname>
          </string-name>
          <string-name>
            <surname>Prokhorenkova</surname>
          </string-name>
          , Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and
          <string-name>
            <given-names>Andrey</given-names>
            <surname>Gulin</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Catboost: unbiased boosting with categorical features</article-title>
          .
          <source>In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems</source>
          <year>2018</year>
          ,
          <article-title>NeurIPS 2018</article-title>
          , December 3-
          <issue>8</issue>
          ,
          <year>2018</year>
          , Montre´al, Canada, pages
          <fpage>6639</fpage>
          -
          <lpage>6649</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Nils</given-names>
            <surname>Reimers</surname>
          </string-name>
          and
          <string-name>
            <given-names>Iryna</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>SentenceBERT: Sentence embeddings using Siamese BERTnetworks</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>3982</fpage>
          -
          <lpage>3992</lpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Riedl</surname>
          </string-name>
          and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Biemann</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>TopicTiling: A text segmentation algorithm based on LDA</article-title>
          .
          <source>In Proceedings of ACL 2012 Student Research Workshop</source>
          , pages
          <fpage>37</fpage>
          -
          <lpage>42</lpage>
          ,
          <string-name>
            <surname>Jeju</surname>
            <given-names>Island</given-names>
          </string-name>
          , Korea. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Wolf</surname>
          </string-name>
          , Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Rush</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Transformers: State-of-the-art natural language processing</article-title>
          .
          <source>In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          , pages
          <fpage>38</fpage>
          -
          <lpage>45</lpage>
          , Online. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Albin</given-names>
            <surname>Zehe</surname>
          </string-name>
          , Leonard Konle, Lea Katharina Du¨mpelmann, Evelyn Gius, Andreas Hotho, Fotis Jannidis, Lucas Kaufmann, Markus Krug, Frank Puppe, Nils Reiter, Annekea Schreiber, and
          <string-name>
            <given-names>Nathalie</given-names>
            <surname>Wiedmer</surname>
          </string-name>
          . 2021a.
          <article-title>Detecting scenes in fiction: A new segmentation task</article-title>
          .
          <source>InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume</source>
          , pages
          <fpage>3167</fpage>
          -
          <lpage>3177</lpage>
          , Online. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Albin</given-names>
            <surname>Zehe</surname>
          </string-name>
          , Leonard Konle, Svenja Guhr, Lea Katharina Du¨mpelmann, Evelyn Gius, Andreas Hotho, Fotis Jannidis, Lucas Kaufmann, Markus Krug, Frank Puppe, Nils Reiter, and
          <string-name>
            <given-names>Annekea</given-names>
            <surname>Schreiber</surname>
          </string-name>
          . 2021b.
          <article-title>Shared task on scene segmentation@konvens2021</article-title>
          .
          <source>In Shared Task on Scene Segmentation.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>