<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluation and Alignment of Movie Events Extracted via Machine Learning from a Narratological Perspective</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Feng Zhou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>FedericoPianzola</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Language and Cognition, University of Groningen</institution>
          ,
          <addr-line>Oude Kijk in 't Jatstraat 26, 9712 EK Groningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <fpage>49</fpage>
      <lpage>62</lpage>
      <abstract>
        <p>We combine distant viewing and close reading to evaluate the usefulness of events extracted via machine learning from audio description of movies. To do this, we manually annotate events from Wikipedia summaries for three movies and align them to ML-extracted events. Our exploration suggests that computational narratology should combine datasets with events extracted from multimodal data sources that take into account both visual and verbal cues when detecting events.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;movie events</kwd>
        <kwd>narrative events</kwd>
        <kwd>computational narratology</kwd>
        <kwd>audio description</kwd>
        <kwd>movie summaries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Related Research</title>
      <p>
        Narrative theory de昀椀nes events as the minimal units of narrative, being states, processes in
time, and changes of state [
        <xref ref-type="bibr" rid="ref22">22, 32</xref>
        ]. Moreover, narrative events are o昀琀en analyzed along the
dimensions of their duration and frequency in chronological terms, and span of their
representation through a medium1[0]. In the case of movies, events are represented as scenes, typically
taking place at a location, involving a set of characters, and spanning a continuous period of
time [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In addition, sub-scenes occur within a single location and time frame, are usually
shorter, contain several shots that do not constitute a complete semantic unit, and may lack a
clear beginning or end8[]. As scenes are sequences of sub-scenes, the shots contained within
a series of sub-scenes can be used to create narrative sub-events or whole even7ts, 1[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        To cope with such granularity, most ML methods extract events through the cross-modal
understanding of audio, visual, and video-related text: e.g. alignment between script and
timestamped subtitle [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], retrieval via semantic graphs of video clips and plot summari3e4s, [
        <xref ref-type="bibr" rid="ref33">35</xref>
        ],
segmentation by detecting audio descriptions between character dialogue3s1,[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. The
evaluation of ML methods is done via manual annotation of movie events and aligning sentences in
movie scripts or summaries with video clip3s0[
        <xref ref-type="bibr" rid="ref33 ref34">, 35, 36</xref>
        ]. User-generated content platforms like
Wikipedia and IMDb provide a vast amount of movie-related textual da3ta, 1[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] that can
provide high-level semantic annotation of movies, such as describing speci昀椀c scenes and events
using text [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and serve as gold standards to evaluate the ability of machine learning methods
for event extraction.
      </p>
      <p>Despite the variety of methods, the relevance of ML-extracted events for summaries that
align with human expectations has not been explored in detail. Therefore, we compared a
dataset of automatically extracted events to events manually extracted from movie summaries,
focusing on alignment, duration, length, and type of events.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <sec id="sec-3-1">
        <title>3.1. Machine Learning Extracted Events from Audio Description</title>
        <p>
          Audio description (AD), also known as Descriptive Video Service (DVS), is conceived for
visually impaired audiences and usually contains the most important visual elements of a movie
frame, such as scene changes, characters’ appearance, actions, and interactions between
characters [31]. AD uses short and precise language inserted between characters’ dialogue and
mixed with the original soundtrack of the movie31[
          <xref ref-type="bibr" rid="ref23">, 23</xref>
          ]. AD has been used to extract movie
events and research the narrative structure of visual materials because of the high quality of
the movie description text and the high degree of alignment with visual conten3t1][. We used
the Montreal Video Annotation Dataset (M-VAD)3[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]1 to evaluate this type of story
representation with respect to narrative event granularity and plot covera2g8e].[ Three movies of
1Since AD appears between character dialogues, the curators of the M-VAD dataset 昀椀rst separated the AD
soundtrack from the movie soundtrack, then they automatically segmented the movie into di昀erent events by
detecting the pauses between ADs, 昀椀nally they transcribed the AD of the events into text using a professional
transcription service. The data including the movie event description text and the corresponding video clips’
timestamps are obtained from the M-VAD Names GitHub repository:
https://github.com/aimagelab/mvad-namesdataset/releases/tag/1.0. The subset used in this article and the manually-annotated data are available in the
reposdi昀erent genres were selected from the M-VAD dataset:500 Days Of Summer (2009, romantic
comedy, 95 minutes), The Social Network (2010, biographical drama, 120 minutes), andFlight
(2012, thriller drama, 138 minutes).
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Manually Annotated Events from Movie Summaries</title>
        <p>
          Wikipedia crowdsourced summaries can be considered a proxy of the events that people 昀椀nd
important in a movie, spanning the whole story arc28[]. They are also used as gold standard
in NLP because they have an event granularity that balances e昀케cient plot recognition and
information loss 2[
          <xref ref-type="bibr" rid="ref35 ref8">8, 37</xref>
          ]. We linked these gold-standard events to the respective video clips
by identifying the timestamps corresponding to the events’ occurrences in the movie. These
events manually extracted from summaries allowed us to evaluate the ML movie event dataset
with respect to key content coverage (condensing important parts of informatio2n9)][and the
temporal localization of such events in the plot. Additionally, events can be classi昀椀ed by type
[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and used to explore the criteria underlying the summarization process and the plot structure
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. We annotated events in summary texts based on the main verb in sentences, according to
four types: stative, process, change of state in the story world, or none of the previo3u2s].[Two
annotators independently annotated events from Wikipedia summary texts using the so昀琀ware
CATMA [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and then marked the start and end timestamps of the events’ corresponding video
clips while watching the movies.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <p>
        To test whether Wikipedia summaries’ events cover the semantic content of most video clips
describing key events in the movie, we 昀椀rst rescaled between 0 and 1 the timestamp (for video
clip index) and annotation’s 昀椀rst character’s position (for plot sentence index) of all annotated
events in the summary text that we were able to align to corresponding video clips. In this
way, we obtained a comparable plot sentence index and video clip inde1x].[2 Second, we
manually aligned human annotated events with the ML-extracted events based on the
respective timestamps to examine whether the ML-extracted events have the same event description
granularity as Wikipedia summary events from text and video perspectiv2e8s][. The criterion
used for the manual alignment is whether the overlapping time interval between one or
several corresponding ML- and human-extracted events is at least 75% of either extracted event(s)
duration [24]. Additionally, a human evaluation was conducted to check whether the events
extracted via ML can represent the semantics of manually annotated events. Subsequently,
to assess the extent of the alignment between ML-extracted and human-annotated events, we
counted the number of events, and compared their time duration and distribution throughout
the movie [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
itory: https://github.com/Marina-Zhou/Movie_Event_Alignment
2To avoid altering the depiction of event distribution in the movie summary, the few summary events without a
timestamp have not been plotted in2. Additionally, there are instances where a manually extracted event relates
to a number of movie clips. Only the segment with the earliest chronological occurrence is kept, the others are
discarded for the sake of clarity of the visualization. The full list of events is anyway available in the released
dataset.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Discussion</title>
      <p>[and with whom Whip had spent the night before the incident].</p>
      <p>
        Another notable phenomenon occurs at the end of the summaries, where many events of the
movie ending are reported. The distribution histograms show that the majority of the events
in the summaries correspond to events from the last third of the movie. A possible reason
is that the end is where usually the most important events of a story occur, sometimes being
the con昀氀uence of multiple events [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. But it could also be due to the interpretative function
of summaries, which need to provide a sense of closure, therefore the short summary text
describes more video clip events at the end of the 昀椀lm to create such e昀ect of closure2[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], even
though this may not be present in the story.
      </p>
      <p>Figure 3 shows that ML-extracted events cover almost the entire duration of the movie,
whereas events in summaries have notable gaps. The width of the bands re昀氀ects the length
of the reported events and, in Figur4e, it is visible how several ML-extracted events can be
encompassed by a longer summary event. There are also some successfully aligned events
belonging to sub-scenes about secondary characters that appear in parallel narrati2v,e7s].[</p>
      <p>Figure5 shows that only part of the ML-extracted events not aligned with summary events
occur within the time spans corresponding to manually-annotated events. These events
cannot be aligned to a summary event either because the overlapping range with summary events
is too small (less than 75%) or because the number of ML-events is insu昀케cient to fully re昀氀ect
the semantics of the manually-annotated events. For example, in the movFielight, the
simple ML-extracted event [Whip turns o昀 the news] overlaps with the much more informative
manually-annotated event [he stays with Charlie until the NTSB hearing]. This is due to the
mere descriptive function of AD.</p>
      <p>The other type of ML-extracted events not aligned with summary events occur outside the
time spans covered by summary events. To some extent, these events can be considered
progressive scenes that set the background for a summary event and serve as event separator1s6][.
However, most of them are unimportant events that do not contribute to the plot development.
They primarily describe the appearance and secondary actions of the characters, and the visual
elements in scenes.</p>
      <p>Our annotated data indicates that events extracted from AD more easily align with
manuallyannotated events when they contain descriptions of visual elements, especially direct
descriptions of character actions, rather than information captured through dialogue or
background sounds. These ML-extracted events are located within manually-annotated events
with longer time duration and may not exhibit direct semantic relevance for such
manuallyannotated events. Nevertheless, they constitute an integral component of the broader
manually-annotated events because they are progressive scenes that separate sub-events. As
such, they have a role in the narrative progression because they introduce new situations where
subsequent events happen. For instance, in the movieFlight, the ML-extracted event [They
reach a clearing, through clouds and bright sunlight, as Whip levels out] serves as a
transition, enabling the audience to experience the severe turbulence of the plane and facilitating a
smoother acceptance of new information and emotions during a plot transition.</p>
      <p>
        Most of the manually-annotated events are of type “process” (67~78%) and “change of state”
(14~16%). Since ML-extracted events contain the description of the characters’ actions, it is
easier to align them with manually-annotated “process” and “change of state” events, than with
“stative” events (see Figure6). Notably, most of the “change of state” events in the summary
are aligned with ML-extracted events, suggesting that these are indeed important plot turns
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. However, we were not able to align some “change of state” events from the movi5e0s0
Days of Summer and The Social Network (see Appendix A). This suggests that AD events could
still miss information that is important for the plot.
6. Limitations
• The dataset used in this research only contains data from three di昀erent movies.
Therefore, clear generalization about the usefulness of events extracted from AD via ML for
the understanding of narrative cannot be drawn.
• The source of ML-extracted events are AD scripts, which are written by humans.
Therefore, this type of data cannot be considered completely machine-generated. A
comparison with captions generated by multimodal Large Language Models could provide a
      </p>
      <p>
        better case study.
• AD only describes the visual elements of the movie, not the events performed through
the character’s dialogue, e.g someone angrily yelling at someone else could be a highly
relevant event, maybe a break-up between two partners. Therefore, we suggest to
combine AD with data sourced from subtitles or movie scripts.
• The amount of manually-annotated data is quite small (only three movies), so for all the
analyses we retained all data annotated by one annotator. The inter-annotator agreement
is reported only for reference. Therefore, some subjectivity is unavoidable.
• There is a 1-2 seconds delay between ADs and the corresponding scenes in the movie, so
the curators of the M-VAD dataset added 2 seconds at the end of the video to
compensate for this deviation 3[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, this method may also cause inaccurate timestamps.
For example, [He takes Whip’s hand] in the movieFlight has the end timestamp 01:26:56
but the event actually ends 01:26:54. Moreover, in the M-VAD dataset, the 昀椀rst 3
minutes of the movie are discarded. All these choices slightly a昀ect the alignment with our
manually-annotated events.
• We removed some manually-annotated events not corresponding to any video clip in
the movie when plotting the visualization, e.g. an event reported by a character but not
shown on screen.
      </p>
    </sec>
    <sec id="sec-6">
      <title>7. Conclusions</title>
      <p>One one hand, most of the events extracted by machine learning from AD can be considered
sub-events rather than narrative events. Because their time duration is very short, they cannot
be considered as complete semantic units and need to be understood in conjunction with the
video context and dialogue. These events o昀琀en include descriptions of visual elements in the
椀昀lm and are placed between character dialogues. Some of these events serve as progressive
scenes, both inside and outside the automatically extracted event, and are used as separators
for other sub-events and events. On the other hand, the events extracted manually from the
summary are a high-level synthesis of the dialogue and actions of the characters in the movie,
and cover longer time spans that may contain multiple sequential events and sub-events.
Moreover, verbs in summary events o昀琀en refer to one of the four categories of narratological events
and are highly relevant for the understanding of the plot, but some verbs in AD events describe
the settings and cannot be considered narratological events.</p>
      <p>Computational narratology can certainly make use of ML-extracted events based on audio
description but there are strong limitations that suggest using them in combination with other
data sources that take into account both visual and verbal cues when detecting events. The
multimodal communicative nature of movies seems to pose a big challenge for computational
narratology because existing datasets used for ML can only provide an extremely simpli昀椀ed
notion of narrative event.</p>
      <p>In order to train ML models to extract narrative events, there is a need for movie events
datasets that include audiovisual information, subtitles, and AD sentences. The annotation
guidelines should also take into account that the narratological concept of event, as unit of a
story, should also be related to other conceptualization of events, namely sub-events as those
observed in AD and macro-events as those observed in movie summaries.</p>
      <p>
        Since events and causal relationship between events play a key role in narrative
understanding [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], our multimodal human evaluation method on the M-VAD movie event dataset should
be replicated to evaluate the ability of other automatic methods for the extraction of narrative
events. This could help to understand how summary events are constructed by smaller units
such as events displayed in the story and sub-events.
      </p>
      <p>
        In addition, we have not yet conducted research on di昀erent types of events extracted by ML
and the story arcs that they form. In the future, we could also explore whether movie events
extracted via di昀erent ML methods contain important narrative events (key plot turns) in the
story arc [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
    </sec>
    <sec id="sec-7">
      <title>A. Appendix</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nagrani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Brown</surname>
          </string-name>
          , and
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          . “Condensed Movies:
          <article-title>Story Based Retrieval with Contextual Embeddings”</article-title>
          .
          <source>InC: omputer Vision - ACCV 2020: 15th Asian Conference on Computer Vision</source>
          , Kyoto, Japan,
          <source>November 30 - December 4</source>
          ,
          <year>2020</year>
          ,
          <string-name>
            <given-names>Revised</given-names>
            <surname>Selected</surname>
          </string-name>
          <string-name>
            <given-names>Papers</given-names>
            ,
            <surname>Part</surname>
          </string-name>
          <string-name>
            <given-names>V.</given-names>
            <surname>Kyoto</surname>
          </string-name>
          , Japan: Springer-Verlag,
          <year>2020</year>
          , pp.
          <fpage>460</fpage>
          -
          <lpage>479</lpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .1007/9 78-
          <lpage>3</lpage>
          -
          <fpage>030</fpage>
          -69541-5\_
          <fpage>28</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bordwell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Thompson</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. SmithF.</surname>
          </string-name>
          <article-title>ilm art: An introduction</article-title>
          . Vol.
          <volume>8</volume>
          .
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          New York,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Burghardt</surname>
          </string-name>
          , A. He昀琀berger, J. Pause,
          <string-name>
            <given-names>N.-O.</given-names>
            <surname>Walkowski</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Zeppelzauer</surname>
          </string-name>
          . “
          <article-title>Film and video analysis in the digital humanities-an interdisciplinary dialog”</article-title>
          <source>D.Iing:ital Humanities Quarterly 14.4</source>
          (
          <year>2020</year>
          ). url: http://digitalhumanities.org:8081/dhq/vol/14/4/000532/0 00532.html.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chaturvedi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Roth</surname>
          </string-name>
          . “
          <article-title>Where Have I Heard This Story Before? Identifying Narrative Similarity in Movie Remakes”</article-title>
          .
          <article-title>InP:roceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</article-title>
          , Volume
          <volume>2</volume>
          (
          <string-name>
            <surname>Short</surname>
            <given-names>Papers). New</given-names>
          </string-name>
          <string-name>
            <surname>Orleans</surname>
          </string-name>
          , Louisiana: Association for Computational Linguistics,
          <year>2018</year>
          , pp.
          <fpage>673</fpage>
          -
          <lpage>678</lpage>
          . doih:ttps://doi.org/10.18653/v1/
          <fpage>N18</fpage>
          -210
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Cohn</surname>
          </string-name>
          . “
          <article-title>Visual Narrative Structure”</article-title>
          .
          <source>InC:ognitive Science 37.3</source>
          (
          <issue>2013</issue>
          ), pp.
          <fpage>413</fpage>
          -
          <lpage>452</lpage>
          . doi: https://doi.org/10.1111/cogs.12016.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Cooper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nascimento</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Francis</surname>
          </string-name>
          . “
          <article-title>Exploring Film Language with a Digital Analysis Tool: the Case of Kinolab</article-title>
          .”
          <source>InD:HQ: Digital Humanities Quarterly 15.1</source>
          (
          <year>2021</year>
          ). url: https://www.digitalhumanities.org/dhq/vol/15/1/000515/000515.ht m.l
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>J. E. Cutting.</surname>
          </string-name>
          “
          <article-title>Event segmentation and seven types of narrative discontinuity in popular movies”</article-title>
          .
          <source>In: Acta Psychologica</source>
          <volume>149</volume>
          (
          <year>2014</year>
          ), pp.
          <fpage>69</fpage>
          -
          <lpage>77</lpage>
          . doi: https://doi.org/10.1016/j.actps y.
          <year>2014</year>
          .
          <volume>03</volume>
          .003.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>J. E. Cutting.</surname>
          </string-name>
          “
          <article-title>Narrative theory and the dynamics of popular movies”</article-title>
          .
          <source>IPns:ychonomic Bulletin &amp; Review 23.6</source>
          (
          <issue>2016</issue>
          ), pp.
          <fpage>1713</fpage>
          -
          <lpage>1743</lpage>
          . doi: https://doi.org/10.3758/s13423-016-10
          <fpage>51</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dunn</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Schumacher</surname>
          </string-name>
          . “Explaining Events to Computers: Critical Quanti昀椀cation, Multiplicity and Narratives in Cultural Heritage.”
          <source>DInH:Q: Digital Humanities Quarterly 10.3</source>
          (
          <year>2016</year>
          ). url: http://www.digitalhumanities.org/dhq/vol/10/3/000262/000262.ht m.l
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Genette</surname>
          </string-name>
          .
          <article-title>Narrative discourse: an essay in method</article-title>
          . Ithaca, N.Y: Cornell University Press,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Gius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Meister</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Meister</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Petris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bruck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jacke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schumacher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gerstorfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Flüh</surname>
          </string-name>
          , and J. HorstmannC. atma.
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.6419805.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Gius</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Vauth</surname>
          </string-name>
          . “
          <article-title>Towards an Event Based Plot Model. A Computational Narratology Approach”</article-title>
          .
          <source>In:Journal of Computational Literary Studies</source>
          <volume>1</volume>
          .1 (
          <year>2022</year>
          ). doi: https://doi .org/10.48694/jcls.11 0.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jänicke</surname>
          </string-name>
          , G. Franzini,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Cheema</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. Scheuermann.</surname>
          </string-name>
          “
          <article-title>On Close and Distant Reading in Digital Humanities: A Survey and Future Challenges”</article-title>
          . IEnu:rographics Conference on
          <string-name>
            <surname>Visualization (EuroVis) - STARs</surname>
          </string-name>
          (
          <year>2015</year>
          ). Ed. by
          <string-name>
            <given-names>R.</given-names>
            <surname>Borgo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ganovelli</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Viola</surname>
          </string-name>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>103</lpage>
          . doi:
          <volume>10</volume>
          .2312/eurovisstar.20151113.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Im</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schriber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gross</surname>
          </string-name>
          , and H. P昀椀ster. “
          <article-title>Visualizing Nonlinear Narratives with Story Curves”</article-title>
          .
          <source>InIE:EE Transactions on Visualization and Computer Graphics 24.1</source>
          (
          <issue>2018</issue>
          ), pp.
          <fpage>595</fpage>
          -
          <lpage>604</lpage>
          . doi: https://doi.org/10.1109/TVCG.
          <year>2017</year>
          .
          <volume>274411</volume>
          8.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>I.</given-names>
            <surname>Laptev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marszalek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Rozenfeld</surname>
          </string-name>
          . “
          <article-title>Learning realistic human actions from movies”</article-title>
          .
          <source>In:2008 IEEE Conference on Computer Vision</source>
          and Pattern Recognition. Anchorage,
          <string-name>
            <surname>AK</surname>
          </string-name>
          , USA: Ieee,
          <year>2008</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . doi:https://doi.org/10.1109/CVPR.
          <year>2008</year>
          .
          <volume>458775</volume>
          6.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Kuo</surname>
          </string-name>
          . “
          <article-title>Content-Based Movie Analysis and Indexing Based on AudioVisual Cues”</article-title>
          .
          <source>In:IEEE Transactions on Circuits and Systems for Video Technology 14.8</source>
          (
          <issue>2004</issue>
          ), pp.
          <fpage>1073</fpage>
          -
          <lpage>1085</lpage>
          . doi: https://doi.org/10.1109/TCSVT.
          <year>2004</year>
          .
          <volume>83196</volume>
          8.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shmilovici</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Last</surname>
          </string-name>
          . “
          <article-title>MND: A New Dataset and Benchmark of Movie Scenes Classi昀椀ed by Their Narrative Function”</article-title>
          . In:Computer Vision - ECCV 2022 Workshops. Ed. by
          <string-name>
            <given-names>L.</given-names>
            <surname>Karlinsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Michaeli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Nishino</surname>
          </string-name>
          . Vol.
          <volume>13804</volume>
          . Cham: Springer Nature Switzerland,
          <year>2023</year>
          , pp.
          <fpage>610</fpage>
          -
          <lpage>626</lpage>
          . doi:https://doi.org/10.1007/978-3-
          <fpage>031</fpage>
          -25069-9\_3
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Magliano</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Zacks</surname>
          </string-name>
          . “
          <article-title>The Impact of Continuity Editing in Narrative Film on Event Segmentation: Cognitive Science”</article-title>
          .
          <source>In:Cognitive Science 35.8</source>
          (
          <issue>2011</issue>
          ), pp.
          <fpage>1489</fpage>
          -
          <lpage>1517</lpage>
          . doi: https://doi.org/10.1111/j.1551-
          <fpage>6709</fpage>
          .
          <year>2011</year>
          .
          <volume>01202</volume>
          . x.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P.</given-names>
            <surname>Papalampidi</surname>
          </string-name>
          , F. Keller, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Lapata</surname>
          </string-name>
          . “
          <article-title>Movie Plot Analysis via Turning Point Identi昀椀- cation”</article-title>
          .
          <source>In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          .
          <source>Hong Kong</source>
          , China: Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>1707</fpage>
          -
          <lpage>1717</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1180.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>V. F.</given-names>
            <surname>Perkins</surname>
          </string-name>
          . “
          <article-title>Where is the world? The horizon of events in movie 昀椀ction”</article-title>
          .
          <article-title>InS:tyle and meaning: Studies in the detailed analysis of 昀椀lm (</article-title>
          <year>2005</year>
          ), pp.
          <fpage>16</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Phelan</surname>
          </string-name>
          and P. J. Rabinowitz, eds.Understanding narrative.
          <article-title>The theory and interpretation of narrative series</article-title>
          . Columbus: Ohio State University Press,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>G.</given-names>
            <surname>Prince</surname>
          </string-name>
          .
          <article-title>A Grammar of Stories: An Introduction</article-title>
          . Berlin, Boston: De Gruyter,
          <year>1974</year>
          . doi: https://doi.org/10.1515/978311081590 0.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23] [24] [25]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tandon</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schiele</surname>
          </string-name>
          .
          <article-title>“A Dataset for Movie Description”</article-title>
          .
          <source>In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          .
          <year>2015</year>
          , pp.
          <fpage>3202</fpage>
          -
          <lpage>3212</lpage>
          . doi: https://doi.org/10.1109/CVPR.
          <year>2015</year>
          .
          <volume>729894</volume>
          0.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Torabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tandon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schiele</surname>
          </string-name>
          . “Movie Description”.
          <source>InI:nternational Journal of Computer Vision</source>
          <volume>123</volume>
          .1 (
          <issue>2017</issue>
          ), pp.
          <fpage>94</fpage>
          -
          <lpage>120</lpage>
          . doi: https://doi.org/10.1007/s11263-016-0987- 1.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Salway</surname>
          </string-name>
          . “
          <article-title>A corpus-based analysis of audio description”</article-title>
          .
          <source>In: Leiden, The Netherlands: Brill</source>
          ,
          <year>2007</year>
          , pp.
          <fpage>151</fpage>
          -
          <lpage>174</lpage>
          . doih:ttps://doi.org/10.1163/9789401209564\_01 2.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Sang</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          . “
          <article-title>Character-based movie summarization”</article-title>
          .
          <source>InP:roceedings of the 18th ACM international conference on Multimedia. Firenze Italy: Acm</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>855</fpage>
          -
          <lpage>858</lpage>
          . doi: https://doi.org/10.1145/1873951.187409 6.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>E.</given-names>
            <surname>Segal</surname>
          </string-name>
          .
          <source>Beginnings and Endings</source>
          .
          <year>2019</year>
          . doi: https://doi.org/10.1093/acrefore/978019020 1098.013.1051.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ji</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Synopses of Movie Narratives: a Video-Language Dataset for Story Understanding</article-title>
          .
          <year>2022</year>
          . url: http://arxiv.org/abs/2203.0571 1.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Syed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yousef</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Al Khatib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jänicke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          . “
          <article-title>Summary Explorer: Visualizing the State of the Art in Text Summarization”</article-title>
          .
          <source>InP: roceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Online and Punta Cana</source>
          , Dominican Republic: Association for Computational Linguistics,
          <year>2021</year>
          , pp.
          <fpage>185</fpage>
          -
          <lpage>194</lpage>
          . doi: https://doi.org/10.18653/v1/
          <year>2021</year>
          .emnlp-demo.
          <volume>2</volume>
          .2 [30] [31]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tapaswi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Stiefelhagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Urtasun</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Fidler</surname>
          </string-name>
          . “
          <article-title>MovieQA: Understanding Stories in Movies through Question-Answering”</article-title>
          .
          <source>I2n0:16 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          .
          <year>2016</year>
          , pp.
          <fpage>4631</fpage>
          -
          <lpage>4640</lpage>
          . doi: https://do i.
          <source>org/10</source>
          .1109/CVPR.
          <year>2016</year>
          .
          <volume>501</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Torabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>CourvilUlsei.ng Descriptive Video</surname>
          </string-name>
          <article-title>Services to Create a Large Data Source for Video Annotation Research</article-title>
          .
          <year>2015</year>
          . arXiv:
          <volume>1503</volume>
          .01070 [cs.CV].
          <string-name>
            <given-names>M.</given-names>
            <surname>Vauth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. O.</given-names>
            <surname>Hatzel</surname>
          </string-name>
          , E. Gius, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Biemann</surname>
          </string-name>
          . “
          <source>Automated Event Annotation in Literary Texts.” In:Chr</source>
          .
          <year>2021</year>
          , pp.
          <fpage>333</fpage>
          -
          <lpage>345</lpage>
          . url: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2989</volume>
          /short%5C%
          <fpage>5</fpage>
          <lpage>Fpaper18</lpage>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>G.</given-names>
            <surname>Vercauteren.</surname>
          </string-name>
          “
          <article-title>A narratological approach to content selection in audio description: towards a strategy for the description of narratological time”M.Inon:TI</article-title>
          . Monografıás de Traducción e Interpretación 4 (
          <year>2012</year>
          ), pp.
          <fpage>207</fpage>
          -
          <lpage>231</lpage>
          . doi: https://doi.org/10.6035/MonTI.20 12.4.9.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>P.</given-names>
            <surname>Vicol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tapaswi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Castrejon</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Fidler</surname>
          </string-name>
          . “
          <article-title>MovieGraphs: Towards Understanding Human-Centric Situations from Videos”</article-title>
          .
          <source>In2:018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>2018</year>
          , pp.
          <fpage>8581</fpage>
          -
          <lpage>8590</lpage>
          . doi: https://doi.org/10.1109/CVPR.20
          <volume>18</volume>
          .00895.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>“A Graph-Based Framework to Bridge Movies and Synopses”</article-title>
          .
          <source>In2:019 IEEE/CVF International Conference on Computer Vision</source>
          (ICCV).
          <year>2019</year>
          , pp.
          <fpage>4591</fpage>
          -
          <lpage>4600</lpage>
          . doi: https://doi.org/10.1109/ICCV.
          <year>2019</year>
          .
          <volume>0046</volume>
          9.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yi</surname>
          </string-name>
          , G. Zhang, J. Liu, and
          <string-name>
            <surname>S. Zhang.</surname>
          </string-name>
          “
          <article-title>Movie Scene Event Extraction with Graph Attention Network Based on Argument Correlation Information”</article-title>
          .
          <source>SInen:sors 23.4</source>
          (
          <issue>2023</issue>
          ), p.
          <fpage>2285</fpage>
          . doi: https://doi.org/10.3390/s2304228 5.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Zacks</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Tversky</surname>
          </string-name>
          . “
          <article-title>Event structure in perception and conception</article-title>
          .”
          <source>InPs:ychological Bulletin 127.1</source>
          (
          <issue>2001</issue>
          ), pp.
          <fpage>3</fpage>
          -
          <lpage>21</lpage>
          . doi: http://doi.apa.org/getdoi.cfm?doi=10.1037/00 33-
          <fpage>2909</fpage>
          .
          <year>127</year>
          .
          <issue>1</issue>
          .3.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>