<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>and Narration in Contemporary Popular Fiction in Swedish - Stylometric Explorations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mats Dahllöf</string-name>
          <email>mats.dahllof@lingfil.uu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Uppsala University</institution>
          ,
          <addr-line>Uppsala</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <fpage>203</fpage>
      <lpage>211</lpage>
      <abstract>
        <p>A fundamental feature of many genres of fiction is the alternation between a narrative frame (NF) and quoted inset (QI) dialogue. Both formal and representational features distinguish NF and QI segments. This paper is an explorative study on the stylistic diferentiation between frame and inset material in recent commercially successful fiction in Swedish. There are mainly two orthographic options as regards this distinction: Explicitly enclosing inset segments within quotation marks is one. Using an initial dash to indicate utterance display is another, in which case frame and inset material typically alternate in a way not made explicit by the orthography. The corpus behind the present study comprised 450 novels. In order to deal with dash orthography data (135 books), we trained a multilayer perceptron classifier to tell NF and QI segments apart. We relied on the fact that native quotation mark text can be converted to annotated dash orthography data, which can then be used for supervised training and validation. A small-scale manual evaluation on the texts we aim to analyze, yielded an accuracy score around 95%. In order to explore the stylometric relations between NF and QI components in the novels, we looked at a selection of basic grammatical features. A characterization of each feature was made by means of recording the fraction of works in which the relative frequency of the feature is higher in QI than in NF. This summarizes how authors tend to “use” that feature to create a contrast between NF and QI. Another way to examine how the NF and QI styles are related is to apply a correlation test. We then saw, for instance, that QI material in 100% of the books are denser in auxiliary verbs, second person pronouns, and interjections, while NF segments in all or almost all cases are denser in nouns, adjectives, third person pronouns, and prepositions. We could also observe that e.g. noun density in NF and QI correlate in a strong way. The same holds for adverbs and cardinal numerals. This suggests that books and authors exhibit stylistic tendencies which afect both narrator and the characters, as far as the kind of fiction we have studied go.</p>
      </abstract>
      <kwd-group>
        <kwd>Explorations</kwd>
        <kwd>ifction</kwd>
        <kwd>dialogue</kwd>
        <kwd>quotation</kwd>
        <kwd>quoted inset</kwd>
        <kwd>narration</kwd>
        <kwd>narratorial frame</kwd>
        <kwd>direct speech</kwd>
        <kwd>stylometry</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The present study is concerned with recent commercially successful fiction in Swedish. It will
focus on a stylometric exploration of the diferentiation between frame narration and quoted
inset material, typically dialogue. A component in this work was the training of a classifier
which can tell the two kinds of content apart in novels written in “dash orthography”, which
makes this a non-trivial task.
CEUR
Workshop
Proceedings</p>
      <sec id="sec-1-1">
        <title>Real life: author ⇒ reader</title>
      </sec>
      <sec id="sec-1-2">
        <title>Narratorial frame (NF): narrator ⇒ addressee</title>
      </sec>
      <sec id="sec-1-3">
        <title>Quoted inset (QI): character ⟺ character</title>
        <sec id="sec-1-3-1">
          <title>Level of action</title>
        </sec>
        <sec id="sec-1-3-2">
          <title>Level of fictional communication</title>
        </sec>
        <sec id="sec-1-3-3">
          <title>Level of nonfictional communication</title>
          <p>
            A fundamental feature of many genres of fiction is the alternation between a narrative frame
(NF) and quoted inset (QI) dialogue, see Figure 1. Dialogue is the mimetic, more or less verbatim
display – using the device of “direct speech” –, of utterances made by the characters in the
story as chains of events play out. QIs are distinguished by both formal and representational
features. One of the former is directness, i.e. “the inset’s syntactic and deictic independence of
the frame” [1, p. 111]. The distinction between quoted dialogue and other modes of narration is
not always a matter of a simple binary opposition, but the analysis here will treat it as such.
This simplification is for the most part adequate for the kind of fiction we analyse here. It is also
possible for authors to use indirect speech and free indirect speech – as illustrated in examples
(3) and (4), respectively, below –, to mention the two most well-known alternatives. Speech
acts can also be described in ways that only summarize their content or focus on other aspects
(cf. e.g. [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] and [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ]). Furthermore, we should remember that both narration and quotation –
and indeed any kind of communication – can be embedded in QI material.
          </p>
          <p>
            In this article we will address two research questions in an exploratory fashion. The first one
is instrumental: How well can we separate inset and frame text automatically? Secondly, we
will investigate whether and how a systematic exploration of stylometric features can be used to
ifnd and illustrate diferences between inset and frame text as well as correlations between the
two embedded styles over a corpus of novels. The method and aims behind this study belong to
the school of “distant reading […] focus[ing] on units that are much smaller or much larger
than the text: devices, themes, tropes” [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ].
          </p>
          <p>In Swedish fiction, there are two main orthographic options for diferentiating between
NF and QI material. The most common and most explicit one is enclosing quoted segments
within quotation marks (””). The quotation marks serve as an explicit markup unambiguously
indicating the start and end of directly quoted material. Example (1) illustrates this, with the
oficial English translation exhibiting the explicit style. Using a dash (–) to indicate the start of
an utterance display paragraph – as in the original (2) – is a less explicit and somewhat less
common (30% of the books in our corpus, see below) orthographic device.</p>
          <p>(1) ”För många träfar”, säger han. ”Tiden rinner ut.” [Adapted.]</p>
          <p>
            “Too many results,” he says. “Time’s running out.” [Oficial English translation. 1]
1Lars Kepler, Stalker, Alfred A. Knopf, 2016. Translation by Neil Smith.
(2) – För många träfar, säger han. Tiden rinner ut. [Lars Kepler, Stalker, Bonniers, 2014.]
– Too many results, he says. Time’s running out. [Adapted.]
(3) He said that there were too many results and that time was running out. [Adapted.]
(4) [He was pessimistic.] There were too many results. Time was running out. [Adapted.]
In dash-style novels we often find an unmarked alternation between frame and inset segments
in the same paragraph, which would have been unambiguous had quotation marks been used.
So, the “inquit” formula he says in (2) belongs to the NF, but could also – quite easily if you
disregard the wider context – be understood as being part of the QI. More extended frame
material surrounded by quotation is known as “suspended quotation” [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. Identifying the spans
of QI segments in dash-style fiction consequently relies on semantic and pragmatic factors and
requires an interpretational efort from the reader. It also constitutes a non-trivial problem for
engineering in natural language processing (NLP).
          </p>
          <p>Supplementary Materials providing details on the corpus and a more complete presentation
of results is available at URL https://github.com/mdahllof/dhnb2022sm.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Previous Research</title>
      <p>
        The analysis of dialogue and narrative goes back to Plato’s opposition between μίμησις and
διήγησις, or between showing and telling, in modern critical parlance. Sternberg [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] lists
ifve representational features associated with direct speech and with the idea of μίμησις as
“speak[ing] in the person of another” [1, p. 111]: It is empathetic, specific, realistic, distinctive,
and reproductive. Quoted speech “is conventionally understood to replicate exactly what the
quoted character is supposed to have said” [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. By contrast, two of the most salient features
of narratives are their temporal relations to the events of the story and the “voice” of the
narrator. Authors thus have a range of options relating to the narration, e.g. first or third person
perspective and past or present tense. (See [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for an extended discussion.)
      </p>
      <p>
        Still, it has often been assumed in corpus-based literary studies that fiction is one register
and that a novel, for instance, exhibits one style. Egbert and Mahlberg [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] reject this idea. They
use multidimensional analysis involving factor analysis on a corpus of 19th-century English
novels to show that there are “extreme diferences between the linguistic characteristics of
ifctional speech and narration” [p. 86], which are two register categories typically “interspersed
throughout a text” [p. 98].
      </p>
      <p>
        Quoted inset recognition as a problem in NLP has been addressed in a number of studies,
which have been motivated by the central importance of the task for computational literary
studies. Recent work, e.g. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], rely on machine learning but also involve
rule-based procedures. A closely related issue for NLP engineering is identifying the speakers
behind inset segments [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <p>
        The corpus behind the present study comprises a collection of bestsellers and beststreamers in
Swedish, including translations, in the Swedish book market 2015–2020, along with a batch of
original productions in Swedish from the online streaming platform Storytel. The bestsellers
are the novels which have earned that status according to the Swedish Publishers’ Association
(SvF), either in hardback or paperback. As beststreamers (cf. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]) we count the
novels (excluding the categories history, biographies, teens and young adult, and children’s
books) which have been among the twenty most streamed audiobooks each year in the Storytel
platform. Storytel Original works are primarily produced for online streaming consumption
(“born-audio”), but are also made available as e-books. The current corpus includes all adult
ifction in Swedish published by Storytel Original from its inception in 2016 until May 2021. 2
The corpus is thus not a sample, but a complete dataset given the parameters defining it. As
regards genre, we distinguished crime, prestige (which does not overlap with crime for any
of the books we have looked at), and other genres. See [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for details on genre assignment.
Table 1 gives an overview of the composition of the corpus. A complete list is provided in the
Supplementary Materials.
      </p>
      <p>In the classifier training and the stylometric analysis we distinguished the following three
categories of paragraph.</p>
      <p>• NF only paragraphs, i.e. those not beginning with a dash (–) or containing quotation
marks (”). Henceforth: NF paragraphs.
• Typical QI paragraphs, i.e. those beginning with QI material as indicated by an intial
quotation mark or dash. They are often QI only, but there is often an alternation between QI
and NF text. Henceforth: QI/NF paragraphs. (Our classifier targeted the QI/NF separation
in the dash-orthography variety of such paragraphs.)
• Other paragraphs, i.e. those with non-initial explicitly quoted text or non-matching
quotation marks. They were excluded from the quantitative analysis.</p>
      <p>
        The machine learning design, as well as the stylometric analysis, departed from data tagged
with part-of-speech (POS) labels. We used the Stanza [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] system and its “talbanken 1.0.0” [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
model for Swedish3 to tag the texts.
2Many of these works are available in “seasons” comprising 10–24 episodes. The length of each season is roughly
that of a novel, and we concatenated all the episodes of each season into one document in this study. The only
Storytel Original among the beststreamers was a first episode of a season, Svart stjärna – S1E1 (2016), by Jesper
Ersgård and Joakim Ersgård. That text is counted as a part of a Storytel Original in the corpus.
3https://stanfordnlp.github.io/stanza/.
4. Quoted Inset vs Narratorial Frame Classification
Separating QI and NF text in dash-style QI/NF paragraphs is a classification problem.
Transforming text written in the quotation mark style, e.g. (1), into the dash style, e.g. (2), is for the
most part trivial. This means that native quotation mark text can be converted into annotated
dash-style data. We can use such minimally artificial data for supervised training and validation
of classifiers working on dash style documents (as do e.g. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]). These classifiers can
then be used to separate NF and QI text in paragraphs originally written in the dash style.
      </p>
      <p>We did not explore the possibilities for classifier design in depth, but managed to find an
approach which can be considered satisfactory, at least for the purpose at hand.</p>
      <p>The paragraphs were segmented at punctuation marks, where QI/NF shifts may occur. This
leaves us with short snippets to be classified. Features were derived from both the snippet text
and the preceding narration context, and involved POS unigrams and bigrams, word forms, as
well as pronoun person and verb tense. This allows the classifier to take advantage of both
lexical information and point of view diferences between the narration and the dialogue. A
multilayer perceptron with three layers, each comprising 15 neurons, was found to be the
optimally performing architecture, with 800 features selected. (The classifier pipeline can be
found among the Supplementary Materials.)</p>
      <p>
        We applied a cross-validation setup which was three-fold over 315 books (see first row in
Table 1) taking, for training, 80,000 randomly selected snippets from 210 books and and as
testing data all snippets from all dialogue paragraphs from the 105 other books. We then saw
accuracy scores from 94.2% to 94.6% (on word token level: 93.2% to 93.4%).4 A small-scale
manual evaluation of the best-performing of the three classifiers on the kind of data we actually
aim to analyze, i.e. dash-orthography novels, yielded an accuracy score of 95.7%, based on 282
randomly selected snippets from the 135 relevant novels in our corpus. This performance is
dificult to compare with previous approaches, but we can note that [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] hold that “an accuracy of
0.9 is remarkable”. The resulting classifier was then used to tag the 135 books for the stylometric
analysis discussed below.
      </p>
    </sec>
    <sec id="sec-4">
      <title>5. Stylometric Exploration</title>
      <p>
        In order to explore the stylometric relations between NF and QI components in the novels, we
took a selection of very basic grammatical features (listed in Table 2) as our point of departure.
They comprise parts of speech, pronouns according to grammatical person, and negation (only
inte). These were quantified as relative frequency among word tokens. We also included the
most important inflectional categories, quantified as ratio in tokens of the applicable part of
speech. (The full table of novels times feature values is available in the Supplementary Materials.)
A survey of a population that includes all elements defined by a set of criteria, like the present
one, does not face a problem of sampling bias. This means that simple descriptive statistics are
relevant, while significance testing is not applicable as in sample-based studies [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
4On the level of single books, the score is between 79.5%, for Kvinna inför rätta (Apple Tree Yard) by Louise Doughty,
and 99.5% for Playground by Lars Kepler and Tårtgeneralen by Filip Hammar and Fredrik Wikingsson. The Doughty
novel is a first/second person singular narrative in present tense, while the two with the highest scores are in a
third person past tense NF.
      </p>
      <p>A direct way of visualizing the QI and NF values for the various books and features is to
scatter-plot them against each other as in Figure 2, which shows four examples of features.
The NF values for preterite clearly distribute the works bimodally, forming a larger high-value
cluster and a smaller low-value cluster. The QI values are clearly less dispersed, suggesting
that authors to a high extent “agree” on the proper average amount of preterite in fictional
speech. This does not correlate to any visible degree with the preterite density in the NF. (The
plot for present tense, see Supplementary Materials, presents a similar shape upside down.)
Looking at another plot, we see that the noun density is quite dispersed both in NF and QI.
However, the NF value is always higher than the QI value, even if the ranges of possible values
overlap. Furthermore, the QI and NF values correlate. In other words, there is a tendency four
noun-dense narrators to be coupled with noun-dense characters. We see the same kind of
correlation for adverbs, but in most books (86% of them) the QI material is denser in adverbs
than the NF. Finally, in Figure 2, there is a plot for negation only showing books by the ten most
prolific Swedish authors in the corpus. The distribution is roughly as for adverbs in general, but
we also see that most authors cluster in particular regions, i.e. they use negation consistently in
their books in both NF and QI. Lars Kepler, for instance, is low-negation in NF, but high in QI.
Sofie Sarenbrant is high-negation in both NF and QI, even extremely so in NF.
Fractions of works in which the QI value exceeds the NF value for the features explored and Spearman’s
Feature
p1 (first person pronoun)
p2 (second person pronoun)
p3 (third person pronoun)</p>
      <p>QI &gt; NF (%)
neg(ation)
VBPRT (preterite)
VBPRS (present)
VBSUP (supine)
VBIMP (imperative)
VBINF (infinitive)
VBSFO (s-form)
NNIND (indefinite)
NNDEF (definite)
NNSIN (singular)
NNPLU (plural)
87
100
1
99
19
78
9
100
65
100
4
0
53
34</p>
      <p>0.19
0.22
0.29
0.53
0.10
0.08
0.44
0.32
0.47
0.56
0.58
0.58
0.58
0.60</p>
      <p>A straightforward characterization of the “behaviour” of each feature was made by means of
recording the fraction of works in which the relative frequency of the feature is higher in QI
than in NF (i.e. above the diagonals in Figure 2). This score summarizes how authors tend to
“use” that feature to create a contrast between NF and QI material. Another way to examine
how the NF and QI styles are related is to apply a correlation test. Not wishing to assume that
these variables are normally distributed, we used Spearman’s correlation coeficient (  ), which
is rank-based and consequently non-parametric (i.e. not based on a normality assumption).
Noun and adverb densities, as shown in Figure 2, will yield quite high  values.</p>
      <p>Table 2, whose values are plotted in Figure 3, shows the results. We find, as it were, a left
column of NF-oriented features and a right one of QI-oriented ones. In eight cases the tendencies
hold for all novels (0% or 100%) in the corpus. The NF-oriented features comprise elements of
general and definite nominal reference, prepositions, passive and past tense verb forms, and
numerals. The QI material, by contrast, is generally denser in interjections, first and second
person pronouns, and present tense and imperative verb forms. Nouns are indefinite to a higher
degree than in the NF. A plausible explanation is that there is less of anaphoric reference to
already introduced entities. This is as can be expected from dialogue. The higher incidences
of auxiliary verbs, adverbs, including negation, and subordinating conjunctions does not as
immediately chime with what can be expected.</p>
      <p>We see that there is a positive correlation between QI and NF values for all features, but the
tendency is of a stronger kind ( &gt; 0.55 ) as regards nouns, cardinal numerals, adverbs, and the
inflectional categories of nouns. The lowest degree of correlation is seen in features related
to deictic reference, i.e. tense and grammatical person, whose use in narration or dialogue is
based on pragmatic principles.</p>
    </sec>
    <sec id="sec-5">
      <title>6. Discussion and Conclusions</title>
      <p>
        We have introduced a method for visualizing and quantifying how authors use various
grammatical features to diferentiate between NF and QI material in fiction. Our observations clearly
agree with the findings of Egbert and Mahlberg [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], according to whom present tense and first
and second person are associated with dialogue, while there is more of past tense and third
person in narration. These results are far from unexpected and confirm the sanity of their and
our methods. Our analysis clearly revealed tendencies which are strong, but also of a fairly
abstract kind. There is surely much more to find in the data.
      </p>
      <p>We have found that NF and QI segments can be told apart with a high degree of accuracy by
means of a multilayer perceptron. It should be stressed that there is a potential methodological
danger in using a classifier that to a large extent rely on the same kind of features that are
explored in the study. As always, patterns revealed by corpus-based “distant reading” should be
seen as an invitation to look closer at the data.</p>
      <p>Another open and interesting issue is to what extent the choice of QI orthography influences
other stylistic options. The absence of quotation marks in the dash orthography is likely to
prompt authors to use other means of making the QI status of segments obvious to the readers.</p>
      <p>The present study has, in a sense, provided a map to the comfort zones of Swedish popular
ifction and to the vast and fascinating regions of stylistic space this genre tends to shun.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was carried out in the project “Patterns of Popularity: Towards a Holistic
Understanding of Contemporary Bestselling Fiction” funded by Vetenskapsrådet (2019-02829). PI Karl
Berglund contributed to the design and data curation for the present study.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sternberg</surname>
          </string-name>
          ,
          <article-title>Proteus in quotation-land: Mimesis and the forms of reported discourse</article-title>
          ,
          <source>Poetics Today</source>
          <volume>3</volume>
          (
          <year>1982</year>
          )
          <fpage>107</fpage>
          -
          <lpage>156</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Jahn</surname>
          </string-name>
          ,
          <article-title>Narratology 2.3: A guide to the theory of narrative, 2021</article-title>
          . URL: https://www. uni-koeln.de/~ame02/pppn.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rimmon-Kenan</surname>
          </string-name>
          , Narrative Fiction: Contemporary Poetics, Routledge, London and New York,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>McHale</surname>
          </string-name>
          ,
          <article-title>Speech representation</article-title>
          , in: P. Hühn,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Schmid</surname>
          </string-name>
          , J. Schönert (Eds.),
          <article-title>the living handbook of narratology</article-title>
          , Hamburg University, Hamburg,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Moretti</surname>
          </string-name>
          ,
          <source>Conjectures on world literature, New Left Review</source>
          <volume>1</volume>
          (
          <year>2000</year>
          )
          <fpage>54</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Herrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piper</surname>
          </string-name>
          ,
          <article-title>Computational stylistics</article-title>
          , in: D.
          <string-name>
            <surname>Kuiken</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jacobs</surname>
          </string-name>
          (Eds.),
          <article-title>Handbook of Empirical Literary Studies</article-title>
          , De Gruyter Reference,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Egbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mahlberg</surname>
          </string-name>
          ,
          <article-title>Fiction - one register or two? speech and narration in novels</article-title>
          ,
          <source>Register Studies</source>
          <volume>2</volume>
          (
          <year>2020</year>
          )
          <fpage>72</fpage>
          -
          <lpage>101</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F.</given-names>
            <surname>Jannidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Konle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zehe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hotho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krug</surname>
          </string-name>
          ,
          <article-title>Analysing direct speech in german novels</article-title>
          ,
          <source>in: 5. Tagung des Verbands Digital Humanities im deutschsprachigen Raum</source>
          ,
          <source>DHd</source>
          <year>2018</year>
          , Köln, Germany,
          <source>February 26 - March 2</source>
          ,
          <year>2018</year>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wirén</surname>
          </string-name>
          ,
          <article-title>Distinguishing narration and speech in prose fiction dialogues</article-title>
          ,
          <source>in: Proceedings of the Digital Humanities in the Nordic Countries 4th Conference</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>124</fpage>
          -
          <lpage>132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kurfalı</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wirén</surname>
          </string-name>
          ,
          <article-title>Zero-shot cross-lingual identification of direct speech using distant supervision</article-title>
          ,
          <source>in: Proceedings of LaTeCH-CLfL</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>105</fpage>
          -
          <lpage>111</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Byszuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Woźniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Leśniak1</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Łukasik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Šeļa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eder</surname>
          </string-name>
          ,
          <article-title>Detecting direct speech in multilingual collection of 19th-century novels</article-title>
          ,
          <source>in: Proceedings of 1st Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>100</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D. K.</given-names>
            <surname>Elson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. R.</given-names>
            <surname>McKeown</surname>
          </string-name>
          ,
          <article-title>Automatic attribution of quoted speech in literary narrative</article-title>
          ,
          <source>in: Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence (AAAI-10)</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>1013</fpage>
          -
          <lpage>1019</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Berglund</surname>
          </string-name>
          ,
          <article-title>Introducing the beststreamer: Mapping nuances in digital book consumption at scale</article-title>
          ,
          <source>Publishing Research Quarterly</source>
          <volume>37</volume>
          (
          <year>2021</year>
          )
          <fpage>135</fpage>
          -
          <lpage>151</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>K.</given-names>
            <surname>Berglund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dahllöf</surname>
          </string-name>
          ,
          <article-title>Audiobook stylistics: Comparing print and audio in the bestselling segment</article-title>
          ,
          <source>Journal of Cultural Analytics</source>
          <volume>11</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          . doi:
          <volume>10</volume>
          .22148/001c.
          <fpage>29802</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bolton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>Stanza: A python natural language processing toolkit for many human languages</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1001</fpage>
          -
          <lpage>1008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Nivre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>UD Swedish</given-names>
            <surname>Talbanken</surname>
          </string-name>
          ,
          <year>2008</year>
          . URL: https://universaldependencies.org/ treebanks/sv_talbanken/index.html.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Hirschauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Grüner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mußhof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Becker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jantsch</surname>
          </string-name>
          ,
          <article-title>Can p-values be meaningfully interpreted without random sampling?</article-title>
          ,
          <source>Statistics Surveys</source>
          <volume>14</volume>
          (
          <year>2020</year>
          )
          <fpage>71</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>