<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detecting Narrativity Across Long Time Scales</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrew Piper</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sunyam Bagga</string-name>
          <email>sunyam.bagga@mail.mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Monteiro</string-name>
          <email>laura.monteiro@mail.mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrew Yang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marie Labrosse</string-name>
          <email>marie.labrosse@mail.mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yu Lu Liu</string-name>
          <email>yu.l.liu@mail.mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>McGill University</institution>
          ,
          <addr-line>688 Sherbrooke St, H2J3B2 Montreal</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <fpage>319</fpage>
      <lpage>332</lpage>
      <abstract>
        <p>Storytelling is a universal human practice that serves as a key site of education, collective memory, fostering social belief systems, and furthering human creativity. It can occur in diferent discursive domains for diferent social purposes with difering degrees of intensity. In this project, we develop computational methods for measuring the degree of narrativity in over 335,000 text passages distributed across two- to three-hundred years of history and four separate discursive domains (fiction, non-fiction, science, and poetry). We show how these domains are strongly diferentiated according to their degree of narrative communication and, second, how truth-based discourse has declined considerably in its utilization of narrative communication. These findings suggest that there has been a long-term historical diferentiation between the practices of knowing and telling, which raises important questions with respect to the social acceptance of both science and the arts.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;narratology</kwd>
        <kwd>history</kwd>
        <kwd>systems theory</kwd>
        <kwd>discourse analysis</kwd>
        <kwd>computational narrative studies</kwd>
        <kwd>digital humanities</kwd>
        <kwd>natural language processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        in the text, just as ostensibly non-narrative documents, such as scientific reports, may also
exhibit degrees of narrativity. Herman [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] has taken this understanding one step further to
suggest that narrativity is not simply a matter of the local interplay of formal and textual
features, but emerges through the interaction between readers and texts. Narrativity can thus
be understood as a potentially rising or falling quality within documents (or other forms of
communication) that depends on the interaction of diferent linguistic or semiotic features
combined with readers’ responses.
      </p>
      <p>
        While a wealth of recent work in the field of natural language processing has engaged with
the detection of diferent dimensions of narrativity (such as causal and temporal relations
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], turning points [
        <xref ref-type="bibr" rid="ref21">21, 2</xref>
        ], reportable events [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], frames [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], etc.), no work to our knowledge
has undertaken the more elementary task of narrativity detection itself. Can we reliably
predict whether a span of text is engaging in narrative communication and if so, to what
degree of intensity? Such work has the potential not only to contribute to our theoretical
understanding of narrativity as a form of communication. It can also provide empirical insights
into the distribution of narrativity across diferent discursive domains and time periods giving
us insights into the social functions of narrative communication over time. The latter will be
our concern here.
      </p>
      <p>
        In this paper, we develop computational models to detect “narrativity” as a local,
multidimensional textual quality across four diferent discursive domains over a two- to
threehundred year time-period. Our aim in doing so is to test the relationship between narrative
communication and the process of functional diferentiation among social systems as theorized
by the sociologist Niklas Luhmann [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. According to Luhmann, social systems are governed
by communicative practices (“codes”) that maintain a system’s internal coherence in
distinction from other systems. As societies modernize, diferentiation strengthens over time as each
system evolves to maintain its internal coherence in distinction from its environment (i.e. other
systems).
      </p>
      <p>
        The question we wish to test here is whether the practice of narrative participates in this
process of functional diferentiation between the social systems Luhmann labels “art” and
“science.” According to Luhmann, art’s function lies in its ability to communicate the sensory
process of observation, to “allow a world to appear within the world” (p. 241), whereas the
function of science is to “structure the field of possible statements [about the world] with the
help of the code true/untrue” (p. 227) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. As a form of communication strongly associated
with the idea of “world-building” [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], we would thus expect narration to be highly associated
with artistic forms of expression but not necessarily negatively associated with scientific
discourse. There is nothing intrinsic to narrative communication that makes it an inappropriate
vehicle for fact-based discourse. After all, one can tell true and untrue stories.
      </p>
      <p>
        And yet according to the data and models used here, we can observe a very clear historical
trajectory of the de-narrativization of truth-based discourse. Our findings bring to light
longstanding and growing tensions between what Hayden White first introduced as the relationship
between knowing and telling [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. For White, the function of narration should be understood
as “a solution to a problem of general human concern, namely, the problem of how to
translate knowing into telling, the problem of fashioning human experience into a form assimilable
to structures of meaning that are generally human rather than culture-specific” (p. 5) [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
Narration is a key communicative mode for White that makes knowledge “assimilable” to
individual human beings and collective societies. As a growing body of research has indicated,
narrative is indeed an efective means of addressing the problem of knowledge sharing across
a variety of social domains, from economics [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] to climate change [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to political polarization
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>Our findings, preliminary as they are, suggest the need for further research into this
disassociation of science and narration and its potential social efects. Is the growing public distrust
in science related to the denarrativization of scientific communication? Do eforts of “public
science” or science journalism have a positive efect on reversing public distrust and are such
efects related to their degree of narrativity? If narration is increasingly seen as belonging
to the domain of art, has this diferentiation from science unintentionally contributed to the
devaluation of the arts (or their study)? How might the arts instead participate more explicitly
in the process of knowledge transfer, i.e. help “recouple” in Luhmann’s terms the practices of
knowing and telling?</p>
      <p>In order to detect narrativity in our historical collections, we undertake the following steps,
which we describe in greater detail in the following sections. First, we construct a data set
of 335,245 documents to represent our two primary social systems of art and science, which
we subdivide into four domains of fiction, poetry, science, and non-fiction. We then develop
a working theory of “narrativity” drawn from existing theoretical literature that informs our
manual annotation of the data. Building a team of three trained student annotators, we
handannotate 401 passages according to a scalar understanding of narrativity derived through
numerous meetings and discussion. This data is then used to train and test our machine
learning models, which we describe in Section 3. We present our results in Section 4 and
include a discussion of their potential implications as well as limitations (Section 5). Finally,
we conclude with a brief discussion of where future work in computational narrative studies
might lead.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data</title>
      <sec id="sec-2-1">
        <title>2.1. Annotated Data</title>
        <p>
          In order to annotate training data for the presence of narrativity, we rely on the following
theoretical schema developed by Herman [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. According to this schema, narrative communication
consists of the following four elements:
        </p>
        <sec id="sec-2-1-1">
          <title>1. Situatedness: narrativity depends on the social context in which it occurs</title>
          <p>
            2. Event sequencing: narrativity depends on temporally ordered events
3. World making: narrativity depends on the fact of disequilibrium such that we can observe
a change in the world
4. Feltness: narrativity captures the experience of events, i.e. “what it is like”
Herman’s categories can be seen as syntheses of previous narratological frameworks,
capturing a good degree of consensus in the field. The emphasis on feltness, for example, is strongly
indebted to the argument by Fludernik [7] that “Experientiality reflects a cognitive schema
of embodiedness that relates to human existence and human concerns” (p. 9). Similarly,
event-sequencing is strongly indebted to the work of theorists like Genette [9], Sternberg [
            <xref ref-type="bibr" rid="ref26">26,
27</xref>
            ], and Ricoeur [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ] and their emphasis on temporality as a central component of narrative
communication, while world making derives from the work of Labov and Waletzky [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ] and
Bruner [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ].
          </p>
          <p>In general then, Herman’s model is guided by the notion that, “Narrative roots itself in the
lived, felt experience of human or human-like agents interacting in an ongoing way with their
cohorts and surrounding environment” (our emphasis). Thus for Herman what matters most
about narrativity is: a) the centralization of one or more agents; b) the sequencing of events
and thus time; and finally, c) the idea of “lived experience in an environment”, i.e. a sense of
world building.</p>
          <p>Based on this theoretical framework, we hand-annotate 401 passages drawn from the
experimental data using the following steps:</p>
          <p>First, we assembled a team of three annotators who all have majors in the humanities. These
are readers who have high levels of education and exposure to training in textual analysis.
Second, over the course of several weeks we engaged in discussions and experiments regarding
the concept of “narrativity” with respect to the theoretical framework discussed above as well as
diferent kinds of text passages. These discussions culminated in a codebook, which is included
in the supplementary material.1 Annotators were then asked to code a given passage across
three dimensions of narrativity, which were defined for the annotators as “agency,” “event
sequencing,” and “world making.” Note how we translated Herman’s “feltness” into “agency”
to better account for the idea of experientiality at the heart of most major narrative theories.</p>
          <p>For each passage, readers were asked to respond to the following statements using a five-point
Likert scale:
• “This passage foregrounds the lived experience of particular agents.” (Agency)
• “This passage is organized around sequences of events that occur over time.” (Event
sequences)
• “This passage creates a world that I can see and feel.” (World making)</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>1. Strongly disagree</title>
          <p>2. Somewhat disagree
3. Unsure
4. Somewhat agree
5. Strongly agree</p>
          <p>Notice how we do not expressly ask if readers felt that the passage was “narrative” or not.
Rather, we ask them to consider their feelings with respect to these three primary narrative
dimensions, which we then average into a single “narrativity” score. We found that this increased
reader agreement and allowed for more nuanced understandings of narrative communication.
For example, it was not uncommon for some types of discourse to emphasize sequential events
but lack an emphasis on agency or building a world.</p>
          <p>We provide a few sample passages that received low and high average narrativity scores by
readers. Note that passages have been truncated from their actual length.</p>
          <p>Non-Fiction - Average Reader Score 1.2 The employment of the uninterpretable symbol
in the intermediate processes of trigonometry furnishes an illustration of what has been said.
Lapprehend that there is no mode of explaining that application which does not covertly assume
the very principle in question. But that principle, though not, as I conceive, warranted by
formal reasoning based upon other grounds, seems to deserve a place among those axiomatic
1Note that we provide the reader-annotated data, annotator’s codebook, metadata, code for all models and
concrete implementation details of custom features from Table 1 in the Supplementary Material. It is available
at https://doi.org/10.7910/DVN/DAWVME
truths which constitute in some sense the foundation of general knowledge, and which may
properly be regarded as expressions of the mind’s own laws and constitution.
Fiction - Average Reader Score 1.44 It is too weak for a shield, too transparent for a screen,
too thin for a shelter, too light for gravity, and too threadbare for a jest. The wearer would be
naught indeed who should misbeseem such a wedding garment. But wherefore does the sheep
wear wool? That he in season sheared may be, And the shepherd be warm though his flock be
cool.</p>
          <p>Science - Average Reader Score 4.55 I assisted at the opening of her Body, and having found
in the matrix a little round mass of the bigness of a great black Cherry, I took the husband
aside, and asked him, Num a tempore fluxus menstruorum uxorem cognevisset? And having
received for answer, that he had, I prayed him to let me carry home with me this little ball,
which I had found in her womb. I was no sooner come home but I opened it, and found, that
nature had wrought with so much activity in so small a time...</p>
          <p>Fiction - Average Reader Score 4.55 Whereupon a sudden outcry arose within the house,
and a head popped angrily out of the aperture so suddenly created. But as instantly it returned
within. For Jorian tossed the lattice to the ground by the door and thrust his spear-head into
the cravat of red which the man had about his throat, shouting to him all the while in the name
of the Prince, of the Duke, of the Emperor, of the Archbishop, of all potentates, lay and secular,
to come down and open the gates.</p>
          <p>
            Because our annotations use a multi-point scale, we assess inter-rater reliability (IRR) using
the average deviation index as discussed in Burke, Finkelstein, and Dusig [
            <xref ref-type="bibr" rid="ref2">4</xref>
            ]. We report an
average deviation of 0.48 (± 0.27). This indicates that on average our annotators’ judgments
per passage fall within just under 0.5 points of each other on our 5-point Likert scale, suggesting
reasonable levels of agreement. A one-way ANOVA was conducted to compare the efect of
genre on average deviation among annotators, with a significant efect observed [F (3, 397) =
7.56, p = 6.26e − 05], with poetry generating significantly more deviation among annotators
than the other genres (mean AD of 0.58). We also note that as seen in Figure 1, annotator
scores were were not normally distributed around 3.0, but rather exhibit a skewed central
tendency between 2.0 and 2.5. 65% of the annotations were below 3, suggesting there were
only a minority of passages where our annotators were confident of the passage’s narrativity.
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Experimental Data</title>
        <p>
          Our experimental data consists of five separate collections that are designed to capture the
two social systems of “art” and “science,” which we represent as the four discursive domains
of fiction, non-fiction, poetry, and science. Doing so allows us to see aggregate behavior across
the two systems as well as potential internal diferences based on discourse type. Our data
consists of:
• Fiction &amp; Non-Fiction. This data is derived from the Hathi Trust Digital Library and
is drawn from Piper and Bagga [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. It encompasses 85,130 passages of fiction and 99,968
passages of non-fiction spanning the years 1800-1999 written in English. The labels are
generated using modified predictive models based on prior work [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. The distribution
        </p>
        <p>
          of the number of passages per year is indicated in Figure 2. Years represent year of
publication, not year of composition or first printing. Our data reflects reading material
available in a given year as archived by academic libraries.
• Poetry. This data is drawn from the Literature Online Poetry database. It consists
of 73,077 poems by 857 authors who wrote in English and who were alive during the
nineteenth and twentieth centuries. To estimate year of publication, we use the author’s
birth-date plus 35 years to capture an estimated career midpoint. Because of the
relatively small number of poets in our dataset, we are not able to capture a consistent
number of poems per year.
• Science. To represent the domain of scientific writing, we use two diferent data sets. The
ifrst is drawn from the Royal Society Corpus (RSC 4.0) based on the first two centuries of
the Philosophical Transactions of the Royal Society of London from its beginning in 1665
to 1869 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Due to copyright restrictions, no data is publicly available after 1869. This
dataset consists of 31,698 documents. To augment this data, we use a collection of 45,439
randomly selected articles drawn from top 100 most common articles in the JSTOR Data
for Research platform organized under the heading “physical sciences” published between
the years 1900 and 2015. The distribution of articles over time is captured in Figure 2.
        </p>
        <p>
          Because our interest is in local narrativity, i.e. the extent to which a span of tokens expresses
narrative communication, we represent our documents as randomly selected sequences of 5
sentences in length. This number has been indicated in prior work as a reasonable frame in
which completed narratives can transpire [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. We can thus assume that “narrativity” can be
present in spans of this length. Future work will want to explore this parameter further.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Models</title>
      <p>For the purposes of our project, we use the predicted probability of a passage’s “narrativity”
as an indicator of the degree of narrative communication present in that passage. In order to
build a model to predict a passage’s narrativity, we train and validate our models using our
reader annotated data. We experiment with three widely-used algorithms (Logistic Regression,
Support Vector Machines, and Random Forests) and multiple combinations of diferent features
to identify the best-performing model. We present our feature components in Table 12 and
present the performance of each model using diferent feature combinations in Figure 3. As
can be seen in Figure 3, Random Forest performs the best out of the three learning algorithms.
Table 2 presents a brief overview of the top-5 performing models using Random Forest.</p>
      <p>We assess model performance according to Pearson’s correlation coefficient rather than the
more traditional F1 score (although we also report traditional classification metrics in Table
2). Because our metric of narrativity is predicted probability and not a binary classification,
the question we want to address is how well our models correlate with the scalar nature of
reader judgments.</p>
      <p>To construct our experimental feature spaces, we aggregate our features into three general
categories: lexical features (ngrams), syntactical features (part-of-speech and dependency
relationships), and higher-level custom features designed to capture specific narratological theories,
including time, concreteness, animate entities and perceptuality. For a full discussion of the
custom features, see the supplementary material.</p>
      <p>
        As we can see in Figure 3, all three classifiers behave similarly and achieve their maximum
performance on a variety of feature combinations. Interestingly, unigrams tend to perform
2Note that the results shown in Figure 3 correspond to a maximum of 100 features per category. This
is why # Features for word-bigrams, for example, is 100 although the complete feature space involved 25,434
word-bigrams. Experiments with other values of max-features yielded similar results.
better than the limited sets of bi- or trigrams for lexemes, pos, and dependency tags. The
best performing model consists of part-of-speech unigrams, % dialog and custom-built features
that aim to capture “event sequences”, “world building”, and “agency” for which we use the
categories tense, mood, and voice. There appears to be a strong grammatical signature to
narrativity that marginally grows in strength when we add in features that capture the notion
of “environment” emphasized in Herman’s theory above [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We leave a deeper exploration of
these issues to future work.
      </p>
      <p>As we can see in Figure 4, the correlation between reader judgments and predicted probability
is approximately linear and indicates a reasonable level of agreement (r = 0.742). We observe
higher levels of variability in the middle range of annotations between 2.0 and 3.0, which
is to be expected. As readers’ judgments become less certain, so too do we observe more
variability in our models’ predictions. Future work will want to explore the extent to which
more annotations lead to higher levels of correlation or whether we achieve some kind of
maximum level of correlation between computational models and human judgments in this
area.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>
        Applying our best performing model (Random Forest with POS-unigrams + Pct-Quoted +
tense, mood and voice features) on the experimental data described above, we generate the
average yearly predicted probability of narrativity across all four domains as shown in Figure
5. According to our models the four domains behave in distinctive fashion with respect to
narrativity, providing support for the idea that narrativity may be another facet underlying
Luhmann’s thesis about functional diferentiation [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Second, with respect to the fact-based
domains of nonfiction and science, we also observe meaningful decays in the estimated intensity
of narrativity over our time period. For science, we see a drop from an average five-year high
of 0.495 estimated narrativity in 1707 to a five-year low of 0.236 in 1994, while for non-fiction
we see a less dramatic decline from 0.428 (in 1844) to 0.338 (in 1996). Because our non-fiction
class can potentially contain scientific reports archived in Hathi, we cannot definitively tell if
this decline of narrativity in non-fiction is due to the growth of science writing in Hathi or the
decline of narrativity in non-scientific forms. While we provide some validation of this in the
discussion, future work in this direction will depend on more fine-grained genre classification
with respect to non-fictionality.
      </p>
      <p>In terms of our two “literary” domains, we see little change over time, suggesting relative
stability of these domains’ relationship to narrativity. While this does not run counter to
expectations with respect to fiction, theorists of poetry might be surprised to see such continuity
given the popularity of long narrative poems in the nineteenth century (for example in the work
of Walter Scott or Longfellow to name two prominent examples). Future work will want to
explore in greater depth whether there is a meaningful break with respect to poetic narrativity
for authors born after the late nineteenth-century that then potentially reverses course for
younger poets born closer to the end of the twentieth-century as indicated in Figure 5. More
domain-specific training data would be needed along with more careful sampling techniques to
gain confidence about any such shifts given how slight they are with respect to our models.</p>
      <p>Taken altogether, our models suggest that narrativity is strongly socially diferentiated across
diferent discursive domains and that at least with respect to fact-based discourses this
diferentiation is increasing strongly over time as both non-fiction writing and specifically scientific
writing exhibit declines in their reliance on narrative communication. We take up the
implications of these findings in our discussion.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>Our study raises a number of questions for future research. Representing the social systems of
“art” and “science” is a challenging task. In our work, we have tried to capture at least two
larger subdomains of writing within these systems to better understand the kind of internal
diferentiation may be at work. While future work will want to experiment with diferent
samples dependent on diferent archival resources, we do observe interesting diferences with
respect to the narrative behavior of our subdomains. For example, we see how the narrativity
of poetry is not only considerably lower than prose fiction, it consistently hovers around the
iffty-percent mark, suggesting that one of the potential social functions of poetry as a genre
could be its ability to communicate narrative ambiguity. Such ambiguity is corroborated by
the higher average deviation among our annotators with respect to the poetry training data.
This suggests an interesting potential theory one could pursue for the future study of poetry
in a larger social context along with the potential increase in narrativity that we observe for
poets born after the 1960s.</p>
      <p>Because our models have been trained to understand cross-domain behavior of narrativity,
our work cannot however speak to within-domain distinctions with respect to narrativity. For
example, an interesting question to be pursued in the future is the intensity and extent of
narrativity at the document level within our diferent discourses. When do we see novels for
instance engage in more explicitly “narrative” communication, are there reliable patterns of
the rise and fall of narrativity, or finally what kinds of novels (genres) indicate greater degrees
of narrativity overall? Similarly, for science documents while we observe an overall process of
denarrativization of scientific documents, are there still portions of articles that engage in more
narrative-like behavior or portions of the field (i.e. disciplines) that engage in more narrative
communication than others? These questions would help provide insights into the relationship
between categories like genre, discipline and narration.</p>
      <p>Further reflection could also be given to our framework of “truth-based discourse,” which
is not exactly synonymous with “science,” which is one of the reasons we also model
“nonifctional” writing as well. Scientific writing is one kind of communication that makes truth
claims, but there are numerous others that belong to diferent institutional frameworks. We
note that in a random sample of two-hundred passages drawn from our non-fiction experimental
data that the number of “scientific” texts moves from 6 in the nineteenth century to 10 in the
twentieth. While this represents a large increase, it is still a very small fraction of all writing in
our non-fiction sample suggesting that the decline in narrativity in non-fiction is not strongly
related to a rise of scientific writing in Hathi Trust. In other words, writing classified as
nonifction is exhibiting similar trends to science but is being produced in diferent institutional
contexts. Future work could explore more deeply why we see this denarrativization of
nonifction along with science writing.</p>
      <p>At the level of data annotation, while we demonstrate solid agreement between readers
and reasonable model correlation with reader judgments, we are not able to annotate large
amounts of data to better calibrate our models. Hand-annotation is a slow and expensive
process and future work will want to explore mechanisms that allow for scaling annotation
while maintaining quality. We assume model accuracy will increase with increased amounts
of annotated data, which may or may not have a bearing on the historical trends we observe
here. The observed declines in our science and non-fiction corpora are so steep and consistent
that we would be surprised if future work indicated significant changes in this regard.</p>
      <p>In terms of our theoretical framework, it is important to underscore that our approximation
of narrativity is just that. While we do not observe significant shifts in the distribution of
narrativity according to feature-space selection, our models are still guided by a particular
theoretical framework with respect to narrativity. Future work will want to explore alternative
theories and feature representations of narrativity to see if the historical trends we are observing
continue to emerge.</p>
      <p>Finally, future work will want to explore the extent to which our findings are or are not
culturally specific, i.e. the extent to which they hold in other language communities and the
extent to which those correlations are driven by social factors such as national wealth or
education levels. Just how universal is this process of functional diferentiation and denarrativization
with respect to truth-based discourses that we have observed here? Is this indeed a marker of
“modernization”?</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>Our work has demonstrated that narrative communication is a detectable linguistic quality
of texts from the perspective of human readers and machine learning. Even with a small set
of training data we can achieve reasonable levels of predictive accuracy and correlation with
trained reader judgments across very diferent kinds of texts over relatively long historical time
spans. We also show that with sufficient training readers can agree quite well regarding the
intensity of a passage’s narrativity.</p>
      <p>Being able to identify the degree of narrativity in large-scale historical document collections
allows us to gain a better understanding of the distribution of narrative communication across
documents that serve diferent social functions. Modeling narrativity at the computational level
suggests that narrative is a form of communication that participates in Luhmann’s theory of
functional diferentiation, at least with respect to the social systems of art and science. While
culturally and historically universal – narration is present in all linguistic communities and
recorded time periods – narration is far from being socially universal. Indeed, it appears that
in modern, highly diferentiated societies narration is increasingly aligned with the particular
social system of the arts as truth-based discourse becomes less and less narrativized over time.
How this may impact urgent large-scale questions such as trust in science or particular collective
responses to social problems such as climate change or health pandemics remains an open, yet
important question for future research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bal</surname>
          </string-name>
          . Narratology:
          <article-title>Introduction to the Theory of Narrative</article-title>
          . University of Toronto Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[4] [7] [8] [9]</source>
          [2]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Boyd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. G.</given-names>
            <surname>Blackburn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Pennebaker</surname>
          </string-name>
          . “
          <article-title>The Narrative Arc: Revealing Core Narrative Structures through Text Analysis”</article-title>
          .
          <source>In: Science Advances 6.32</source>
          (
          <year>2020</year>
          ),
          <year>eaba2196</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bruner</surname>
          </string-name>
          . “
          <article-title>The Narrative Construction of Reality”</article-title>
          .
          <source>In: Critical Inquiry 18.1</source>
          (
          <issue>1991</issue>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>M. J. Burke</surname>
            ,
            <given-names>L. M.</given-names>
          </string-name>
          <string-name>
            <surname>Finkelstein</surname>
            , and
            <given-names>M. S.</given-names>
          </string-name>
          <string-name>
            <surname>Dusig</surname>
          </string-name>
          . “
          <article-title>On Average Deviation Indices for Estimating Interrater Agreement”</article-title>
          .
          <source>In: Organizational Research Methods</source>
          <volume>2</volume>
          .1 (
          <issue>1999</issue>
          ), pp.
          <fpage>49</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bushell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Buisson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Workman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Colley</surname>
          </string-name>
          . “
          <article-title>Strategic Narratives in Climate Change: Towards a unifying narrative to address the action gap on climate change”</article-title>
          .
          <source>In: Energy Research &amp; Social Science</source>
          <volume>28</volume>
          (
          <year>2017</year>
          ), pp.
          <fpage>39</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N.</given-names>
            <surname>Chambers</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          . “
          <article-title>Unsupervised Learning of Narrative Event Chains”</article-title>
          .
          <source>In: Proceedings of ACL-08: HLT</source>
          . Columbus, Ohio: Association for Computational Linguistics,
          <year>2008</year>
          , pp.
          <fpage>789</fpage>
          -
          <lpage>797</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Fludernik</surname>
          </string-name>
          .
          <article-title>Towards a 'Natural' Narratology</article-title>
          . Routledge,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>G.</given-names>
            <surname>Genette</surname>
          </string-name>
          . “
          <article-title>Boundaries of Narrative”</article-title>
          .
          <source>In: New Literary History 8.1</source>
          (
          <issue>1976</issue>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>G.</given-names>
            <surname>Genette. Narrative Discourse</surname>
          </string-name>
          :
          <article-title>An Essay in Method</article-title>
          . Vol.
          <volume>3</volume>
          . Cornell University Press,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Giora</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          . “
          <article-title>Degrees of Narrativity and Strategies of Semantic Reduction”</article-title>
          .
          <source>In: Poetics 22.6</source>
          (
          <issue>1994</issue>
          ), pp.
          <fpage>447</fpage>
          -
          <lpage>458</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D.</given-names>
            <surname>Herman</surname>
          </string-name>
          . Basic Elements of Narrative. John Wiley &amp; Sons,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kermes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Degaetano-Ortlieb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khamis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Knappen</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Teich.</surname>
          </string-name>
          “
          <article-title>The Royal Society Corpus: From Uncharted Data to Corpus”</article-title>
          .
          <source>In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)</source>
          .
          <year>2016</year>
          , pp.
          <fpage>1928</fpage>
          -
          <lpage>1931</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kubin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Puryear</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schein</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Gray</surname>
          </string-name>
          . “
          <article-title>Personal Experiences Bridge Moral and Political Divides Better than Facts”</article-title>
          .
          <source>In: Proceedings of the National Academy of Sciences 118.6</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>W.</given-names>
            <surname>Labov</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Waletzky</surname>
          </string-name>
          . “
          <article-title>Narrative Analysis: Oral Versions of Personal Experience</article-title>
          .” In: (
          <year>1967</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Luhmann</surname>
          </string-name>
          . Die Kunst der Gesellschaft. Suhrkamp,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>N. Luhmann. Social</given-names>
            <surname>Systems</surname>
          </string-name>
          . Stanford University Press,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mostafazadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Chambers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Vanderwende</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kohli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Allen</surname>
          </string-name>
          .
          <article-title>“A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories”</article-title>
          .
          <source>In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          .
          <year>2016</year>
          , pp.
          <fpage>839</fpage>
          -
          <lpage>849</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mostafazadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Grealish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Chambers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Allen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Vanderwende</surname>
          </string-name>
          . “
          <article-title>CaTeRS: Causal and Temporal Relation Scheme for Semantic Annotation of Event Structures”</article-title>
          .
          <source>In: Proceedings of the Fourth Workshop on Events</source>
          . San Diego, California: Association for Computational Linguistics,
          <year>2016</year>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>61</lpage>
          . doi:
          <volume>10</volume>
          . 18653 / v1 /
          <fpage>W16</fpage>
          - 1007. url: https://www.aclweb.org/anthology/W16-1007.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ochs</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Capps</surname>
          </string-name>
          . Living Narrative:
          <article-title>Creating Lives in Everyday Storytelling</article-title>
          . Harvard University Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>McKeown</surname>
          </string-name>
          . “
          <article-title>Modeling Reportable Events as Turning Points in Narrative”</article-title>
          .
          <source>In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          . Lisbon, Portugal: Association for Computational Linguistics,
          <year>2015</year>
          , pp.
          <fpage>2149</fpage>
          -
          <lpage>2158</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D15</fpage>
          -1257. url: https://www.aclweb.org/anthology/D15-1257.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>P.</given-names>
            <surname>Papalampidi</surname>
          </string-name>
          , F. Keller, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Lapata</surname>
          </string-name>
          . “
          <article-title>Movie Plot Analysis via Turning Point Identification”</article-title>
          .
          <source>In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          .
          <source>Hong Kong</source>
          , China: Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>1707</fpage>
          -
          <lpage>1717</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1180. url: https://www.aclweb. org/anthology/D19-1180.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pianzola</surname>
          </string-name>
          . “
          <article-title>Looking at Narrative as a Complex System: The Proteus Principle”</article-title>
          .
          <source>In: Narrating Complexity</source>
          . Springer,
          <year>2018</year>
          , pp.
          <fpage>101</fpage>
          -
          <lpage>122</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Piper</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bagga</surname>
          </string-name>
          . HATHI 1M:
          <article-title>Million Page Historical Prose Data in English from the Hathi Trust</article-title>
          .
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ricoeur</surname>
          </string-name>
          .
          <source>Time and Narrative</source>
          , Volume
          <volume>1</volume>
          . University of Chicago press,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Shiller</surname>
          </string-name>
          . Narrative Economics:
          <article-title>How Stories Go Viral</article-title>
          and Drive Major Economic Events. Princeton University Press,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26] [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sternberg</surname>
          </string-name>
          . “
          <article-title>Telling in Time (I): Chronology and Narrative Theory”</article-title>
          .
          <source>In: Poetics Today 11.4</source>
          (
          <issue>1990</issue>
          ), pp.
          <fpage>901</fpage>
          -
          <lpage>948</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Sternberg</surname>
          </string-name>
          . “
          <article-title>Telling in Time (II): Chronology, Teleology, Narrativity”</article-title>
          .
          <source>In: Poetics Today 13.3</source>
          (
          <issue>1992</issue>
          ), pp.
          <fpage>463</fpage>
          -
          <lpage>541</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>T.</given-names>
            <surname>Underwood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kimutis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Witte</surname>
          </string-name>
          . “
          <source>NovelTM Datasets for English-Language Fiction</source>
          ,
          <fpage>1700</fpage>
          -
          <lpage>2009</lpage>
          ”.
          <source>In: Journal of Cultural Analytics</source>
          <volume>5</volume>
          .2 (
          <issue>May</issue>
          28,
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .22148/ 001c.
          <fpage>13147</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>H. White.</surname>
          </string-name>
          “
          <article-title>The Value of Narrativity in the Representation of Reality”</article-title>
          .
          <source>In: Critical Inquiry 7.1</source>
          (
          <issue>1980</issue>
          ), pp.
          <fpage>5</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zeman</surname>
          </string-name>
          . “
          <article-title>Grammatik der Narration”</article-title>
          .
          <source>In: Zeitschrift für germanistische Linguistik 48.3</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>457</fpage>
          -
          <lpage>494</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>