<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Faces, Fights, and Families: Topic Modeling and Gendered Themes in Two Corpora of Swedish Prose Fiction</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Uppsala University</institution>
          ,
          <addr-line>Uppsala</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <fpage>92</fpage>
      <lpage>111</lpage>
      <abstract>
        <p>This paper explores topic modeling (TM) as a tool for “distant reading” of two Swedish literary corpora. We investigate what kinds of insight and knowledge a TM-based approach can provide to Swedish literary history, and which methodological dificulties are associated with this endeavour. The TM is based on 12- and 24-term chunks of selected verb and common noun lemmas. We generate models with 20, 40, and 100 topics. We also propose a method for a quantitative and qualitative gendered thematic analysis by combining TM with a study of how the topics relate to gender in characters and authors. The two corpora contain, respectively, Swedish classics (1821-1941) and recent bestsellers (2004-2017). We find that most of the topics proposed by the TM are easy to interpret as conceptual themes, and that the “same” themes appear for the two corpora and for diferent TM settings. The study allows us to make interesting observations concerning diferent aspects of gender and topic distribution.</p>
      </abstract>
      <kwd-group>
        <kwd>Topic Modeling</kwd>
        <kwd>Distant Reading</kwd>
        <kwd>Gender Analysis</kwd>
        <kwd>Literary Methodology</kwd>
        <kwd>Swedish Prose Fiction</kwd>
        <kwd>Bestsellers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The aim of this paper is to explore topic modeling (TM) as a tool for “distant
reading” of Swedish literary corpora. We want to investigate what new kinds of
insight and knowledge a TM-based approach can help us gain as regards Swedish
literary history. We also want to discuss some of the methodological dificulties
associated with this endeavour. In particular, this article proposes an approach
to quantitative and qualitative gendered thematic analysis by combining TM
with a study of how the topics relate to gender in characters and authors. This
has, as far as we know, not been tried before.</p>
      <p>Our aims are exploratory and mostly focused on methodical investigation,
but results will also be reported and discussed. The study is concerned with
two corpora: prose fiction with modern(ized) Swedish spelling from
Litteraturbanken (mainly Swedish classics, 1821–1941); and prose fiction from
contemporary Swedish bestseller charts (2004–2017).</p>
      <p>Our research questions can be stated as follows:
– How well does TM work as a tool for extracting content themes from Swedish
literary corpora?
– How robust and reproducible are the results of a TM system?
– Is it possible to find connections between topics and the gender of characters
and authors?
– What are the advantages of using TM as a tool for the analysis of (Swedish)
literature in comparison with other methods?
2</p>
    </sec>
    <sec id="sec-2">
      <title>Literature and Topic Modeling – State of the Art</title>
      <p>In traditional studies on themes in literature, the researcher approaches the texts
having some predefined theme as his or her point of departure. This choice might
be motivated by e.g. its historical, stylistic/aesthetic, or political significance.
The established methodology in finding themes to investigate is to rely on already
read books, on books that could or should be relevant, or on previous research.
In recent times, free-text search engines have come to be more and more used
for locating instances of themes in literary corpora. Still, such procedures rely
on the researcher’s assumptions about which terms are indicative of the relevant
themes.</p>
      <p>
        As several scholars of literature have already pointed out, topic modeling
(TM) makes it possible to approach thematic literary analysis from another
angle. Instead of first deciding on a theme and a material and then search for
passages expressing the theme, the researcher selects a collection of texts and
makes use of an algorithm to find which “topics” – as components of a latent
probabilistic model – best explain the structure of the texts, e.g. [
        <xref ref-type="bibr" rid="ref10 ref33">10, 33</xref>
        ]. Such a
bottom-up approach to thematic literary analysis will generate proposals which
are quantitatively justified by the data in a way that is not possible in the
traditional frameworks of literary studies.
      </p>
      <p>
        The present work belongs to a paradigm of computational quantitative
largescale analysis of literature, which Franco Moretti [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] wittily has characterized
as distant reading. There are a few exploratory examples of this kind of
literary criticism on Swedish literary material, e.g. [
        <xref ref-type="bibr" rid="ref12 ref4 ref6">12, 6, 4</xref>
        ], but only one uses TM
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The shift from manual qualitative narrow-scale methods to computational
distant reading has been criticized by researchers in the humanities for being
reductive, positivist, white male-centred, and not critical enough, e.g. [
        <xref ref-type="bibr" rid="ref1 ref15 ref18 ref20">1, 15, 20,
18</xref>
        ]. Objections have also been raised in a Swedish context, e.g. by [
        <xref ref-type="bibr" rid="ref16 ref5">5, 16</xref>
        ], despite
the fact that there are only few examples of this kind of research on Swedish
material.
      </p>
      <p>TM and similar algorithmic methods thus appear to be both productive
and provocative to literary scholars. Although we believe in the usefulness of
distant reading approaches, we will discuss our results with an awareness of the
methodological problems.</p>
      <sec id="sec-2-1">
        <title>Generative Model</title>
        <p>
          A topic model (TM, but notice the diference between model and modeling here)
is a model of a collection of text pieces, which can be segmented according to
diferent criteria. In our approach, the text segments (called chunks) are
(typically) smaller than paragraphs and only contain a selection of lexical terms. A
TM of the kind used here is a probabilistic model of how the text chunks are
generated by a hypothetical stochastic process. The idea is that we can view
each chunk as a sequence of terms, which is generated in such a way that we
ifrst draw a topic according to the topic distribution of the chunk, and then each
term in the chunk given the term probabilities of the topic [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ].
chunk data
topic model
        </p>
        <sec id="sec-2-1-1">
          <title>Corpus</title>
          <p>Chunking
TM (Mallet)
“Presentation”</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Report</title>
          <p>The modeling process consists in finding a model that as well as possible fits
the data, i.e. the chunks, which reflect the surface facts of the corpus. The last
step, in an approach like the present one, is to turn the statistical model into
something which can be read and interpreted. We follow common practice in
generating lists of the most significant words associated with each topic. The
pipeline is shown in Figure 1.</p>
          <p>In applications of TM, we typically want the topics to capture content
themes. Of course, the concept of theme is a vague and open one. Themes can be
more or less specific, as well as partially overlapping. Saying which theme a topic
represents (based on the corresponding list of significant words) is consequently
a matter of qualitative analysis. The noun topic is often used as a synonym to
theme in everyday language, but we will only use topic in the technical sense
of TM here. So, topics correspond to themes, if the TM is successful, but they
can also fail to do so. Furthermore, topics can not be expected to be
conceptually “pure”, as minor subthemes can be associated with topics even if a main
thematic label is clearly justified.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>TM and Literature Research</title>
        <p>
          Several literary scholars have recognized the relevance of TM for “distant
reading”. As can be expected, this research has often turned to English literature,
such as poetry [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], 19th century prose fiction [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ], or contemporary
American bestsellers [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Jockers [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] makes use of TM to perform narratological
analysis. Among other influential entries we find work on the French Encyclopédie
(18th century) [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], French classical drama (17th–18th century) [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ], Spanish
Golden-Age sonnets [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], Danish 19th century literature [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ], and a meta-study
of American research on literature [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>
          The only previous study using TM on Swedish literature is Barakat’s [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
master thesis (in statistics and data mining). He tries to map statistically what
makes audiobooks popular. Although the study is based on a large corpus (3077
books) and poses interesting questions, its focus is on statistical matters rather
than literary ones. There are also a few studies using TM on non-literary Swedish
material, such as governmental oficial reports [
          <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
          ] and Swedish parliamentary
debates and the discourse on immigration [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Method and Data</title>
      <p>
        The concept of topic is a statistical one. A topic captures a “probability
distribution over terms”, intended to model what can be understood as
“recurring themes” in a collection of text chunks [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Each chunk is associated with a
probability distribution over the topics. TM is a process of estimating the two
distributions. A TM system is guided by a number of parameters and the user’s
decisions on data and parameter settings are crucial for the result it will produce.
      </p>
      <p>
        Our approach involves three steps. First, we propose a procedure, of our own
invention, for selecting the terms and capturing the chunks to be fed to the TM
module. Secondly, there is the TM per se, which relies on a TM module of the
Mallet [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] software package. Finally, the results of the TM are presented to the
user for qualitative interpretation. The discussion is based on our own reading
of the results.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Gender-related Labeling</title>
        <p>The paragraphs from which the text chunks are generated are labeled by two
kinds of gender-related information. The first kind indicates whether a paragraph
is only about female or only about male characters. We also label paragraphs
according to the gender of the author(s). When there are more than one author,
and they belong to diferent sexes, we do not use either label.</p>
        <p>We operationalize the character gender distinction by labeling paragraphs
containing at least one singular third person feminine pronoun (hon [subject
form], henne [object form], hennes [genitive]), without containing a singular
third person masculine pronoun (han, honom, hans). This gives us a simple
way of identifying passages which are about female characters. We also label the
paragraphs involving a singular third person masculine pronoun (but not any
feminine one) in the corresponding way. The two categories are thus by
definition mutually exclusive. However, many paragraphs will not belong to either of
the two categories, by not containing any such pronoun or by involving both
genders. Also, note that pronouns are not among the lexical terms preserved in
the chunking (see next section), so they will not be “seen” by the TM.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Chunking and Selection of Lexical Terms</title>
        <p>The first step in the processing of the texts is part-of-speech tagging. 1 Lexical
terms are then formed by combining the base form and the part-of-speech tag.
This means that all inflected forms are grouped as one term (lemma) and that
the part-of-speech tag disambiguates some lemmas, e.g. röra, giving us röraNN
(jumble) or röraVB (touch). Only a subset of the lexical terms is used for the
TM. First, we remove all instances of the 100 most frequent ones. (They form a
list of stop words, as it were.) Furthermore, we require, for the Classics corpus,
that the terms have at least 10 instances distributed over at least 10 diferent
books. The corresponding requirement for the Bestsellers corpus is 40 instances
over 20 diferent books, which is proportional to the relevant sizes. Finally, we
only include nouns and verbs among the terms, assuming that these carry a high
“semantic weight”, and are the most useful indicators of recurring themes.
supaVB veckaNN skötaVB tjänstNN församlingNN måsteVB klagaVB
prostNN biskopNN biskopNN sockenNN hållaVB
tjänstNN församlingNN måsteVB klagaVB prostNN biskopNN biskopNN
sockenNN hållaVB korNN bröstNN prästNN
visshetNN fiendeNN kyrkaNN fiendeNN bänkNN bondeNN kyrkaNN korNN
fiendeNN fiendeNN fiendeNN trampaVB</p>
        <p>
          After the tagging and the compilation of a frequency dictionary for the
corpus, each text is processed paragraph by paragraph. These are converted to lists
of lexical terms, t1; t2; : : : ; tp (when the paragraph yields p terms). From these
we generate the chunks which will form the data fed to the TM module. The
chunks are of a predefined length, c 2 f12; 24g, in our experiments. In
comparison with e.g. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], our chunks are very small, intended to focus on themes
which are prominent at that level of textual “resolution”. The “window” of the
“chunker” moves forward three terms2 for each capture. So, t1; t2; : : : ; tc will be
the first chunk, t4; t5; : : : ; tc+3, the second one, t7; t5; : : : ; tc+6, the third one, and
so on, until t3n+1; t3n+2; : : : ; t3n+c, for the largest n such that 3n + c p. (So,
this procedure gives us n + 1 chunks for a paragraph term sequence of length p.)
See Table 1 for an example. We do not allow chunks to extend over paragraphs,
which we assume are often associated with shifts in thematic content. We ignore
sentence boundaries. Also note that the chunks overlap and that the number of
term instances in the chunk set will be considerably larger than in the actual
text.
1 We use Stagger [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ] A few obvious frequent tagging errors are corrected.
2 Three rather than one for “economic” reasons.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Topic modeling</title>
        <p>
          The TM is done by means of the ParallelTopicModel class of the Mallet [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]
software. The class is characterized (in a code comment) as providing a “[s]imple
parallel threaded implementation of LDA [Latent Dirichlet Allocation], following
[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], with SparseLDA sampling scheme and data structure from [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]”. The output
is strongly influenced by the setting of a number of parameters. We generated
a fairly small number of topics, 20, 40, or 100. Our settings are consequently
geared towards the extraction of general themes (see [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]). The Mallet settings
were based on tuning on the Classics corpus and manual inspection of the results.
k 2 f20; 40; 100g [a.k.a numTopics]
alphaSum = 2:0 k
beta = 0:001
symmetricAlpha = false
numIterations = 2000
burninPeriod = 200
optimizeInterval = 100
        </p>
        <p>
          The high alphaSum = 2:0 k is motivated by a wish to avoid “favoring just
a few topics” [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ], whereas beta, which “smoothes the word distribution in every
topic” [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] is given a low value (0:001). Since the set-ups stipulate that number
of topics is fairly small, these will be of a general nature. The chunking will
furthermore promote the extraction of themes that manifest themselves locally
in texts. The small beta will promote models where terms tend to be specific to
a few topics.
3.4
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Presentation of the TM Output</title>
        <p>The result of the TM is essentially an assignment of a topic (identified by an
integer) to each instance of a term in the chunks. This means that we can
estimate, for each term (type) and topic, the probability that the term represents
the topic, p(topicjterm), for the complete corpus. Similarly, the result defines,
for each chunk and topic, a ratio saying to what degree the chunk represents the
topic, p(topicjchunk). As the chunks derive from literary works, and are labelled
with gender-related information, we can also compute how large a share a topic
has in a particular book, and in male and female authors, and in passages with
male and female characters.</p>
        <p>
          As is common practice, we present the topics for “reading” as lists of “top”
terms. We rank the terms according to the 2 (chi square) statistics,3 which
quantifies the strength of association between a term and a topic [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. We think
that this is better than looking at the more “elementary” conditional
probabilities: p(termjtopic) is high for frequent, but consequently general words. And
p(topicjterm) will be close to 100% for many very rare – and consequently
3 It is computed in this way: 2 =
∑
∑
(NF T EF T )2 , where NF T is the
        </p>
        <p>EF T</p>
        <p>F 2ff;fg T 2ft;tg
actual number of observations and EF T the expected number of instances under the
assumption that T and F are independent. Here, 2 is used to generate word lists
from the TM output. (We do not use it for significance testing.)
not very significant – words. We do however use colour and boldface to
indicate p(topicjterm), since this probability tells us how specific a term is for a
topic/theme. The styles are to be read as follows: TermPOS 90%, otherwise
TermPOS 75%, otherwise TermPOS 50%, otherwise (plainly) TermPOS
(&lt; 50%), see e.g. Table 3.
3.5</p>
      </sec>
      <sec id="sec-3-5">
        <title>Data: the “Classics” and the “Bestsellers” Corpora</title>
        <p>
          We have applied our method for TM and gender analysis on two quite
diferent corpora of Swedish prose fiction, “Classics” and “Bestsellers”. The “Classics”
corpus is curated from Litteraturbanken (LB) (www.litteraturbanken.se), which
is a collection of Swedish literature mainly from the 19th century and the first
decades of the 20th century. The focus of LB is on classics and literature of
particular historical or aesthetical importance. It also contains translations and
reference literature [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. LB comprises more than 700 e-texts, as well as facsimile
editions. The Classics corpus is a subset of the e-texts, which covers the prose
ifction with modern(ized) (post-1906) spelling, as spelling variation would
interfere with the TM. With duplicate editions removed, the corpus consists of
121 Swedish novels or collections of short stories (6.6 million words), originally
published between 1821–1941, the majority after 1900. Male writers are
overrepresented: Only 36% of the works are by female authors. (The full list of works
is available as an appendix to the TM presentations among the Supplementary
Materials.)
        </p>
        <p>Since LB is centred around some influential Swedish authors and their works,
the corpus is not representative of the full range of literary production in Sweden
during the period in question. Authors like Hjalmar Bergman (14 works), Selma
Lagerlöf (16), and August Strindberg (23) are over-represented, while others are
not found at all. Nevertheless, the LB corpus allows us to study the recurring
thematic content in some of the most prominent Swedish writers of the time.
The Classics corpus was compiled from the system-internal XML-files of LB.</p>
        <p>
          The second corpus studied in this article, “Bestsellers”, is based on the
Swedish bestseller charts compiled by the book trade magazine Svensk Bokhandel
[
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]. We collected all prose fiction on the lists 2004–2017 in the two categories
“bestselling hardbound fiction” and “bestselling paperbacks”. We then excluded
all non-fiction, all fiction that does not contain prose (just a few works of
poetry), and all duplicates. (Duplicate entries are common since bestsellers tend
to sell well both in hardbound and paperback editions.) This gives us 280
bestselling novels and collections of short stories, of which 231 works (82; 5%) were
available in digital form. The raw text from these, some 25:7 million words, thus
constitutes the corpus.4 This corpus covers more than four fifths of all bestselling
prose fiction published in Swedish during the first decades of the 21st century.
4 The extraction and processing of the text were conducted within the confines of
the activities of the university library. Everything except the literary text proper
has been removed, including meta-data, dedications, forewords, afterwords,
acknowledgements, extra material, etc.
The 49 titles that were not available in digital form seem to be quite evenly
distributed over time and genres. However, Swedish crime fiction is presumably
somewhat over-represented in the corpus and many titles are missing from 2017.
        </p>
        <p>The bestseller corpus is dominated by novels (there is only one collection
of short stories), Swedish originals (75%), and crime fiction ( 62%). However, it
also includes translated fiction, genre fiction other than crime (romance, science
ifction, fantasy), as well as literary and more aesthetically experimental fiction.
In contrast to the Classics corpus, the gender distribution in the Bestseller corpus
is balanced: 48% of the works are written by female authors.</p>
        <p>Classics Corpus Bestsellers Corpus
Author WC FP MP Chunks RF RM WC FP MP Chunks RF RM
Female 2.4 25.3 26.8 51508 12082 14917 11.1 25.5 24.0 233733 62755 59886
Male 4.2 10.2 29.6 148517 13757 74045 13.5 15.7 27.4 253228 40166 110616
“Both” - - - - - - 1.1 26.2 25.9 10625 2496 3613
Sum 6.6 - - 200025 25839 88962 25.6 - - 497586 105417 174115</p>
        <p>If we look at the sizes of the data converted into 12-word chunks, we get the
numbers in Table 2. We immediately see that the female authors write about
male and female characters in a balanced way, whereas the male authors write
about male individuals three to four times more often than about women.</p>
        <p>Considering their sizes and principled composition, the two corpora provide
a solid point of departure for a study on thematic trends in Swedish literature
(including recent translated bestsellers). They also allow us to compare, from a
thematic point of view, literature from the beginning of the 20th century with
literature from the beginning of the 21st century.
3.6</p>
      </sec>
      <sec id="sec-3-6">
        <title>Experiments</title>
        <p>The result of a TM experiment is crucially dependent on the user’s design choices,
such as chunk size, number of topics, and assumptions about distribution priors.
Our design and parameter tuning were based on experiments with the Classics
corpus, 12- and 24-term chunks, and k 2 f20; 40g. We then applied the settings
that seemed useful to the Bestsellers corpus and to experiments with k = 100.
At URL https://stp.lingfil.uu.se/~matsd/dhn2019, we provide as
Supplementary Materials for this artice, the automatically generated reports from our
TM experiments, with listings of the works included in the corpora.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussion</title>
      <p>We would like to claim that the topics proposed by our experiments to a high
degree are possible to interpret in a “natural” way as expressions of one main
general theme. However, many topics seem to contain traces of subordinate themes.
Consider, for instance, topic #7 from the Bestsellers corpus when k = 20 and
c = 24, as exhibited in Table 3. It seems (also when we look at the longer
wordlist) that fighting and violence is the central theme. But we also find nouns
for categories of animal. This can be explained by the narratives in the corpus,
but the topic is probably also associated with chunks featuring peaceful animals
and fights without animals. Because of this thematic polysemy and vagueness, it
is often not obvious which terms to use in labeling the topics. We have mixed two
“strategies”: In some cases we use a general term that covers the central part of
the topic. For other topics, we prefer a conjunction of more specific terms. (The
reader is invited to look at the Supplementary Material to get a fuller picture.)</p>
      <p>In the experimental set-ups with k = 20, for both corpora, 15 intuitively
identified themes on diferent levels of specificity appeared as topics in all cases.
So at k = 20, three out of four themes were the “same” for c 2 f12; 24g. We also
tried our scheme on the 2010–2017 subset of the Bestsellers corpus: With k = 20
and c = 24, again roughly three out of four topics seemed to capture the “same”
ones as those extracted from the full Bestsellers corpus. Our experiments also
allowed us to find remarkable thematic similarities between the two corpora, see
Section 4.2. These results suggest that the current TM approach has a certain
degree of stability.</p>
      <p>Changes in the number of topics obviously lead to changes in the output.
In many cases we see that more general themes in a conceptually reasonable
way split into more specific ones when we increase the number of topics from
k = 20 to k = 40, and then further at k = 100. For instance, a persistent theme
surfacing for the Bestsellers corpus is one corresponding to action-loaded fighting
scenes. With k = 20 this theme is captured by one topic (see Table 3), which
seem to derive from many kinds of fighting scene, mostly from crime fiction, but
also from other genres (e.g. historical fiction, which is a plausible explanation for
the high ranking of horse). The theme is quite clearly split in two at k = 40, one
related to blunt force and another to firearm violence (see Table 3). And when
we set k = 100, this general theme yields four diferent topics (see Table 3). Here,
topic #99 is clearly related to firearms ( c = 24, henceforth). Topic #55 seems
to relate to more serious violence, involving e.g. knives or subsequent injuries.
Topics #19 and #78 both seem to capture fist fights, but at slightly diferent
stages: Topic #19 is related to the middle stages of fights, whereas topic #78
corresponds to their consequences.</p>
      <p>These results show that the TM has the ability to discover similarities
between diferent kinds of fight when the number of topics is small, but that it
is also able to separate diferent kinds or aspects of fights from one another as
the number of topics grows larger. There are other examples. Topics inviting
the labels “Outdoor settings” and “Eating and drinking” from the Bestsellers
corpus behave in a similar fashion: Content related to outdoor settings is
gathk = 20
k = 40
#7 skjutaVB, hästNN, skrikaVB, knivNN, markNN, vapenNN, skottNN,
fallaVB, springaVB, kulaNN, djurNN, pistolNN, sparkaVB, pilNN, slåVB,
hundNN, fotNN, gevärNN, slitaVB, skrikNN. [shoot, horse, shout, knife, ground,
weapon, shot, fall, run, bullet, animal, pistol, kick, arrow, hit, dog, foot, gun, tear, cry.]
#8 skrikaVB, slåVB, sparkaVB, tagNN, fallaVB, slitaVB, armNN, tappaVB,
kastaVB, huvudNN, fotNN, benNN, balansNN, gripaVB, snubblaVB, näveNN,
sparkNN, vrålaVB, kraftNN, skrikNN. [shout, hit, kick, hold, fall, tear, arm, drop,
throw, head, foot, leg, poise, seize, stumble, fist, kick, roar, force, shout.]
#22 skjutaVB, springaVB, stegNN, sekundNN, skottNN, vapenNN,
pistolNN, kulaNN, vändaVB, revolverNN, stannaVB, gevärNN, fotNN,
meterNN, ögonblickNN, siktaVB, rörelseNN, närmaVB, hinnaVB, rusaVB. [shoot,
run, step, second, shot, weapon, pistol, bullet, turn, revolver, stop, rifle, foot, meter,
moment, aim, movement, approach, find time, rush.]</p>
      <p>k = 100
#19 fallaVB. kastaVB, slåVB, golvNN, markNN, ramlaVB, knäNN, landaVB,
kantNN, krossaVB, brytaVB, vältaVB, hoppaVB, rullaVB, stötaVB, studsaVB,
bakhuvudNN, vikaVB, träfaVB, slungaVB. [fall, throw, hit, floor, ground, tumble,
knee, land, edge, crush, break, overthrow, jump, roll, push, bounce, back of the head,
fold, hit, hurl.]
#55 tagNN, smärtaNN, knivNN, gripaVB, halsNN, skäraVB, tandNN,
armNN, käkeNN, knäckaVB, muskelNN, huggaVB, ormNN, bitaVB, ryggNN,
sparkaVB, snaraNN, sprutaNN, ryggradNN, greppNN. [grip, pain, knife, grab,
throat, cut, tooth, arm, jaw, break, muscle, chop, snake, bite, back, kick, noose, syringe,
spine, grip.]
#78 fotNN, benNN, balansNN, tåNN, tappaVB, trampaVB, krypaVB,
släppaVB, vacklaVB, greppNN, rockNN, resaVB, vristNN, fläckNN,
spindelNN, högerhandNN, svajaVB, stampaVB, haltaVB, hasaVB. [foot, leg,
balance, toe, loose, tread, crawl, drop, falter, grip, coat, travel/rise, ankle, spot, spider,
right hand, swing, stomp, limp, shamble.]
#99 skjutaVB, vapenNN, skottNN, pistolNN, kulaNN, gevärNN,
revolverNN, siktaVB, magasinNN, patronNN, kornNN, avlossaVB,
ammunitionNN, riktaVB, avtryckareNN, avfyraVB, kaliberNN, laddaVB,
containerNN, bågeNN, kolvNN. [shoot, weapon, shot, pistol, rifle, revolver, aim,
magazine, cartridge, sight (firearm), trigger, ammunition, aim, trigger, trigger, calibre,
load, container, bow, cylinder.]
ered in one topic #6 when k = 20 (again c = 24), split in three when k = 40,
one capturing settings adjacent to water #20, one houses and gardens #28, and
one woods and fields #37. It yields four topics when k = 100. We also find a
clear “Eating and drinking” topic #14, when k = 20. This is split into two when
k = 40 and into three when k = 100: cofee breaks #48, setting the table #83,
and drinking #87.</p>
      <p>Outcomes like these indicate that selecting a certain number of topics does
not necessary make the result more or less conceptually “correct”, but rather
guides the TM process to capture themes on diferent levels of specificity.
Generally, using TM with fewer topics will capture broader thematic categories (such
as fights or depictions of landscapes), whereas a larger number of topics will
allow the TM to capture more specific categories (such as shootings or depictions
of woods).
4.1</p>
      <sec id="sec-4-1">
        <title>Topics, Themes and Gender</title>
        <p>
          There have been previous attempts to connect themes retrieved by TM to author
gender [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. There are also large-scale studies of character gender in literary
text based on pronoun evidence [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. However, no one seems to have combined
these modes of analysis. Our experimental design allows us to find gender-biased
themes both as regards characters and authors.
        </p>
        <p>The corpus data and the TM output make it possible, for each topic t, to
compute its relative frequency in character-based female chunks rt;F , and the
corresponding value, rt;M , for male chunks. The ratio between the two values,
i.e. rt;F =rt;M , is a measure of how much more common the topic is in female
compared to male contexts. We can use this number to rank the topics on a
scale from highly female to highly male. (We should remember, that masculine
pronouns and male character paragraphs are much more common than their
female counterparts in both corpora, notably in the Classics – see Table 2.)</p>
        <p>For the topics retrieved from the Classics corpus with k = 20 and c = 24, the
most female-biased ones can be labeled “Dance, music, entertainment”, “Family
relations” and “Mental states, existential reflection” (see Figure 2). The three
leading topics on the male end can be put under the headings “Authority”
(church and nation), “God, religion, faith”, and “Money, work, trade” (buying,
selling, and monetary matters).</p>
        <p>If we carry out the same exercise for the Bestsellers corpus (again with k = 20
and c = 24), we end up with the picture in Figure 3. The most female topic
collects terms relating to dress, hair, and make-up, which we label “Appearance”.
The following two invite the labels “Family circle. Disease, health care”, as there
appears to be a strong connection between family members and hospitals, and
“Eating and drinking”, for which the top terms agree with that label in an
obvious way. The topics most strongly associated with male characters concern
“War, nations, global politics”, “Police work, investigations, law enforcement”
and “Fighting, violence. Animals” (discussed above). We also find topics with a
gender-independent distribution in both corpora, for instance, “Telephone calls”
#10 in the Bestsellers corpus or “Disease, health care. Accidents” #5 in the
Classics corpus.</p>
        <p>
          ) 10
%
(
y
c
n
eu 5
q
e
r
f
.
l
e
R 0
[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] Dance, [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] Family [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] Mental [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] Money, [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] God,
music relations states work religion
[
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]
Authority
#18 spelaVB, flickaNN, dansaVB, sjungaVB, dansNN, fruNN, mammaNN,
musikNN. [play, girl, dance, sing, dance, wife/Mrs, mother, music.]
#10 barnNN, morNN, farNN, sonNN, hemNN, förälderNN, årNN. [child,
mother, father, son, home, parent, year.]
#2 tankeNN, drömNN, själNN, livNN, känslaNN, minneNN, verklighetNN.
[thought, dream, soul, life, emotion, memory, reality.]
#8 pengNN, köpaVB, betalaVB, säljaVB, arbeteNN, skafaVB, kronaNN.
[money, buy, pay, sell, work, get, crown.]
#11 gudNN, människaNN, gärningNN, helveteNN, andeNN, djävulNN,
nådNN. [God, man, deed, hell, spirit, devil, grace.]
#17 kyrkaNN, prästNN, konungNN, kungNN, församlingNN, folkNN,
svenskNN. [church, priest, king, parish, people, Swede.]
Fig. 2. Most gender-biased topics in the Classics corpus (k = 20 and c = 24). Red
staples: rt;F . Blue staples: rt;M . Grey staples: relative frequency of the topic in the full
corpus. Topics numbered according to their relative frequency. The top seven term of
the topics are listed.
        </p>
        <p>The patterns we find seem to reflect traditional gender roles and stereotypes
to a high degree: To deal with authority, God, and trade was primarily
responsibilities for men around the turn of the century 1900, whereas entertainment,
the family circle, and reflection on mental states were more closely connected to
female concerns. In contemporary bestsellers, men dominate (or are expected to
dominate) in “arenas” like crime, in particular violent crime, the police force,
politics, and the military. Women, on the other hand, are to a higher degree
depicted in contexts where appearance, family life, and meals play a prominent
role. Hence, gender roles have changed – yet they seem to remain the same in
many respects.</p>
        <p>
          In general, the gender bias of characters and authors point in the same
direction. This is true for the examples we have given above. As can be expected, men
[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] War,
nations
#17 hårNN, klänningNN, skjortaNN, skoNN, kjolNN, byxaNN, jeansNN.
[hair, dress, shirt, shoe, skirt, trousers, jeans.]
#19 pappaNN, mammaNN, svaraVB, frågaVB, mormorNN,
sjuksköterskaNN, morfarNN. [dad, mum, answer, question, maternal
grandmother, nurse, maternal grandfather.]
#14 kafeNN, drickaVB, ätaVB, matNN, vinNN, koppNN, glasNN. [cofee,
drink, eat, food, wine, cup, glass.]
#7 skjutaVB, hästNN, skrikaVB, knivNN, markNN, vapenNN, skottNN.
[shoot, horse, shout, knife, ground, weapon, shot.]
#15 polisNN, mordNN, utredningNN, mördareNN, oferNN, brottNN,
vittneNN. [police, murder, investigation, murderer, victim, crime, witness.]
#5 krigNN, landNN, mrNN, presidentNN, distriktNN, arbetsgivareNN,
nådNN. [war, country, mr, president, district, employer, grace.]
        </p>
        <p>Fig. 3. Most gender-biased topics in the Bestseller corpus. Read as Figure 2.
[C19] Eating
and drinking
[BS14] Eating
and drinking</p>
        <p>[C13]
Reading and writing
[BS18]
Reading and writing
#C19 drickaVB, kafeNN, matNN, ätaVB, vinNN, middagNN, glasNN. [drink,
cofee, food, eat, wine, dinner, glass. ]
#BS14 kafeNN, drickaVB, ätaVB, matNN, vinNN, koppNN, glasNN. [cofee,
drink, eat, food, wine, cup, glass.]
#C13 läsaVB, bokNN, skrivaVB, tidningNN, professorNN, författareNN,
brevNN. [read, book, write, newspaper, professor, author, letter.]
#BS18 skrivaVB, bokNN, läsaVB, brevNN, bildNN, textNN, papperNN.
[write, book, read, letter, picture, text, paper.]
Fig. 4. Themes were gender biased has changed. C for Classics and BS for Bestsellers;
otherwise as Figure 2.
generally tend to write more about male-biased themes and women more about
female-biased themes. This also agrees what is known from previous research.
There are however a few exceptions to this pattern. The theme “Agriculture,
animals” #16 in the Classics corpus and “Money and trade” #20 in the Bestsellers
corpus are clearly male-biased in the texts, but have higher or equal
representation among female authors. The opposite pattern can be noted in “Dance,
music, entertainment” #18 in the Classics corpus and “Family circle. Disease,
health care” #19 in the Bestsellers corpus, where the topics are more prevalent
in female character paragraphs while being more frequent in male authors.</p>
        <p>To interpret these instances of thematic “gender-crossing” would require a
more thorough analysis than is possible here. And we should be cautious about
what we read into these results, as our corpora are small and do not represent
balanced selections of text. For instance, some patterns might be due to
individual authors or works. However, these observations support our conclusion
that the present TM approach leads to sound results: the gender biases relating
to characters and authors correspond to each other, but they do not align
entirely. Investigations into such discrepancies can provide insight into how literary
themes connect to gender in diferent ways.</p>
        <p>Another thing we can see is that gender bias is slightly more pronounced in
the Classics corpus, both as regards characters and authors. Such comparisons
can also be made for specific themes. For instance, “Reading and writing” has
moved from a male-dominated arena in the Classics corpus to a gender-neutral</p>
        <p>Topic
one among the contemporary bestsellers. “Eating and drinking” has in a similar
fashion moved from a male-dominated arena in the Classics corpus to a
femaledominated one in the Bestsellers corpus (see Figure 4). These results are easy
to relate to well-known societal changes that have taken place in the time span
between the two corpora.</p>
        <p>Although the gendered patterns discussed above are indeed evident, we should
not overestimate their magnitude. In the most clearly biased themes (found when
k = 100), the diference is a matter of a factor of 2, both as regards characters
and authors. But most topics do not exhibit a pronounced gender bias. The
gendered patterns are overall less distinct than we expected, which is in itself
an interesting result and something that should be addressed more carefully in
the future. We also expected that an increase in k would generate more
genderbiased topics (since larger k-values generate more specific themes in the TM),
but such a pattern could not be observed.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Thematic Comparison between the Two Corpora</title>
        <p>If we look at the topics from the two corpora when the TM has been asked to
deliver broad thematic categories (k = 20, c = 24), we find a striking number
of thematically related topics. We think that a qualitative analysis which shows
that most topics are thematically related to a topic in the other corpus (see Table
4) is justified. So, for instance, eating and drinking, reading and writing, money
and trade, and existential reflection and mental states, are four themes which
Classics #8 pengNN, köpaVB, betalaVB, säljaVB, arbeteNN, skafaVB,
kronaNN, arbetaVB, arbetareNN, kostaVB, prisNN, lönNN. [money, buy, pay,
sell, work, get, crown, work, worker, price, salary.]
Bestsellers #20 pengNN, betalaVB, säljaVB, köpaVB, kronaNN, kundNN,
företagNN, kostaVB, tjänaVB, prisNN, bankNN, miljonNN. [money, pay, sell,
buy, crown, customer, cost, earn, price, bank, million.]
Classics #19 drickaVB, kafeNN, matNN, ätaVB, vinNN, middagNN, glasNN,
brännvinNN, ölNN, brödNN, bordNN, dryckNN. [drink, cofee, food, eat, wine,
dinner, glass, liquor, beer, bread, table, beverage.]
Bestsellers #14 kafeNN, drickaVB, ätaVB, matNN, vinNN, koppNN, glasNN,
brödNN, bordNN, teNN, kokaVB, flaskaNN. [ cofee, drink, eat, food, wine, cup,
glass, bread, table, tea, boil, bottle.]
in a clear way surface as topics for both corpora. Table 5 shows the remarkable
overlap among the top terms for the topics “Money and trade” and “Eating and
drinking”. So, themes like these seem to be persistent in Swedish literature, even
if the corresponding topics also reflect the historical context.</p>
        <p>The topics which seem to capture themes which only appear for one corpus
do so in a conceptually straightforward way. Themes like telephone calls, driving,
and police investigations are retrieved only for the contemporary Bestsellers. For
the Classics corpus, by contrast, some topics capture themes which were more
important for the economy and society of the time period it represents, such as
agriculture, maritime activity, titles, and religion.</p>
        <p>Comparisons of this kind lead to interesting observations, which would be
possible to follow up in a deeper qualitative analysis. In this way, the use of TM
for the exploration of literary corpora suggests new ways for the literary scholar
to track themes, changes in themes, and the decline of themes over extended
periods, throughout literary history.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper, we have proposed an experimental procedure whose central
component is TM to extract themes from Swedish corpora of prose fiction. The topics
and the term lists derived from them are the products of algorithms whose
output is based on statistical processing of text segments in parallel. Which topics
are established depends on patterns of co-occurrence among words. These can
be due to both conceptual and factual associations of a more general nature, as
well as to more idiosyncratic circumstances, e.g. patterns deriving from specific
authors.</p>
      <p>Our analysis of the results is tentative and sketchy because of the size of
the material and the limited scope of the study. We still wish to conclude that
the experiments are successful in consistently producing easy-to-interpret and
conceptually plausible topics. We think that our use of 2 (chi square) to rank
the terms in the topic word lists facilitate the qualitative interpretation, as it lifts
the “core” terms, and – so to speak – highlights the central thematic content.</p>
      <p>The method works with both small k values, for the retrieval of broad
thematic categories, and with larger k values, in which case we get more specific
themes. This makes the method flexible and possible to use for diferent research
questions. We have not systematically investigated diferent parameter settings
for the various modules (other than compared choices for k 2 f20; 40; 100g). This
would be a challenging task, because of the number of parameters involved, and
the lack of precise criteria for assessing the output. The chunking principles are
also in need of further evaluation: For instance, there is a need to compare the
consequences of using chunks captured by a window-based procedure with a
design based on paragraphs or larger units as chunks.</p>
      <p>We would like to highlight three areas in which “distant reading” based on
TM can be of value to scholars of literature:</p>
      <p>First, the use of TM allows us to investigate themes in a bottom-up fashion,
where the computer-generated output is the point of departure for a qualitative
analysis. This approach stands in contrast to methods driven by the researcher’s
preexisting ideas about the nature of given themes. In this way, it prompts us to
read familiar literary material in new ways, by giving us condensed pictures of
the topics/themes which are quantitatively justified by the textual surface facts.
In particular, we hope that we have shown that the use of TM allows us to make
plenty of interesting observations about Swedish literature.</p>
      <p>A second advantage is that TM with labeled chunks supports analysing
topics and their dependence on other properties. In the present study, we performed
a gendered thematic analysis which connects thematic content to biases
concerning the gender of characters and authors. Our use of pronoun evidence to connect
paragraphs to character gender is admittedly a far from perfect
operationalization. It could be improved by the use syntactic parsing and retrieval of more
ifne-grained labeling of literary characters. It is easy to imagine studies in the
same vein where the focus would lie on other categories, such as race/ethnicity,
class, age, etc.</p>
      <p>A third interesting kind of analysis supported by TM is the comparison of
diferent corpora. The collections which we have discussed here, classics from
the decades around the turn 1900 and recent titles from the bestseller charts,
would hardly invite thematic comparison in traditional literary studies.
However, we were able to align the topics derived from the two collections and find
remarkable similarities, as well as interesting diferences. A related application
would be to track how themes develop and recede through time. In this article,
this comparison of the topic sets has been a matter of qualitative analysis. A
prospect for future work is to develop computational mechanism for aligning
topics.</p>
      <p>
        There are also other ways in which TM can serve literary research, which are
beyond the scope of this paper. One is to use TM as a component in a search tool,
which assists the scholar in finding occurrences of a given theme in a corpus (see
e.g. [
        <xref ref-type="bibr" rid="ref17 ref32">32, 17</xref>
        ]). Another application to use TM for basic narratological analysis as
regards the linear distribution of thematic content in the course of novels, as has
been done by Jockers [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Performed on a larger scale and combined with broad
thematic categories, such analyses can help us find recurring plot structures and
allow us to see how narratives work. This is yet another promising idea for future
studies.
      </p>
      <p>It is important to remember that the TM output always has to be interpreted
qualitatively by the literary scholar in the kind of research we propose. This
interpretation is far from a trivial task and requires an understanding of the
corpus. As we see it, TM does not provide a method for delivering definite
reports about the set of themes inherent in a literary corpus. Rather, the TM
output is in need of literary interpretation of a traditional kind to yield new
insights and knowledge.</p>
      <p>In her book Reductive Reading (2018), Sarah Allison develops a closely related
point. She argues that one of most important lessons literary historians can learn
from computational criticism is that it forces the researcher to lay bare what is
taken into account in the analysis and what is not. Such methods are in this way
more solid and more transparent than traditional methods for literary research.
Allison writes:</p>
      <p>
        Reductive reading, by contrast, clears space for reading that is not
reductive. […] My argument for reductive reading is also a defense of descriptive
– call them “weak” – findings: a very strong opening claim can shelter
more nuanced claims than a project that open with a more textured or
qualified polemic. ([
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]: 2–3)
      </p>
      <p>This is similar to how we understand the possibilities of TM for scholars
of literature: If we take the caveats of the method seriously and discuss them
transparently, TM output can provide a solid quantitative backdrop to further
qualitative literary analyses. We consequently reject the idea – that sometimes
is loudly expressed in debates – that quantitative and qualitative methods stand
in opposition to each other.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>We would like to thank Litteraturbanken, and its Head, Professor Mats Malm,
for making data available to us. We are also grateful to the Disciplinary Domain
of Humanities and Social Sciences at Uppsala University for funding.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Allington</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          , Brouillette,
          <string-name>
            <surname>S</surname>
          </string-name>
          , Golumbia,
          <string-name>
            <surname>D</surname>
          </string-name>
          (
          <year>2016</year>
          ), “
          <article-title>Neoliberal Tools (and Archives): A Political History of Digital Humanities</article-title>
          ,” Los Angeles Review of Books, May
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Allison</surname>
            ,
            <given-names>S</given-names>
          </string-name>
          (
          <year>2018</year>
          ), Reductive Reading: A Syntax of Victorian Moralizing, Baltimore: Johns Hopkins University Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Archer</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          , Jockers,
          <string-name>
            <surname>M</surname>
          </string-name>
          (
          <year>2016</year>
          ),
          <article-title>The Bestseller Code: Anatomy of the Blockbuster Novel</article-title>
          , New York: St. Martin's Press.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Barakat</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          (
          <year>2018</year>
          ), “
          <article-title>What Makes an (Audio)Book Popular?,” master thesis in statistics and machine learning</article-title>
          , Department of Computer and Information Science, Linköping University.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bergenmar</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          (
          <year>2017</year>
          ), “
          <article-title>Har feminism och vithetskritik en plats inom digital humaniora?,” in Erixon, P-O, Pennlert</article-title>
          , J (ed.)
          <article-title>Digital humaniora - humaniora i en digital tid</article-title>
          ,
          <source>Göteborg: Daidalos</source>
          , pp.
          <fpage>77</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Berglund</surname>
            ,
            <given-names>K</given-names>
          </string-name>
          (
          <year>2017</year>
          ), “Killer Plotting.
          <article-title>Typologisk intriganalys utifrån fjärrläsningar av 113 samtida svenska kriminalromaner,” Tidskrift för litteraturvetenskap, (3</article-title>
          <issue>-4</issue>
          ), pp.
          <fpage>41</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Blatt</surname>
            ,
            <given-names>B</given-names>
          </string-name>
          (
          <year>2017</year>
          ),
          <article-title>Nabokov's Favourite Word is Mauve: The Literary Quirks and Oddities of Our Most-Loved Authors</article-title>
          , London &amp; New York: Simon &amp; Schuster.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Blei</surname>
          </string-name>
          , DM (
          <year>2012</year>
          )
          <article-title>“Topic Modeling</article-title>
          and Digital Humanities,
          <source>” Journal of Digital Humanities</source>
          , Vol.
          <volume>3</volume>
          , No.
          <volume>2</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Goldstone</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          , Underwood,
          <string-name>
            <surname>T</surname>
          </string-name>
          (
          <year>2014</year>
          ), “
          <article-title>The Quiet Transformation of Literary Studies: What Thirteen Thousand Scholars Could Tell Us,” New Literary History</article-title>
          , vol.
          <volume>45</volume>
          , pp.
          <fpage>359</fpage>
          -
          <lpage>384</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jockers</surname>
          </string-name>
          ,
          <string-name>
            <surname>Matthew</surname>
          </string-name>
          (
          <year>2013</year>
          )
          <article-title>Macroanalysis: Digital Methods</article-title>
          and
          <string-name>
            <given-names>Literary</given-names>
            <surname>History</surname>
          </string-name>
          , Urbana: University of Illinois Press.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Jockers</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mimno</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          (
          <year>2013</year>
          ), “
          <article-title>Significant Themes in 19th-</article-title>
          <string-name>
            <surname>Century</surname>
            <given-names>Literature</given-names>
          </string-name>
          ,” Poetics, vol.
          <volume>41</volume>
          (
          <issue>6</issue>
          ), pp.
          <fpage>750</fpage>
          -
          <lpage>769</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kokkinakis</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          , Malm,
          <string-name>
            <surname>M</surname>
          </string-name>
          (
          <year>2015</year>
          ), “
          <article-title>Detecting Reuse of Biblical Quotes in Swedish 19th Century Fiction using Sequence Alignment,” Corpus-based Research in the Humanities workshop</article-title>
          (CRH), pp.
          <fpage>79</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>Limin</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L</given-names>
            ,
            <surname>Mimno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            ,
            <surname>McCallum</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          (
          <year>2009</year>
          ) “
          <article-title>Eficient methods for topic model inference on streaming document collections.” In Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining (KDD '09)</article-title>
          . ACM, New York, NY, USA,
          <fpage>937</fpage>
          -
          <lpage>946</lpage>
          . DOI: https://doi.org/10.1145/1557019.1557121.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Litteraturbanken</surname>
          </string-name>
          , “The Swedish Literature Bank,” URL: https://litteraturbanken.se/om/english.html.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          (
          <year>2012</year>
          ), “
          <article-title>Where is Cultural Criticism in the Digital Humanities?,” in Debates in the Digital Humanities, Matthew</article-title>
          K. Gold (ed.), Minneapolis: University of Minnesota Press, pp.
          <fpage>490</fpage>
          -
          <lpage>509</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lundblad</surname>
            ,
            <given-names>K</given-names>
          </string-name>
          (
          <year>2017</year>
          ), “
          <article-title>Digital humaniora - en pleonasm i den digitala kulturen,” in Erixon, P-O, Pennlert</article-title>
          , J (eds.)
          <article-title>Digital humaniora - humaniora i en digital tid</article-title>
          ,
          <source>Göteborg: Daidalos</source>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Magnusson</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          , Öhrvall,
          <string-name>
            <surname>R</surname>
          </string-name>
          , Barrling,
          <string-name>
            <given-names>K</given-names>
            ,
            <surname>Mimno</surname>
          </string-name>
          <string-name>
            <surname>D</surname>
          </string-name>
          (
          <year>2018</year>
          ),
          <article-title>“Voices From the Far Right: A Text Analysis of Swedish Parliamentary Debates</article-title>
          ,” working paper,
          <source>SocArXiv Papers</source>
          ,
          <year>April 2018</year>
          , URL: https://osf.io/preprints/socarxiv/jdsqc/.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Mandell</surname>
            ,
            <given-names>L</given-names>
          </string-name>
          (
          <year>2015</year>
          ), “Gendering Digital Literary History,” in Schreibman, S, Siemens,
          <string-name>
            <surname>R</surname>
          </string-name>
          , Unsworth, J (ed.)
          <article-title>A New Companion to the Digital Humanities</article-title>
          , Chichester: Wiley, pp.
          <fpage>511</fpage>
          -
          <lpage>523</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>CD</given-names>
          </string-name>
          , Raghavan,
          <string-name>
            <surname>P</surname>
          </string-name>
          , Schütze,
          <string-name>
            <surname>H</surname>
          </string-name>
          (
          <year>2008</year>
          ) Introduction to Information Retrieval, Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>McPherson</surname>
            ,
            <given-names>T</given-names>
          </string-name>
          (
          <year>2012</year>
          ), “
          <article-title>Why Are the Digital Humanities So White? or Thinking the Histories of Race and Computation,” in Debates in the Digital Humanities</article-title>
          , Gold, MK (ed.), Minneapolis: University of Minnesota Press, pp.
          <fpage>139</fpage>
          -
          <lpage>160</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>AK</given-names>
          </string-name>
          (
          <year>2002</year>
          )
          <article-title>MALLET: A Machine Learning for Language Toolkit</article-title>
          . http://mallet.cs.umass.edu.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Moretti</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          (
          <year>2000</year>
          ), “Conjectures on World Literature,” New Left Review, vol.
          <volume>40</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>54</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Navarro-Colorado</surname>
            ,
            <given-names>B</given-names>
          </string-name>
          (
          <year>2018</year>
          ), “
          <article-title>On Poetic Topic Modeling: Extracting Themes and Motifs From a Corpus of Spanish Poetry,” Frontiers in Digital Humanities</article-title>
          , vol.
          <volume>5</volume>
          (
          <year>June 2018</year>
          ), URL: https://doi.org/10.3389/fdigh.
          <year>2018</year>
          .
          <volume>00015</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Newman</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asuncion</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          , Smyth,
          <string-name>
            <surname>P</surname>
          </string-name>
          , Welling,
          <string-name>
            <surname>M</surname>
          </string-name>
          (
          <year>2009</year>
          )
          <article-title>“Distributed Algorithms for Topic Models</article-title>
          .
          <source>” Journal of Machine Learning Research</source>
          <volume>10</volume>
          (
          <year>December 2009</year>
          ), pp.
          <fpage>1801</fpage>
          -
          <lpage>1828</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Norén</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          (
          <year>2016</year>
          ), “
          <article-title>Information som lösning, information som problem: En digital läsning av tusentals statliga utredningar</article-title>
          ,
          <source>” Nordicom Information</source>
          , vol.
          <volume>38</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>9</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Norén</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          , Snickars,
          <string-name>
            <surname>P</surname>
          </string-name>
          (
          <year>2017</year>
          ), “
          <article-title>Distant reading the history of Swedish film politics in 4500 governmental SOU reports</article-title>
          ,
          <source>” Journal of Scandinavian Cinema</source>
          , vol.
          <volume>7</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>155</fpage>
          -
          <lpage>175</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Rhody</surname>
          </string-name>
          , LM (
          <year>2012</year>
          ), “
          <article-title>Topic Modeling and Figurative Language”</article-title>
          ,
          <source>Journal of Digital Humanities</source>
          , vol.
          <volume>2</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>19</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Roe</surname>
            ,
            <given-names>G</given-names>
          </string-name>
          , Gladstone,
          <string-name>
            <surname>C</surname>
          </string-name>
          , Morrissey,
          <string-name>
            <surname>R</surname>
          </string-name>
          (
          <year>2016</year>
          ),
          <article-title>“Discourses and Disciplines in the Enlightenment: Topic Modeling the French Encyclopédie,” Frontiers in Digital Humanities</article-title>
          , vol.
          <volume>3</volume>
          (
          <year>January 2016</year>
          ), article 8, URL: https://doi.org/10.3389/fdigh.
          <year>2015</year>
          .
          <volume>00008</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Schöch</surname>
            ,
            <given-names>C</given-names>
          </string-name>
          (
          <year>2017</year>
          ), “
          <article-title>Topic Modeling Genre: An Exploration of French Classical and Enlightenment Drama,” Digital Humanities Quarterly</article-title>
          , vol.
          <volume>11</volume>
          (
          <issue>2</issue>
          ), URL: http://www.digitalhumanities.org/dhq/vol/11/2/000291/000291.html.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Steyvers</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          , Grifiths,
          <string-name>
            <surname>T</surname>
          </string-name>
          (
          <year>2007</year>
          ) “
          <article-title>Probabilistic topic models.” In Landauer, TK, McNamara</article-title>
          ,
          <string-name>
            <surname>DS</surname>
          </string-name>
          , Dennis,
          <string-name>
            <surname>S</surname>
          </string-name>
          , Kintsch, W (Eds.),
          <article-title>Handbook of latent semantic analysis</article-title>
          , Mahwah, NJ, US: Lawrence Erlbaum Associates Publishers, pp.
          <fpage>427</fpage>
          -
          <lpage>448</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Svensk</surname>
            <given-names>Bokhandel</given-names>
          </string-name>
          , ”Top lists”, URL: http://www.svb.se/toplists.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Tangherlini</surname>
            ,
            <given-names>T</given-names>
          </string-name>
          , Leonard,
          <string-name>
            <surname>P</surname>
          </string-name>
          (
          <year>2013</year>
          ), “
          <article-title>Trawling in the Sea of the Great Unread: Sub-corpus Topic Modeling</article-title>
          and Humanities Research,” Poetics, vol.
          <volume>41</volume>
          (
          <issue>6</issue>
          ), pp.
          <fpage>725</fpage>
          -
          <lpage>749</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Underwood</surname>
            ,
            <given-names>T</given-names>
          </string-name>
          (
          <year>2014</year>
          ), “Theorizing Research Practices We Forgot to Theorize Twenty Years Ago,” Representations, vol.
          <volume>127</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>64</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Östling</surname>
            ,
            <given-names>R</given-names>
          </string-name>
          (
          <year>2013</year>
          ) “
          <article-title>Stagger: an Open-Source Part of Speech Tagger for Swedish</article-title>
          .
          <source>” Northern European Journal of Language Technology</source>
          , Vol.
          <volume>3</volume>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>