<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>What Do We Talk About When We Talk About Topic?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Joris J. vanZundert</string-name>
          <email>joris.van.zundert@huygens.knaw.n</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MarijnKoolen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>JuliaNeugarten</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Boot</string-name>
          <email>peter.boot@huygens.knaw.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Willem vanHage</string-name>
          <email>w.vanhage@esciencecenter.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ole Mussmann</string-name>
          <email>marijn.koolen@gmail.com</email>
          <email>o.mussmann@esciencecenter.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DHLab, KNAW Humanities Cluster</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>KNAW Huygens Institute</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Radboud University Nijmegen</institution>
          ,
          <addr-line>Nijmegen</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>eScience Center</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <fpage>398</fpage>
      <lpage>410</lpage>
      <abstract>
        <p>We apply Top2Vec to a corpus of 10,921 novels in the Dutch language. For the purposes of our research we want to understand if our topic model may serve as a proxy for genre. We 昀椀nd that topics are extremely narrowly related to an existing genre classi昀椀cation historically created by publishers. Interestingly we also 昀椀nd that, notwithstanding careful vocabulary 昀椀ltering as suggested by prior research, various other signals, such as author signal, stubbornly remain.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;literary 昀椀ction</kwd>
        <kwd>novels</kwd>
        <kwd>computational literary studies</kwd>
        <kwd>topic models</kwd>
        <kwd>top2vec</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>genres, like horror, romance, crime, etc. Our analysis invites re昀氀ection on the usefulness of
topic modelling as a tool for computational literary studies (CLS).</p>
      <p>Because of copyright, we cannot share the texts of the novels on which the topic models
are based. However, we can and will eventually share the intermediate results and our code
through the GitHub repository of the proje1ct.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Topic modelling and literature</title>
      <p>
        Various technologies have been used in CLS and other 昀椀elds to establish the semantic value
of topics. In CLS, for example, researchers have used frequent word20s][, (some variety of)
keywords [
        <xref ref-type="bibr" rid="ref15">3, 19</xref>
        ], top-down methods such as the UCREL Semantic Analysis System1[
        <xref ref-type="bibr" rid="ref9">8</xref>
        ], and
systematic manual procedures1[].
      </p>
      <p>
        A commonly applied technique is topic modelling, which algorithmically identi昀椀es groups
of words that tend to co-occur in a large collection of documen10ts].[Although introduced in
a non-昀椀ction context, this has been applied to 昀椀ction (e.g. [
        <xref ref-type="bibr" rid="ref18 ref21 ref7 ref8">6, 14, 22, 15, 12, 7, 11, 25</xref>
        ]). Topic
modelling has also been applied to the discourse of literary stud9i]easn[d to online literature
reviews [28].
      </p>
      <p>Some previous research in CLS suggests that topic modelling literary works leads to
semantically cohesive topics and an overarching understanding of the subject matter represented in
those works [2]. In Macroanalysis [13], Jockers used “Themes” as the title for his chapter about
topic modelling. This title suggests that topic modelling is a technique for unearthing “the”
theme of a novel, viz. “a salient abstract idea that emerges from a literary work’s treatment of
its subject-matter” (the 昀椀rst meaning of “theme” de昀椀ned in [ 5]). Schröter and Du 2[3] suggest
that “sujet” is similar to literary topic. Lund1y5[] uses topic modelling on a corpus of ca. 1,000
recent popular U.S. novels. He creates one set of general topics, analyzing entire novels at a
time. Additionally, he creates a set of more speci昀椀c topics using individual sentences from
these books. The focus of the study is on how these topics are distributed over genres. Jautze
et al. [12] topic model 400 recent Dutch works of 昀椀ction, and use the resulting topics to predict
a reader response variable (perceived literariness). Others (e1.g1.])[argue that topic models
do not generate semantically coherent descriptions of topics meaningful to CLS.</p>
      <p>
        Of particular interest here are22[], where it is argued that “strong genre signals exist […] on
the levels of function words, content words and syntactic structure” but that they “also exist
on the level of theme or topic”;2[5] arguing that topics relate to structural rather than content
elements of text; and [
        <xref ref-type="bibr" rid="ref20">24</xref>
        ] where it is shown that topics strongly correlate to meta-textual
features such as author and genre.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Problem</title>
      <p>
        There is no hard scienti昀椀c consensus on what textual features constitute genre, and as, for
instance, [
        <xref ref-type="bibr" rid="ref24">29</xref>
        ] argues there may be good reason to question canonical genre classi昀椀cations.
Topics from topic models on the other hand are notoriously hard to clearly relate to literary
constructs or categories2[
        <xref ref-type="bibr" rid="ref20 ref21">2, 25, 24</xref>
        ]. If we are interested in the relation between quanti昀椀able
textual features and readers’ preferences, our question becomes how bottom up topics from a
topic model relate to given genre metadata.
      </p>
      <p>
        We operationalise genre categories using Dutch NUR-codes (Nederlandse Uniforme
Rubrieksindeling or Dutch Uniform Categories classi昀椀cation). These were introduced in 2002 as a market
monitoring instrument and succeed the comparable NUGI codes that have been in use for the
same reason since 1987. NUR is a practical marketing instrument devised and applied by
publishers. It is largely ignored in Dutch literary culture, and it goes largely unnoticed by readers
who mostly encounter it because bookshops tend to sort and arrange their product range
according to the system [
        <xref ref-type="bibr" rid="ref22 ref23">26, 27</xref>
        ]. NUR can be regarded as a rough approximation of the concept
of genre as it is understood by booksellers and readers.
      </p>
      <p>The examination of the correlation between genre (NUR) and topic can be broken down
into several sub-questions. How is a topic related to speci昀椀c NUR codes? What is the topic
distribution of a NUR code? What is the NUR distribution of the books most associated with a
topic? As a 昀椀rst step, we determine how strongly topics are associated with NUR codes. Ideally,
such associations are unrelated to topics that emerge from our model as the artefacts of corpus
features like author signal, translation, and corpus composition as these are irrelevant with
regard to canonical genre categories.</p>
      <p>We make the results of our topic modelling concrete and useful to CLS in two ways: by
examining the correlations between the topics detected in our corpus and one possible
operationalisation of genre, and by re昀氀ecting on the usefulness of topic modelling for literary
analysis in light of our results.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Method</title>
      <p>4.1. Data
Courtesy of an agreement with the Dutch national library and seven Dutch publishing houses
(representing multiple publishers) we have access to the full text of 10,921 Dutch-language
novels, published between 2009 and 2019 in the Netherlands. Tab1lelists a number of general
statistics for the corpus used in this research.</p>
      <p>Element
Novels</p>
      <p>5000
Paragraphs
Sentences
Words</p>
      <p>Number</p>
      <p>Min</p>
      <p>Max</p>
      <p>The collection is based on the EPUBs deposited by publishers at the national library.
Therefore, it is a subset of all books published in the Netherlands in this period. The collection is
skewed towards more recent books (see Figur2e), primarily because more books are now being
made available as ebooks, not because of an increase in publications. The corpus consists of
both Dutch novels and at least 2,199 translated novels, mainly from English, German, French
and various Scandinavian languages, as well as smaller representations of languages such as
Spanish, Italian, and Japanese.</p>
      <p>Figure 1 shows the distribution of book lengths in number of words per book. There is
one sharp peak around 50,000 words and a lower, less pronounced peak around 90,000 words,
roughly coinciding with the conventional lengths of novellas (~80-120 pages) and “full” novels
(~300-500 pages). A small number of books are very short. Of the 10,921 novels, 216 novels (2%)
are shorter than 1,000 words, and 463 novels (4%) are between 1,000 and 10,000 words. Some of
these may be picture books, children’s books, or collections of poetry. Some unusually short
books are regular-sized novels for which the text extraction step did not work properly. These
are mostly EPUB “incunables”. Apparently publishers needed to get used to the EPUB format
and so early EPUBs o昀琀en have poor 昀椀le and content structure.</p>
      <sec id="sec-4-1">
        <title>4.2. Preprocessing</title>
        <p>
          Based on existing research1[
          <xref ref-type="bibr" rid="ref21">2, 25</xref>
          ], we pre-process the novel texts to remove person names and
use lemmas rather than full word forms. We tokenise and parse all novels using SpaC2y 3.3
and remove all word tokens that are part of person entities. SpaCy inserts underscores in
lemmas for certain compound words, but sometimes fails to split the words correctly. E.g. for the
Dutch word “boekhandel” (en:book shop), the SpaCy lemma is almost always “boek_handel”,
but sometimes SpaCy assigns the lemma “boekh_andel”. Therefore, we post-process the
lemmas by removing underscores. Based on a sample of the 1,000 most common lemmas
containing underscores and the variants containing underscore in a di昀erent position, we found that
removing the underscores rarely con昀氀ates words with di昀erent meanings.
        </p>
        <p>We remove common words based on their document frequency, as such words tend to
appear in the majority of topics and thus have no discriminating e昀ect between topics. We also
remove lemmas that occur in very few books. These tend to include many speci昀椀c names
(many character names are not recognised by SpaCy as person named entities), and very rare
vocabulary used by only one or a few authors. This means that we remove lemmas that occur
in fewer than 1% or more then 10% of books (fewer than 103 books or more than 1,030 books).
This leaves a vocabulary of 36,927 lemmas.</p>
        <p>This setup represents a trade-o昀 between corpus coverage and the ability to generate
di昀erentiating topics. For future evaluation we aim to vary this bandwidth to gauge the stability of
2See https://spacy.io
the topic models inferred (see als6o).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. Segmentation</title>
        <p>Next, we need to choose a unit of measure. Based on prior research We test two di昀erent unit
sizes: the whole novel as a document, and documents constructed from joining a sequence of
paragraphs in a novel into segments containing at least 5,000 words. Using whole documents
yields rather few topics (95) to relate to 731 existing NUR codes, while the latter choice results
in more and smaller documents, and 1,182 topics. Further details and analysis of these di昀erent
unit are discussed in AppendixA.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.4. Topic modelling</title>
        <p>
          Given the number and size of documents (or segments) in a corpus, it is di昀케cult to decide the
minimum number of relevant topics.1[5] used LDA3 on a set of 1,136 novels and chose 60 as
the optimal number of topics based on the fraction of topics they could meaningfully interpret.
[12] used LDA on lemmatised 1,000-word segments of 401 Dutch novels and set the number
of topics to 50. [
          <xref ref-type="bibr" rid="ref21">25</xref>
          ] segmented novels into 300 to 500 word segments, and used a pointwise
mutual information (PMI) and a cosine based coherence measure to observe that “as a rule of
thumb [...] the number of topics should lie between 100 and 1502”5[, p.66].
        </p>
        <p>
          In this research we applied Top2Vec4[], which in recent studies has compared favorably
to LDA and other techniques 2[1], [
          <xref ref-type="bibr" rid="ref9">8</xref>
          ]. Models generated using LDA or PLS4Agenerate a
pre-determined number topics as distributions of words. This o昀琀en means that topically
uninformative words have high probabilities in the topics, since they make up a large proportion
of all text. In Top2Vec, joint document and word embeddings are learned in a 300-dimensional
space, which is projected onto a low-dimensional space using UMAP17[], a昀琀er which
HDBSCAN [
          <xref ref-type="bibr" rid="ref12">16</xref>
          ] is used to detect dense document clusters, which determine the number of topics.
This ensures that the words nearest a topic vector best describe the topic and its surrounding
documents [4, p.3].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In this 昀椀rst analysis we use the full corpus, lemmatise tokens, drop persons names (PER
according to SpaCy), and cull lemmas appearing in less than 1% or over 10% of all novels. Top2Vec
was used in its ’fastlearn’ mode with 8 workers and a standard multilingual universal sentence
encoder5. When using full novels as unit of measure, this resulted in 95 topics. For 5,000 token
windows, Top2Vec generated 1,182 topics.</p>
      <p>Labeling the topics that result from topic modelling is a complicated and subjective process
which we believe requires annotation and an assessment of inter-annotator agreement. Such
3Latent Dirichlet allocation, a common approach to topic modelling,hsteteps://en.wikipedia.org/wiki/Latent_Dir
ichlet_allocation
4See https://en.wikipedia.org/wiki/Probabilistic_latent_semantic_analysis
5https://tfhub.dev/google/universal-sentence-encoder-multilingual/3
assessment falls outside the scope of the current paper. For this reason, we have chosen to
number topics here, rather than give them a semantically meaningful label.</p>
      <p>Each book has a nearest topic neighbour, which Top2Vec uses to determine the size of a
topic. This way, each topic has a size, expressed as the number of books for which that topic
is the nearest neighbour. This allows us to investigate how strongly topics are associated with
the NUR codes: we look at the distribution of NUR codes as the percentage of a topic’s most
associated books that are assigned that NUR code. For instance, the 昀椀rst topic is the nearest
topic for 1,655 of the novels, and of these 1,643 (99%) are labelled with NUR code 3R4o3mance.
This topic is therefore strongly associated with a single NUR code. We investigate the
association of topics and NUR code using a heatmap (see 昀椀gure 4). The darker a cell, the stronger
the topic (x-axis) is associated with a genre (y-axis). Many topics strongly associate with one
or two NUR codes. One observation we derive from this is that topic as inferred by Top2Vec
is strongly associated with publishers’ choices for NUR genre code. This observation can be
further corroborated if we use UMAP to create a visualisation of topic clusters where we color
each vector according to the NUR it most closely associates with (see 昀椀gur6e).</p>
      <p>As we have seen in the previous section, the number of topics increases if the unit of measure
decreases. If we topic model at the level of whole novels we 昀椀nd only 95 topics. With 5,000
word segments, the resulting model has 1,182 topics, and the association between topic and
NUR code is even stronger (see Figure5). Both the size of the segments (roughly comparable
to the size of chapters) and the number of topics found, coincide more closely with intuitions
about how literary topics may function. That is, from a literary criticism point of view we
would expect topics to be bound more narrowly to chapters or paragraphs rather than to a
book as a whole, while 95 topics would seem a tiny set of topics to cover a corpus of over
10,000 novels.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>What becomes clear from the results as depicted in Figur4easnd 5, is that topics generated by
Top2Vec are extremely narrowly related to genre, with many topics almost exclusively related
to one genre. Furthermore, topics that are strongly associated with the genres “Dutch
literary novel” and “Translated literary novel” turn out to contain a large number of geographical
indicators (cf. appendixB).</p>
      <p>
        Our results con昀椀rm 昀椀ndings from [
        <xref ref-type="bibr" rid="ref18 ref20 ref21">22, 25, 24</xref>
        ]. In all this means that topics as generated
by Top2Vec across our corpus will be an adequate proxy for genre in the course of our project.
Thus our current result can be summarised as “when we talk about topic modelling we actually
talk about genre”.
      </p>
      <p>
        Like the results of 2[
        <xref ref-type="bibr" rid="ref20 ref21">2, 25, 24</xref>
        ], our results give pause to consider that topics generated
through topic modelling techniques are much more related to signals of genre than to semantic
椀昀elds that literary researchers would consider topical and relevant. Similarly we consider that
although geography can be topical for a novel, geography related signals seem much stronger
than their relevance for literary analysis would warrant. Most salient is the observation that,
even though we followed2[5] in carefully removing function words and author speci昀椀c
vocabulary, we still 昀椀nd that topics strongly coincide with author if we recolor 昀椀gu6raeccording to
author (see 昀椀gure 7). On the one hand this may mean that authors keep within genre, on the
other it means it is still hard to decide what topic model topics relay to us.
      </p>
      <p>So far, we have only looked at one topic modelling technique (Top2Vec) and two segment/document
sizes. Additionally, our 昀椀ndings are limited because we used only Dutch NUR coding as a genre
target. For now, we have also disregarded the skewed makeup of our corpus, in which the
romance genre is severely over-represented (cf. 昀椀gure3). We still need to evaluate the e昀ects
of using a di昀erent corpus balancing, di昀erent document sizes, isolating subgenres, di昀erent
topic modelling techniques such as classic LDA, and di昀erent genre labels.</p>
      <p>In our current corpus, topic turns out to be strongly associated with genre as labeled by
Dutch publishers. Our next step will be to determine the distribution of topics across di昀erent
NUR genres. A昀琀er that we aim to gauge how features of reader reviews relate to the topics we
found.</p>
      <p>A. Goldstone and T. UnderwoodW.hat Can Topic Models of PMLA Teach Us About the
History of Literary Scholarship? 2012. url:
https://tedunderwood.com/2012/12/14/whatcan-topic-models-of-pmla-teach-us-about-the-history-of-literary-scholars.hip/
[11] R. Heuser and L. Le-Khac.A Quantitative Literary History of 2,958 Nineteenth-Century</p>
      <p>British Novels: The Semantic Cohort Method, Literary Lab Pamphlet 4. 2018.
[12] K. Jautze, A. van Cranenburgh, C. Koolen, et al. “Topic Modeling Literary Quality”. In:
Dh. 2016, pp. 233–237.</p>
      <p>M. L. Jockers.Macroanalysis: Digital Methods and Literary History. University of Illinois
Press, 2013.</p>
      <p>M. Lundy. “Text Mining Contemporary Popular Fiction: Natural Language
ProcessingDerived Themes Across Over 1,000 New York Times Bestsellers and Genre Fiction
Novels”. PhD thesis. University of South Carolina, 2020.</p>
    </sec>
    <sec id="sec-7">
      <title>A. Unit of Measure</title>
      <sec id="sec-7-1">
        <title>A.1. Segmentation</title>
        <p>From the perspective of literary studies, it is illogical to bind topic to the full text of a novel. A
novel likely touches on a multitude of topics, so a division into chapters, sections, paragraphs
or even sentences might yield more useful topics. However, segmenting whole novels into
minimum-sized windows has consequences for the co-occurrence of words and the number of
topics that will be detected. By segmenting a whole novel, the words within one segment no
longer co-occur with the words in another segment from the same novel. Therefore, the word
co-occurrence matrix becomes more sparse, so topic modelling algorithms identify more and
smaller dense clusters, resulting in more topics.</p>
        <p>We investigate the impact of segmenting on the co-occurrence of words by analysing how
the number of co-occurring pairs of lemmas increases as we iterate over novels, using either
the whole novels as document boundaries, or segments constructed from joining a sequence of
paragraphs in a novel into segments containing at least 5,000 words. Figu8rsehows the result.
The X-axis shows the number of lemmas (a昀琀er 昀椀ltering out the most and least frequent lemmas
as described above) and the Y-axis shows how many distinct co-occurring pairs of lemmas are
found.</p>
        <p>The di昀erence between segmenting and not segmenting results generates a clear picture. At
the level of whole novels, well over 196 million co-occurrence pairs are indexed a昀琀er having
seen 300,000 lemmas (corresponding to about 250 novels). For the segmented novels, at the
same 300,000 lemma point, there are only 18.3 million pairs. That is a full order or magnitude
less. Regarding a choice for unit of measure, this means that we have to carefully investigate
the trade-o昀 that exists between the size of documents fed to a topic modelling algorithm (i.e.
whole novels or smaller segments) and the number of topics returned.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>B. Topic examples</title>
      <p>The following are examples from topics, generated by Top2Vec at the level of whole novels,
that show a high concentration of coherent geographical lemmas. Lemmas strongly related to
a coherent geographical location are in bold.</p>
      <p>Topic number: 4
Topic number: 5
nochtans, komaan, stilaan, gsm, parking, job, vooraleer, miserkieo,t, brussels, antwerps, euh,
proper, gsmnummerg,oesting, bijgevolg, verdict, plezanta, ntwerpen, vlaams, voormiddag,
speurder, voordienz,aventem, oostende, 昀氀ik, ontgoocheling, zogezegde, deurgat,gent,
autosnelweg, voorhand, pvg,ents, crapuul, evolueren, zijkt,nokke, vlaming,
sukkelaar,mechelen, recupereren, meermaals, contacteren, evident, nonkel, allez, klasseren, rijkswaschhte,lde
vondelpark, schiphol, leidseplein, bitterbal, shag, know, like, lullen, please, gadverdamme,
it, amsterdams, snot, never, grachtenpand, can, kroket, amsterdam, goor, spuug, there,
veegt, kut, ie, see, bh, amstelveen, 昀椀etspad, plee, geilheid, wc, gezeik, rouwkaart, quote,
almere, sure, tje, fucking, is, ehm,hilversum, only,randstad, drop, lacherig,koninginnedag,
too, zeiken, opschuden, poep
Topic number: 11
zo, polder, hbs, jeneverh,aarlem, zoldering, stationsplein, ballpoint, tramhalstceh,evenings,
verveloos,scheveningen, vondelpark, shag, rotterdam, waartussen,
grammofoon,zandvoort, rijksdaalder, arnhem, hongerwinter, jongensboeki,jsselmeer, gymnasium, schemer,
schoolschri昀琀, ijl, schemren, brokkelig, sto昀樀as,amstel, klomup, vitrage, bakeliet, sigarenwinkel,
schrijfmachine, vergelen, bovenhuis, rui, plantsoen, brillegglarso,ningen, celluloid, windstil,
trapper,leidseplein, vooroorlogs, veraf, allengws,assenaar
Topic number: 15
stockholm, zweden, kronen, kopenhagen, oslo, noorwegen, denemarken, deens, noors,
midzomer,昀椀ns , zweed, vooronderzoek, line, verhoren, ordner, politieacademie, lichtkegel,
legitimatie, volvo,昀椀nland , strafregister, rechercheteam昀樀 o,rd , kelderruimte, nor, wide,
politiemen, hoofdbureau, moordzaken, geweldsdelict, villawsicjka,ndinavisch, messing, goddomme,
avondkrant, rijkweg, smeriss, opsporing, bewakingscamersac,andinavie, freule, pagekapsel,
nova, zomerhuis,oostzee, sankt, thomas, moordonderzoek, sterfgeval</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>M. J. Adler</surname>
          </string-name>
          .
          <article-title>The Great Ideas: A Syntopicon of Great Books of the Western World</article-title>
          . Vol.
          <volume>2</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Encyclopaedia</given-names>
            <surname>Britannica</surname>
          </string-name>
          ,
          <year>1952</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Algee-Hewitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Heuser</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>MorettiS</surname>
          </string-name>
          .tanford Literary Lab Pamphlet 10:
          <string-name>
            <given-names>On</given-names>
            <surname>Paragraphs. Scale</surname>
          </string-name>
          , Themes, and
          <string-name>
            <given-names>Narrative</given-names>
            <surname>Form</surname>
          </string-name>
          .
          <year>2015</year>
          . url: https://litlab.stanford.edu/Literar yLabPamphlet10.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Allington</surname>
          </string-name>
          . “
          <article-title>Customer Reviews of 'Highbrow' Literature: A Comparative Reception Study of The Inheritance of Loss and The White Tiger”</article-title>
          .
          <source>InA:merican Journal of Cultural Sociology 9.2</source>
          (
          <issue>2021</issue>
          ), pp.
          <fpage>242</fpage>
          -
          <lpage>268</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Angelov</surname>
          </string-name>
          . “
          <article-title>Top2Vec: Distributed Representations of Topics”</article-title>
          . Ianr:Xiv preprint arXiv:
          <year>2008</year>
          .
          <volume>09470</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [10] [13] [14] [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Baldrick</surname>
          </string-name>
          .
          <source>The Oxford Dictionary of Literary Terms [online]</source>
          . Oxford University Press,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Bode</surname>
          </string-name>
          . “
          <article-title>”Man people woman life” - “Creek sheep cattle horses”: In昀氀uence, Distinction, and Literary Traditions”</article-title>
          .
          <source>InA: World of Fiction: Digital Collections and the Future of Literary History</source>
          . University of Michigan Press,
          <year>2019</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>197</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Buurma</surname>
          </string-name>
          . “
          <article-title>The Fictionality of Topic Modeling: Machine Reading Anthony Trollope's Barsetshire Series”</article-title>
          .
          <source>InB:ig Data &amp; Society 2.2</source>
          (
          <issue>2015</issue>
          ), p.
          <fpage>2053951715610591</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Egger</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          . “
          <string-name>
            <given-names>A Topic</given-names>
            <surname>Modeling Comparison Between</surname>
          </string-name>
          <string-name>
            <surname>LDA</surname>
          </string-name>
          , NMF, Top2Vec, and BERTopic to Demystify Twitter Posts.”
          <source>In:Frontiers in sociology 7</source>
          (
          <year>2022</year>
          ), p.
          <fpage>886498</fpage>
          . doi:
          <volume>10</volume>
          .3389/fsoc.
          <year>2022</year>
          .
          <volume>886498</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Goldstone</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Underwood</surname>
          </string-name>
          . “
          <article-title>The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us”</article-title>
          .
          <source>NIne:w Literary History 45.3</source>
          (
          <issue>2014</issue>
          ), pp.
          <fpage>359</fpage>
          -
          <lpage>384</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>M. L. Jockers</surname>
            and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Mimno</surname>
          </string-name>
          . “
          <article-title>Signi昀椀cant Themes in 19th-Century Literature”</article-title>
          .
          <source>InP:oetics 41.6</source>
          (
          <issue>2013</issue>
          ), pp.
          <fpage>750</fpage>
          -
          <lpage>769</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>McInnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Healy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Astels</surname>
          </string-name>
          . “
          <article-title>HDBscan: Hierarchical Density Based Clustering”</article-title>
          .
          <source>In: Journal of Open Source So昀琀ware 2</source>
          .11 (
          <year>2017</year>
          ), p.
          <fpage>205</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>McInnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Healy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Saul</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Großberger</surname>
          </string-name>
          . “UMAP:
          <article-title>Uniform Manifold Approximation and Projection”</article-title>
          .
          <source>In:Journal of Open Source So昀琀ware 3</source>
          .29 (
          <year>2018</year>
          ), p.
          <fpage>861</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>McIntyre</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Archer</surname>
          </string-name>
          .
          <article-title>“A Corpus-based Approach to Mind Style”</article-title>
          . In: (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Misset</surname>
          </string-name>
          . “
          <article-title>Replete with instruction and rational amusement”?: Unexpected Features in the Register of British Didactic Novels</article-title>
          ,
          <fpage>1778</fpage>
          -
          <lpage>1814</lpage>
          .
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pianzola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rebora</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. Lauer.</surname>
          </string-name>
          “
          <article-title>Wattpad as a Resource for Literary Studies. Quantitative and Qualitative Examples of the Importance of Digital Social Reading and Readers' Comments in the Margins”</article-title>
          .
          <source>In:PloS one 15.1</source>
          (
          <issue>2020</issue>
          ),
          <year>e0226708</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>E.</given-names>
            <surname>Saral</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. G.</given-names>
            <surname>Alhama.</surname>
          </string-name>
          “
          <string-name>
            <given-names>A Topic</given-names>
            <surname>Modeling</surname>
          </string-name>
          <article-title>Study of the COVID-19 Impact in an Online Eating Disorder Community in Reddit”</article-title>
          . In: Tilburg: Tilburg University,
          <year>2022</year>
          . uhrttlp: s: //clin2022.uvt.nl/a-topic%
          <fpage>20</fpage>
          -modeling
          <article-title>-study-of-the-covid-19-impact-in-an-online-eati ng-disorder-community-in-redd</article-title>
          .it/
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>C.</given-names>
            <surname>Schöch</surname>
          </string-name>
          . “
          <source>Topic Modeling Genre: An Exploration of French Classical and Enlightenment Drama.” In: DHQ: Digital Humanities Quarterly 11.2</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Schröter</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Du</surname>
          </string-name>
          . “
          <article-title>Validating Topic Modeling as a Method of Analyzing Sujet and Theme</article-title>
          .”
          <source>In: Journal of Computational Literary Studies</source>
          <volume>1</volume>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>L.</given-names>
            <surname>Thompson</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Mimno</surname>
          </string-name>
          . “Authorless Topic Models:
          <article-title>Biasing Models Away from Known Structure”</article-title>
          .
          <source>In:Proceedings of the 27th International Conference on Computational Linguistics. Santa Fe</source>
          , New Mexico, USA: Association for Computational Linguistics,
          <year>2018</year>
          , pp.
          <fpage>3903</fpage>
          -
          <lpage>3914</lpage>
          . url: https://aclanthology.org/C18-132.
          <fpage>9</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>I.</given-names>
            <surname>Uglanova</surname>
          </string-name>
          and
          <string-name>
            <surname>E. Gius.</surname>
          </string-name>
          “
          <source>The Order of Things. A Study on Topic Modelling of Literary Texts.” In: Chr 18-20</source>
          (
          <year>2020</year>
          ), p.
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [26]
          <string-name>
            <surname>K. Van Rees</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Janssen</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Verboord</surname>
          </string-name>
          . “
          <article-title>Classi昀椀catie in het culturele en literaire veld 1975-2000: Diversi昀椀catie en nivellering van grenzen tussen culturele genres”</article-title>
          . IPnr:oductie van literatuur.
          <source>Het literaire veld in Nederland</source>
          <year>1800</year>
          -2000. Ed. by
          <string-name>
            <given-names>G.</given-names>
            <surname>Dorleijn and K. Van Rees</surname>
          </string-name>
          .
          <source>Nijmegen: Vantilt</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>239</fpage>
          -
          <lpage>283</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [27] “
          <article-title>Nur”</article-title>
          . In:Algemeen letterkundig lexicon. Ed. by G. Vis,
          <string-name>
            <given-names>P.</given-names>
            <surname>Verkruijss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Van</given-names>
            <surname>Gorp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Delabastita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Van Bork</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bernaerts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Willaert</surname>
          </string-name>
          , E. Op de Beek, and
          <string-name>
            <given-names>N.</given-names>
            <surname>Geerdink</surname>
          </string-name>
          . Digitale Bibliotheek voor de Nederlandse Letteren,
          <year>2012</year>
          . urhlt:tps://www.dbnl.
          <source>org/tekst/dela0 12alge01%5C%5F01/dela012alge01%5C%5F01%5C%5F01441.ph</source>
          .p [28]
          <string-name>
            <given-names>M.</given-names>
            <surname>Walsh</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Antoniak</surname>
          </string-name>
          . “The Goodreads “
          <article-title>Classics”: A Computational Study of Readers, Amazon, and Crowdsourced Amateur Criticism”</article-title>
          .
          <source>IJno:urnal of Cultural Analytics</source>
          <volume>4</volume>
          (
          <year>2021</year>
          ), pp.
          <fpage>243</fpage>
          -
          <lpage>287</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wilkens</surname>
          </string-name>
          . “Genre,
          <article-title>Computation, and the Varieties of Twentieth-Century U.S. Fiction”</article-title>
          .
          <source>In: CA Journal of Cultural Analytics 2.2</source>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .22148/16.009. url: https://cultura lanalytics.
          <source>org/article/1106</source>
          .5
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>