<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Aarhus, Denmark
∗Corresponding author.
£ peter.boot@huygens.knaw.nl(P. Boot); j.daza@esciencecenter.nl(A. Daza); c.schnober@esciencecenter.nl
(C. Schnober);w.vanhage@esciencecenter.nl(W. v. Hage)
ȉ</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>In the Context of Narrative, we Never Properly Defined the Concept of Valence</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Peter Boot</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Angel Daza</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carsten Schnober</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Willem vanHage</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Huygens Institute for the History and Culture of the Netherlands (KNAW)</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Netherlands eScience Center</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Valence is a concept that is increasingly being used in the computational study of narrative texts. We discuss the history of the concept and show that the word has been interpreted in various ways. Then we look at a number of Dutch tools for measuring valence. We use them on sample fragments from a large collection of narrative texts and find only moderate correlations between the valences as established by the various tools. We discuss these diferences and how to handle them. We argue that the root cause of the problem is that Computational Literary Studies never properly defined the concept of valence in a narrative context.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;valence</kwd>
        <kwd>polarity</kwd>
        <kwd>sentiment</kwd>
        <kwd>word-embedding</kwd>
        <kwd>narrative</kwd>
        <kwd>computational literary studies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The study of emotion and sentiment is increasingly popular in computational literary studies
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. In this paper we will look specifically at the concept of valence, the positive or negative
sentiment associated with a word or a text passage. The concept is used in a number of recent
studies, but it seems to be used in quite diferent ways. This calls for a deeper look at the
history of the concept.
      </p>
      <p>
        One of the most-quoted studies in the field is the article by Reagan et al. about six basic
shapes in the emotional arcs in stories3[0]. The emotional arcs that the authors create
represent the flow of what they call ‘sentiment’, measured by the ‘Hedonometer’, a dictionary-based
tool that assigns sentiment to words1[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. There is no discussion in the article about the
status of this sentiment: does it correspond to the sentiment that readers experience? Does it
correspond to sentiment of the characters or the narrator? The paper includes an annotated
emotion arc forHarry Potter and the Deathly Hallows which seems to show that the highs and
lows of the story correspond to the maxima and minima in the arc. But the paper treats the
construction of the arcs on the basis of the sentiment data as a purely technical problem, without
asking what it is exactly that these arcs are modelling.
      </p>
      <p>Bizzoni and Feldkamp 2[] do address the issue in a case study onThe Old Man and the Sea.
For each sentence in the novel, two annotators rated ‘the sentiment expressed by the sentence’.
They were instructed ‘to avoid rating how a sentence made them feel and to try to report
only on the sentiments actually embedded in the sentence, i.e., to think about the valence
of each sentence individually, without overthinking the story’s narrative to reduce contextual
interpretation’. This is an interesting instruction, in that it explicitly states the sentiment is not
in the reader and it shouldn’t relate to the story events. It assumes that there is such a thing
as ‘the sentiment embedded in a sentence’. The paper then goes on to check whether LLM or
dictionary-based sentiment models correlate with the annotators’ ratings. We will come back
to this study below.</p>
      <p>
        Rebora, in his survey of sentiment analysis in literary studi3es1][, also asks where the
sentiment is supposed to reside: in the text or in the reader. He notes that in literary studies,
narratologists would opt for the text, while students of reader response would look at the reader.
But he also points at a third possibility: the characters as vehicles of emotion. Nalisnick and
Baird, e.g., have applied sentiment analysis to Shakespeare’s plays in order to study the
relations between the plays characters2[7]. And there are other possibilities: the sentiment that
one finds in the text can also be used to gauge the sentiment of the author, as in Stirman and
Pennebaker’s study of suicidal poets3[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. It is even possible to study the sentiment in novels
and other texts not out of an interest in anything that has to with the book itself, but to read
larger social attitudes. For example, in their study of the perception of coal and oil in recent U.S.
works [
        <xref ref-type="bibr" rid="ref13">14</xref>
        ], Grubert and Algee-Hewitt are interested in the perception of these energy sources
in contemporary U.S. society.
      </p>
      <p>It is clear that all of the above approaches to sentiment can be enlightening. But there is a
danger that we forget that sentiment in texts can have these diferent aspects, and our current
tooling is certainly not able to distinguish them. In an efort to create some clarity, we will
in this paper briefly recount the history of the concept of valence, focusing on the various
definitions researchers have used as well as on how it was established or computed. In a second,
empirical part, we look at a number of tools for computing valence in Dutch. In so far as
they are dictionary-based we look at the overlap between dictionaries and the degree to which
they assign the same values to their shared words. Then we use a sample of fragments from
Dutch fiction to assess how the various tools compare in their assignment of valence to these
fragments.1</p>
      <p>
        Note that our main interest here is not in the sentiment arcs that can be derived from book
segments’ valences (see also 1[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). What we want to contribute to is the much more elementary
question: what is a word or chunk valence in the first place, and how do we compute it?
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background: Valence, Sentiment and Polarity</title>
      <sec id="sec-2-1">
        <title>2.1. Valence in psycho-linguistic studies</title>
        <p>
          We start our short history of the concept of valence with the 1957 booTkhe measurement of
meaning, by Osgood, Suci and Tannenbaum 2[
          <xref ref-type="bibr" rid="ref8">9</xref>
          ]. Osgood and his colleagues wanted to describe
1The notebooks and datasets underlying this paper are availablehattps://doi.org/10.5281/zenodo.1394221.8
the meaning of certain concepts, such as the word ‘lady’, and asked test subjects to associate
these nouns with positions on a scale between opposite adjectives, such as good vs. bad, hard
vs. soft or kind vs. cruel. They used a Likert scale, and the form might look as in Figu1r.e
        </p>
        <p>After averaging, this gave them, for twenty concepts and fity pairs of adjectives, 1000
measurements. A factor analysis identified the hidden dimensions underlying the measurements.
A fragment of the resulting table is reproduced in Figu2r.eWe see that the results of some
adjective pairs, such as good vs. bad and beautiful vs. ugly, are almost completely explained
by the hidden variable I, the scores on strong vs. weak are mostly explained by hidden variable
II, etc. But what are these hidden variables? Osgood et al. then write ‘The problem of labelling
factors is somewhat simpler here than in the usual case. (...) The first factor is clearly
identifiable asevaluative (...). The second variable identifies itself fairly well aspaotency variable (...).
The third factor appears to be mainly aanctivity variable (...)’ (italics original).</p>
        <p>Later, these dimension would become known by other names: evaluativeness as pleasure or
valence, activity as arousal, potency as power or dominance. This is not the whole story, but
Osgood and colleagues made a fundamental contribution to the three-factor theory of emotion.
It is interesting to note that they already mention that these factor loadings depend on
cultures: e.g., for Japanese and Korean respondents, the adjective pair delicate vs. rugged clearly
belonged to the evaluative dimension, for U.S. respondents it did not.</p>
        <p>
          We continue with a look atThe General Inquirer, one of the first text analysis programs [
          <xref ref-type="bibr" rid="ref36">36</xref>
          ].
        </p>
        <p>The 1966 book describes among others the Harvard III dictionary, developed for use with the
General Inquirer. Four groups of words were included to account for the high and low ends
of the evaluativeness and potency variables found by Osgood (pp. 176, 185). We see here that
what for Osgood were hidden variables, the result of a computational process that required the
interpretative step of labelling, have now become measurable entities underlying texts.</p>
        <p>A next step in our tale was set by Bradley and Lang in 1999, in their paper ‘Afective Norms
for English Words’7[]. As prompts in psychological research they needed words with known
values for the dimensions of (what they called) Pleasure, Arousal and Dominance. They asked
subjects to rate the words using a ‘self-assessment mannikin’ (see Figur3efor an example). For
pleasure, the instruction that they gave their subjects was ‘At one extreme of this scale, you
are happy, pleased, satisfied, contented, hopeful. When you feel completely happy you should
indicate this by bubbling in the figure at the left. The other end of the scale is when you feel
completely unhappy, annoyed, unsatisfied, melancholic, despaired, or bored’ (italics ours). The
valence associated with a word now equals (possibly) complete happiness or unhappiness of
the subject, i.e., a subjective feeling.</p>
        <p>
          The creation of ever larger dictionaries of words with associated valence and other variables
has been a constant in psycholinguistic research ever since. We mention a few: Warriner and
colleagues created a list of almost 14000 English lemmas with Valence, Arousal and Dominance
[
          <xref ref-type="bibr" rid="ref40">40</xref>
          ]. They used Likert scales rather than the self-assessment mannikins, but they stuck to the
subjective language: ‘At one extreme of this scale, you are happy, pleased, satisfied, contented,
hopeful. When you feel completely happy you should indicate this by choosing rating 1’. One
reason why their work is remarkable is that they found that there are systematic diferences
between men and women in how they rate various categories of words. The last study for
English words that we mention here is that by Mohammad24[]. Mohammad used a procedure
called best-worst scaling where respondents were asked: ‘Which of the four words below is
associated with the MOST happiness / pleasure / positiveness / satisfaction / contentedness /
hopefulness OR LEAST unhappiness / annoyance / negativeness / dissatisfaction / melancholy
/ despair?’ The choice of words ‘is associated with’ is less subjective and seems to ask for a
more objective relation between the words and the associated values than formulations such as
‘when you feel completely happy’. Still, Mohammad too found important diferences between
various groups of people (by gender, age and self-assessed Big 5 personality characteristics)
with respect to the values they associate with the various words.
        </p>
        <p>
          One study, especially relevant for the second part of this paper, was done by Moors and
colleagues [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]. They created a list of 4300 Dutch words with associated valence and other values.
Participants ‘were asked to judge the extent to which the words in the study referred to
something that is positive/pleasant (“positief/aangenaam”) or negative/unpleasant
(“negatief/onaangenaam”)’. Introducing a third perspective, rather than giving a subjective response to a word
(Bradley) or a judgment about the language system (Mohammad), subjects were now asked
about properties of the objects in the world that the words represent. Moors presents separate
valence scales for the general population, for women and for men. The examples that we could
give of diferences in evaluation between women and men are all distressingly predictable.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Valence or sentiment in consumer reviews</title>
        <p>With the appearance of consumer reviews, the problem of establishing the opinions that they
expressed and the related sentiments became important subjects for marketeer2s1][. As an
example of early studies we mention Hu and Liu’s work15[]. Hu and Liu find sentences in
reviews that contain both a product feature and an opinion. The opinion words and their
polarity (= whether they express a positive or negative sentiment) are deduced from WordNet
by starting with some seedwords. Over the years, their work has produced a 6800-word opinion
lexicon. In the study of consumer reviews, the question is no longer how a word strikes the
reader, but what was the opinion that a writer wanted to express. That also means that the
attention moves to adjectives rather than the nouns that the tradition established by Bradley
[7] was typically interested in.</p>
        <p>
          A Dutch sentiment dictionary with a focus on review sentiment was created by De Smedt
and Daelemans [
          <xref ref-type="bibr" rid="ref7">8</xref>
          ]. They used frequently occurring adjectives from book reviews, and asked
annotators ‘to classify each adjective in terms of positive-negative polarity and subjectivity’.
The question here no longer refers to subjective feeling or properties of an object, but to a
property of the sentiment word. The dictionary is special (among sentiment dictionaries) in
that it distinguishes between multiple word senses. E.g. the word ‘scherp’ (sharp) applied to
a sound has negative polarity, but in ‘a sharp thinker’ the word’s polarity is positive. Using
computational tools and based on several linguistic resources, the initial list of words has been
expanded to 5500 word senses. De Smedt and Daelemans’ dictionary can be used as part of the
Pattern toolset 9[].
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Valence from word embeddings</title>
        <p>Under the name of SentiArt, Arthur Jacobs introduced a completely diferent way of estimating
valence in the field of computational literary studies17[]. The method uses a set of positive
and negative seed words in combination with a word embedding. The valence of a word is then
computed summing its similarities to the positive seed words and subtracting the similarities
to the negative ones. The method is based on the assumption that, in the word embedding,
words with similar meanings cluster together. Because there is no need for human ratings, this
seems like an objective procedure. However, the choice of the texts that are used to create the
word embedding, the procedure that is used for its computation as well as the choice of seed
words are to some extent arbitrary. It is also not self-evident that the procedure should work
at all. Jacobs 1[7, p. 3] states that the method was able to explain 34% of the variance between
words found by Warriner40[], which isn’t exactly promising.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Valence beyond the word</title>
        <p>The various approaches to the valence concept that we have discussed all assign valence at the
word level. It is not self evident that it is possible to define valence at a higher level, e.g. that
of sentences, paragraphs or even book chapters.</p>
        <p>
          Bradley and Lang extended their work on word valence to small texts (one to a few sentences)
[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] These small texts describe, in the second person, situations that would probably cause an
emotional state with a certain valence (as well as arousal and dominion). We give an example
with low valence: ‘You gag, seeing a roach moving slowly over the surface of the pizza. You
knock the pie on the floor. Warm cheese spatters on your shoes’.
        </p>
        <p>
          Specifically in the context of narrative, Rebora asked students to rate the sentiment in
paragraphs of a story by Pirandello32[]. Their agreement was very weak. As we saw, in 2[] two
researchers rated sentiment in sentences inThe Old Man and the Sea. Their result was better
than Rebora’s: the correlation between their ratings was strong, after detrending even very
strong. Kaakinen and colleagues report on a multilingual database of short stories (ca. 1,000
characters) [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], for which raters established valence and arousal using self-assessment
mannikins. Finally, Jacobs 1[6, pp. 117-119] briefly reports on a number of experiments where
readers were asked to rate the valence of sentiments or short sections of among others a Harry
Potter novel andPippi Longstocking. He does not report agreement measures, but the resulting
sentiment arcs could be predicted reasonably well using his SentiArt toolset.
        </p>
        <p>
          Outside of the domains of psychology or narrative, there exists a plethora of studies on the
polarity of especially reviews and social media texts. We just mention work on the orientation
of tweets [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ], targeted not so much at establishing a text’s valence, but at classifying a text
as positive, negative or neutral, and work on product reviews, where beyond establishing a
text’s overall valence the aim is to find the aspects of a product that reviewers are positive or
negative about [
          <xref ref-type="bibr" rid="ref41">41</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Provisional conclusion</title>
        <p>We have seen how the concept of valence morphed from a hidden dimension of meaning into
something that could be measured in text using just a few adjectives, and further into a
subjective feeling, a property of the language or a property of the things that we talk about. It is
clear that in practice, these are not unrelated. If I consider peace a good thing, I’ll probably feel
good about the word and I’ll use positive words in talking about it. But conceptually they are
distinct, and the extent to which they agree in practice is an empirical question. Another thing
that we learned in this short review is that valence judgments vary by gender, age, personality
and culture. Anyone who confidently writes aboutth‘e sentiment’ of a narrative work should
be aware of that.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <p>After this brief look at the history and the operationalisation of the notion of valence, we turn
to a comparison of various tools that assign valence to Dutch words and texts. Except for the
approaches already discussed, we also include two transformer-based language models. These
tools by themselves have no notion of valence, but have been trained to fulfill classification
tasks that assign short texts to evaluative categories.</p>
      <p>Our interest here is not in finding the tool that best approximates the ‘true’ sentiment value,
if such a thing exists, or the gold labels, which we don’t have. What we are interested in is
whether these tools agree or disagree and what that says about the current state of sentiment
analysis for narrative, at least for Dutch.</p>
      <p>The tools that we use are:
LiLaH the sentiment mapping of LiLaH [22], a manually corrected version of an automatic
translation of the NRC emotion lexicon25[], see A.2.2.</p>
      <p>LIWC a general-purpose tool for text analys5is][based on an underlying dictionary, seAe.2.3.</p>
      <sec id="sec-3-1">
        <title>Moors discussed above, see2.1.</title>
      </sec>
      <sec id="sec-3-2">
        <title>Pattern discussed above, see2.2.</title>
      </sec>
      <sec id="sec-3-3">
        <title>SentiArt discussed above, see2.3.</title>
      </sec>
      <sec id="sec-3-4">
        <title>VanRoy a transformer-based model trained on book reviews, sAe.e2.5.</title>
        <p>xlm a transformer-based multilingual mod1e]l b[ased on Twitter, seeA.2.5.</p>
        <p>For a further description of the tools that we will use, as well as for the settings that we apply
in computing the SentiArt valences, we refer to the AppendixA(.1).</p>
        <p>For the dictionary-based tools (that includes the tools with curated dictionaries as well as
SentiArt), we first compare the dictionaries. We report a number of measures:
• dictionary size;
• number of shared words;
• word-level agreement of assigned values;
• words where the dictionaries disagree.</p>
        <p>After the dictionary comparison, we compare the result of the tools on a sample of Dutch
novels. We create the sample as follows: from a collection of 10,921 recently published Dutch
books we remove non-fiction and books with less than 5000 words, then select every fith
book. For every selected book (n=2087) we select a random newline character, which usually
corresponds with a paragraph start. Starting from that location we select the first hundred
tokens.</p>
        <p>Then, for each of the tools mentioned in the Appendix we compute the valence assigned
to the fragment. For Moors, LiLaH and SentiArt, the valence of the text fragment is the
average lemma valence for all words whose lemma occurs in the relevant lexicon. Words whose
lemma does not occur in the resource are ignored. For the other tools, see the Appendix for
the computational procedure.</p>
        <p>We then compute correlations between the tools’ valence assignments. For some pairs of
tools we also look at fragments where the two tools produce very diferent results.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>In describing correlations, we use the labels proposed by Evan13s][: 0.00 - 0.19: very weak,
0.20 - 0.39: weak, 0.40 - 0.59: moderate, 0.60 - 0.79: strong, 0.80 - 1.00: very strong.</p>
      <sec id="sec-4-1">
        <title>4.1. Comparison at dictionary level</title>
        <p>For simplicity’s sake, in this section we ignore LIWC15, as it did much worse than LIWC07 in
the analysis of the fragment valences.
4.1.1. Overlap at dictionary level
4.1.2. Agreement at dictionary level
We look first at the agreement between the continuously valued (SentiArt, Moors and Pattern)
and the binary-valued tools (LIWC07 and LiLaH), given in tabl1e. These look as one would
expect. In all cases there is a clear distinction between the mean values of the positive and
the negative words in LiLaH and LIWC07. For LIWC07, the diference between the means is
somewhat larger than for LILaH. LIWC07 seems to agree better with the continuous tools than
LiLaH does.</p>
        <p>Then we look at the correlations between the continuously valued valences (Ta2b)l.eThe
correlations of Moors and Pattern with SentiArt are strong, the correlation of the two curated
dictionaries Moors and Pattern is very strong. We should be aware, however, that these are
correlations over a relatively small number of words.</p>
        <p>Finally, for all dictionary pairs we also checked whether there are words where the
dictionaries disagree, and if so, whether there is an obvious culprit. For LIWC07 and Moors and to
a lesser extent Pattern we hardly found apparent mistakes. LiLaH contains a number of clear
errors, maybe due to the limited availability of the Dutch translat2o2r, p[. 154]. In the
SentiArt dictionary, there are countless misclassifications. See TablAe.5 for examples. We also
note here that among the top positive words in the SentiArt valences, there appears a curious
group of words related to hospitality, such as dinner, hostess, sommelier, catering, service and
culinary, which also raises some questions about the adequacy of the procedure.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Comparison at text fragment level</title>
        <p>After computing the valences assigned to the fragments we computed their correlations (see
Figure4). We see that the correlations between the results of the various sentiment analysis
tools that we have looked at doesn’t get better than moderate (the only strong correlation is
between the two LIWC flavours). As we don’t know the ‘true’ valences, we’re not in a position
to say which is the best tool, but we can say something.</p>
        <p>1. The correlation of the XLM transformer-based model is with the others is at best weak,
for the Van Roy model there is no correlation. But we already knew that these tools were
very diferent than the other tools, and it confirms (if confirmation were needed) that the
type of training text is really important.
2. Of the other dictionaries, the agreement of Pattern with the rest is at best moderate but
mostly weak. The reason is probably Pattern’s background in consumer review analysis.
3. From the two LIWC dictionaries, LIWC15 performs noticeably less than LIWc07. This is
probably due to the Dutch LIWC15 being an automatic translation of the English
dictionary.
4. The remaining curated dictionaries (LiLaH, LIWC07 and Moors) and the SentiArt
approach have moderate correlations with each other.</p>
        <p>For the tools Moors, SentiArt and LIWC07 we also did an analysis of text fragments that
scored high on one tool and low on another. The main causes for diferences were apparent
errors in the SentiArt word valence, homonymies, the word ‘niet’ (not) in Moors being assigned
a low valence, and words not being present in the curated dictionaries. Some of those are
unavoidable in the context of a dictionary approach, others could be avoided by better curation.
Some of the options that we used for the SentiArt computations resulted from preliminary
testing with these diferently-rated fragments. See the Appendix for details.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>As the main results of the empirical part of this paper we see:
1. The SentiArt procedure to assign valence to large collections of words has serious
limitations, even when computed on the basis of a domain-specific word embedding.
2. These limitations can to some extent be overcome by computing distances to a centroid
vector rather than to the individual seed words, by only looking at the top and bottom
quartile of the resulting valence distribution, and by excluding punctuation and function
words (see Appendix for details). However, while leaving out punctuation and function
words from the fragment valence computation helps, it is not a real solution to the
problem that apparently the word embedding-based valence assignment is producing flawed
results.
3. The agreement between other tools for computing valence of Dutch narrative text are
never better than moderate.</p>
      <p>With respect to the first two items in this list, this suggests that we may have to look
beyond word2vec for a better answer to questions of semantic relatedness between lemmas (e.g.
to contextual text representations as provided by Transformer-based language mod1e0l]s).[
The limited number of pretty arbitrary seed words seems another limitation of the SentiArt
approach. A better way of obtaining valence ratings for many more words than can be manually
curated might be machine learning with as target the Moors valences, and as features (among
others) the word2vec distances to some of the top and bottom Moors words.</p>
      <p>With respect to the last item, the question is: how bad is it that these tools only agree
moderately? If we knew that one of the tools is mostly correct, it wouldn’t matter, we could
just stop using the others. But we suspect that this is not the case. We have seen enough
limitations in all of the tools and in the dictionary approach as such that it is unlikely that any
of these tools presents us with more than a rough approximation of correctness.</p>
      <p>
        That might lead us to asking why we have focussed here on dictionary-based approaches. In
their survey of sentiment analysis in literary studies, Kim and Kling2e0r] [wrote that ‘much
digital humanities research (especially dealing with text) uses the methods of text analysis that
were in fashion in computational linguistics twenty years ago’. And in a direct comparison,
current machine learning tools usually perform better in predicting human valence ratings
(see e.g. [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ] for a study by Van Atteveldt and colleagues in the field of politicology). So why
do we still study these methods, rather than follow the lead of computational linguistics?
      </p>
      <p>
        One answer to that question could come from Teodorescu and Mohamma3d7][, who show
that, in spite of instance-level inaccuracy, dictionary-based methods work very well for larger
bins of texts (e.g. for groups of 30 or a 100 tweets). They argue that ‘[f]or applications where
simple, interpretable, low-cost, and low-carbon-footprint systems are desired, the
lexiconbased systems [...] are often more suitable’. That suggests that dictionary-based methods might
be better suited to study sentiment at the chapter than at the sentence level of a novel. Another
answer comes from Öhman 2[
        <xref ref-type="bibr" rid="ref7">8</xref>
        ], who argues that for the large texts that in the humanities we
are often interested in, the annotation eforts on which machine learning tools depend are
are just not feasible. But the best answer is maybe one that Öhman also hints at when she
proposes to leave the term ‘sentiment analysis’ to the computational linguists and argues that
even if it is not computational sentiment, diferences between texts that are made visible by
dictionary-based tools, if statistically significant, are still relevant research findings.
      </p>
      <p>We wouldn’t go so far as as to say that what dictionary-based methods can do is not
sentiment analysis. But it is true that most of the work in computational linguistics has been on the
detection of sentiment in the sense of stance, where the aim is to detect the view that a text’s
author expresses about some object. In narrative, and especially literary narrative, the aim of
the text is not to convey the author’s view about the characters or the events, and if it were, it
wouldn’t necessarily be the aim of researchers to uncover that view. This doesn’t mean that we
don’t need to work on well-annotated corpora of narrative on which we can apply the tools of
machine learning, far from that, but it does mean that current pre-trained sentiment analysis
tools have been trained on corpora so diferent from the corpora that we are interested in that
they may not be very relevant to the analysis of narrative.</p>
      <p>Returning to the question of how much of a problem we have with these moderate
correlations, and assuming, for the sake of argument, that the correlation of our tools with the ‘true’
valence is about equal to the best of their mutual correlations, that is .51, what is it can we do
with a measurement that misses so much information? Maybe we could look at some patterns,
very carefully. But it certainly would not make sense to use these measurements as ingredients
in e.g. predictive modelling or the construction of narrative arcs.</p>
      <p>We see some ways of moving forward:
1. creating a much larger dictionary than the present curated dictionaries for Dutch along
the lines sketched earlier in this section;
2. using some sort of ensemble measure, in the hope that the tools can compensate for each
other’s weaknesses;
3. only using sentiment analysis in a narrative context on larger text segments;
4. starting an annotation efort for valence in narrative fragments.</p>
      <p>
        All of these, however, are only stop-gap measures for what we believe is the real problem,
which is that we have as a discipline not really defined the concept of valence in a narrative
context. As we saw in our overview of the history of the concept of word valence in section
2, many completely diferent definitions and operationalisations have have been proposed. It
has been possible to get away with these diferences in the analysis of by and large simple and
straightforward texts such as social media posts. In his survey of sentiment analysis in literary
studies [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ], Rebora writes ‘S[entiment] A[nalysis], in fact, can be performed by selecting or
combining an ample variety of approaches [...]. Choosing one approach over the other means
also defining the very nature of the object under examination’. We might add that to ‘define
the very nature of the object under examination’ is what, with respect to valence, we have up
to now, and to our peril, shied away from.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments References</title>
      <p>This research was funded by the Netherlands eScience Center, grant number ASDI.2020.032.</p>
      <p>M. M. Bradley and P. J. Lang.Afective Norms for English Words (ANEW): Instruction
Manual and Afective Ratings . Tech. rep. Technical report C-1, the center for research in
psychophysiology, 1999.
[22]</p>
    </sec>
    <sec id="sec-7">
      <title>A. Appendix</title>
      <sec id="sec-7-1">
        <title>A.1. Tools</title>
      </sec>
      <sec id="sec-7-2">
        <title>A.2. Moors: Norms of valence, (...) for 4,300 Dutch words</title>
        <p>The Moors approach is based on the Moors et al. articl2e6][ discussed in the text. The valences
reported by Moors vary from 1 (lowest) to 7 (highest).</p>
        <p>A.2.1. Pattern: A Subjectivity Lexicon for Dutch Adjectives
We use the De Smedt and Daelemans subjectivity dictionary discussed in the tex8t][. For the
dictionary comparison, if a word occurs in the dictionary multiple times (because of multiple
word senses) we take the average valence. For the computation of the fragment valence, we do
not use the dictionary directly, but apply the Pattern toolset to the unlemmatised text fragment.
Pattern does not just look at individual words, but uses some aspects of the context, such as
the presence of intensifying adverbs (‘awfully good’).</p>
        <p>
          A.2.2. LiLaH: The LiLaH Emotion Lexicon of Croatian, Dutch and Slovene
The LiLaH dictionary 2[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is a manually corrected version of an automatic translation of the
NRC emotion lexicon 2[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] for three languages. It assigns words to positive or negative
sentiment (+1 or -1) as well as to specific emotions. For Dutch, however, only the sentiment values
are available. It contains 5746 Dutch words, 2519 are positive, 3431 are negative and 204 are
both positive and negative. In the computation of the fragment valence, if a word is both
positive and negative, we count its value as 0.
        </p>
        <p>
          A.2.3. LIWC: Linguistic Inquiry and Word Count
LIWC is a general-purpose tool for text analysis created by psychologist James Pennebaker
and colleagues. Its latest version is LIWC 20225][. Underlying the tool is an English-language
dictionary that assigns words to (multiple) categories, including categories for positive and
negative emotion. There exist Dutch translations for the 20074][and 2015 [
          <xref ref-type="bibr" rid="ref39">39</xref>
          ] versions of LIWC.
The translation of the LIWC 2007 dictionary is a manual translation. It includes wildcards. The
translation of LIWC 2015 is an automatic translation that resolved wildcards.
        </p>
        <p>What LIWC reports is the relative frequency in a text of words in the various categories. We
compute the relative frequencies using the LIWCTools python packag3e][. For computation of
the fragment valence we subtract the relative frequency of negative emotion from the relative
frequency of positive emotion.</p>
        <p>
          A.2.4. SentiArt: a word-embedding based computation
For SentiArt, we use a word2vec-computed23[] word embedding. We used the gensim
package [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] to do the computation; we pre-tokenized, lemmatized, and lowercase the text with a
window size of 8 words and only counted lemmas that appear in the corpus 5 or more times.
        </p>
        <p>To compute or not compute the centroids (average value) for the
positive words and the negative seed words before we compute and
subtract the similarities.</p>
        <p>To use a list of noun-only seed words or to include also corresponding
adjectives (and one verb, where there is no corresponding adjective).</p>
        <p>To use or not to use the lens method [18, p. 22], which excludes the
second and third quartiles of the valence distribution from the
computation.</p>
        <p>To exclude or not punctuation characters.</p>
        <p>To exclude or not function words. We use the function words as
defined by Dutch LIWC 2007.</p>
        <p>The texts that we used for the word-embedding are the full texts of 13,210 novels taken from a
larger collection of 18,467 books in Dutch.</p>
        <p>We use a number of diferent options in the SentiArt computation of the word and the
fragment valences (see tableA.1). Table A.2 gives the seed words that we use for computing the
SentiArt valence. The fact that we included among the SentiArt options the possibilities to
exclude punctuation and/or function words is the result of preliminary testing. We saw in these
tests that punctuation and function word valences had a sizable efect on the SentiArt-assigned
fragment valences, even under the ‘lens’ condition. E.g. the comma, the full stop and the
indefinite article ‘een’ (‘a’) all have a SentiArt valence in the upper quartile of the distribution.</p>
        <p>After computing the four SentiArt valence dictionaries, we look at their distributions. We
also list the 50 words with the highest and lowest valence, in order to check whether the
computation makes sense.</p>
        <p>A.2.5. Transformer-based models
As mentioned, we use two transformer-based models on the fragments. One model is the
cardifnlp/twitter-xlm-roberta-base-sentiment model1][. This is a multilingual model, trained
on tweets. It does not assign a continuous valence value, but classifies a text as positive,
negative or neutral2. The other model is robbert-v2-dutch-base-hebban-reviews5. It is a model
trained on Dutch book reviews from the book discussion site Hebban, and aims to predict the
rating associated with the review3. For both tools, the text types that they were trained on are
very diferent from the book fragments that we will use them on. For this reason, we did not
expect that they would agree with the other tools, but were willing to be surprised.
Table A.2
Valence seed words</p>
        <p>Condition</p>
        <p>Positive
Basic
Extended
tevredenheid (contentment)
blijdschap (joy)
genot (pleasure)
trots (pride)
opluchting (relief)
voldoening (satisfaction)
verrassing (surprise)
Basic seed words &amp;
tevreden (contented)
blij (glad)
genieten (enjoy)
trots (proud)
opgelucht (relieved)
voldaan (satisfied)
verrast (surprised)</p>
        <p>Negative
walging (disgust)
verlegenheid (shyness)
angst (anxiety)
verdriet (sadness)
schaamte (shame)
walgend (disgusted)
verlegen (shy)
angstig (anxious)
verdrietig (sad)
beschaamd (ashamed)</p>
      </sec>
      <sec id="sec-7-3">
        <title>A.3. Computing the SentiArt valences</title>
        <p>We computed four SentiArt valences, as described above. Here we used only the 17,306 lemmas
that occur in the novel fragments that we will analyse.</p>
        <p>As a first check of the results, we looked at the 50 words with highest or lowest valence
for each of the computations. The words with the reportedly lowest valence are for all four
computations indeed words that describe very unpleasant things. As an example, here are
(English translations of) the 10 words with lowest valence for the centroid - nouns and adjectives
condition: fear, distraught, anger, anxious, rage, shame, misunderstood, powerlessness, anger,
confounded. For the words with highest valence, the picture is somewhat diferent. There are
many words that no doubt represent a positive evaluation (excellent, fantastic, great), but there
also appears a curious group of words that seem somehow related to hospitality, such as dinner,
hostess, sommelier, catering, service and culinary; words that certainly are far removed from
the positive seed words that went into the process.</p>
        <p>Next we look at the distribution of the computed valences. FigurAe.1 shows that the
centroid-based computations, and especially the one with nouns and adjectives as seed words,
have a somewhat wider distribution. That seems an attractive property, as it provides stronger
distinctive power to the valence assignment.</p>
        <p>The correlations between the SentiArt valences are very strong (TabAl.e3). We see there
is a 2 to 3 percent disagreement between the centroid and non-centroid versions, and a 6 to 7
percent disagreement between the nouns versus the nouns and adjectives seed words.</p>
        <p>For each dictionary pair, we selected 10 words that get a high valence rating in one dictionary
but a low rating in another. In none of the pairs, this created word lists where intuitively
we would consider one of the dictionaries wrong, except for the pair centroid and nouns /
noncentroid and nouns. Here, the words that were rated high in the centroid but low in the
Figure A.1: Distribution of the four SentiArt valences. ‘ct’: centroid, ‘nct’: non-centroid, ‘n’: noun
labels, ‘na’: noun and adjective labels.
non-centroid condition were all obviously positive: humour, eagerness, liveliness, surrender,
delight, self-assurance, cheerfulness, approval, passion, lust. This provides another argument
in favour of the centroid-based computation.</p>
        <p>As explained in the previous section, for SentiArt we have five binary options, and
therefore 32 diferent results. In initial testing, it appeared that the computation without centroid
consistently led to results with lower correlations to the other tools than the computation with
centroid. We dropped the computation without centroid from further consideration. From
the remaining sixteen SentiArt valences, the best correlation with the other tools was reached
with the options lens, just the original (noun) labels, and not considering punctuation and
stop words (see FigureA.2 for Pearson correlation with the other dictionary-based tools.). We
continue with this SentiArt valence, which in the rest of the paper we just call SentiArt.</p>
      </sec>
      <sec id="sec-7-4">
        <title>A.4. Other tables and figures</title>
        <p>Table A.4
Overlap in words between the tools’ dictionaries. A limitation to take into account: in LIWC07, some
terms include wildcards. The number of words that it covers are therefore larger than the number
reported here. The computation of the overlap does not take into account the wildcards.
tool 1</p>
        <p>tool 2
sentiart
sentiart fake sorrowful rigid boring disdain- mindful forgive kiss compassion
ful sucky unattractive loved ode innocence
m ad ,
n
o
e ab :’</p>
        <p>p
h l ‘
t</p>
        <p>n ,s
n . io
o n t</p>
        <p>c
r ,
ex ie d</p>
        <p>s</p>
        <p>no fu
fo it e</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Barbieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Anke</surname>
          </string-name>
          , and J.
          <string-name>
            <surname>Camacho-Collados</surname>
          </string-name>
          .
          <article-title>“XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond”. IPnr:oceedings of the Thirteenth Language Resources</article-title>
          and Evaluation Conference. Marseille, France,
          <year>2022</year>
          , pp.
          <fpage>258</fpage>
          -
          <lpage>266</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bizzoni</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Feldkamp</surname>
          </string-name>
          . “
          <article-title>Comparing Transformer and Dictionary-Based Sentiment Models for Literary Texts: Hemingway as a Case-Study”</article-title>
          .
          <source>InP:roceedings of the Joint 3rd International Conference on Natural Language Processing for Digital Humanities and 8th International Workshop on Computational Linguistics for Uralic Languages</source>
          . Tokyo, Japan,
          <year>2023</year>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>228</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Boot</surname>
          </string-name>
          .
          <source>LIWCTools. Version 1.3.3</source>
          .
          <year>2016</year>
          . url:https://github.com/pboot/LIWCtool.s
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Boot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zijlstra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Geenen</surname>
          </string-name>
          . “
          <article-title>The Dutch Translation of the Linguistic Inquiry and Word Count (LIWC) 2007 Dictionary”</article-title>
          .
          <source>InD:utch Journal of Applied Linguistics 6.1</source>
          (
          <issue>2017</issue>
          ), pp.
          <fpage>65</fpage>
          -
          <lpage>76</lpage>
          . doi:
          <volume>10</volume>
          .1075/dujal.6.1.04boo.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Boyd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ashokkumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Seraj</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Pennebaker</surname>
          </string-name>
          . “
          <article-title>The Development</article-title>
          and Psychometric Properties of LIWC-
          <volume>22</volume>
          ”. InA: ustin, TX: University of Texas at Austin 10 (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. M.</given-names>
            <surname>Bradley</surname>
          </string-name>
          and
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Lang</surname>
          </string-name>
          . “
          <article-title>Afective Norms for English Text (ANET): Afective ratings of text and instruction manual”</article-title>
          .
          <source>InT:echical Report. D-1</source>
          , University of Florida, Gainesville, FL (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <surname>T. De Smedt</surname>
            and
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Daelemans</surname>
          </string-name>
          . “”Vreselijk mooi!”
          <article-title>(terribly beautiful): A Subjectivity Lexicon for Dutch Adjectives</article-title>
          .” In:Lrec. Istanbul, Turkey,
          <year>2012</year>
          , pp.
          <fpage>3568</fpage>
          -
          <lpage>3572</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <surname>T. De Smedt</surname>
            and
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Daelemans</surname>
          </string-name>
          . “
          <article-title>Pattern for Python”</article-title>
          .
          <source>InT:he Journal of Machine Learning Research 13.1</source>
          (
          <issue>2012</issue>
          ), pp.
          <fpage>2063</fpage>
          -
          <lpage>2067</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          . “BERT:
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding”. PInro:ceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)</article-title>
          . Minneapolis, Minnesota,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . doi:
          <volume>10</volume>
          .48550/arXiv.
          <year>1810</year>
          .
          <volume>04805</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Dodds</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Desu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Reagan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Williams</surname>
          </string-name>
          , L. Mitchell,
          <string-name>
            <surname>K. D. Harris</surname>
            ,
            <given-names>I. M.</given-names>
          </string-name>
          <string-name>
            <surname>Kloumann</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          <string-name>
            <surname>Bagrow</surname>
          </string-name>
          , et al. “
          <article-title>Human Language Reveals a Universal Positivity Bias”</article-title>
          .
          <source>In:Proceedings of the national academy of sciences 112.8</source>
          (
          <issue>2015</issue>
          ), pp.
          <fpage>2389</fpage>
          -
          <lpage>2394</lpage>
          . doi:
          <volume>10</volume>
          .1073/pnas.1411678112.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Elkins</surname>
          </string-name>
          .
          <article-title>The shapes of Stories: Sentiment Analysis for Narrative</article-title>
          . Cambridge University Press,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .1017/9781009270403.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Evans</surname>
          </string-name>
          .
          <source>Straightforward Statistics for the Behavioral Sciences. Thomson Brooks/Cole Publishing Co</source>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E.</given-names>
            <surname>Grubert</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Algee-Hewitt</surname>
          </string-name>
          .
          <article-title>“Villainous or Valiant? Depictions of Oil and Coal in American Fiction and Nonfiction Narratives”</article-title>
          .
          <source>In:Energy research &amp; social science 31</source>
          (
          <year>2017</year>
          ), pp.
          <fpage>100</fpage>
          -
          <lpage>110</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.erss.
          <year>2017</year>
          .
          <volume>05</volume>
          .030.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15] [16] [17] [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hu</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          . “
          <article-title>Mining and Summarizing Customer Reviews”</article-title>
          .
          <article-title>InP:roceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <year>2004</year>
          , pp.
          <fpage>168</fpage>
          -
          <lpage>177</lpage>
          . doi:
          <volume>10</volume>
          .1145/1014052.1014073.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          .
          <article-title>Neurocomputational Poetics: How the Brain Processes Verbal Art</article-title>
          . Anthem Press,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          . “
          <article-title>Sentiment Analysis for Words and Fiction Characters from the Perspective of Computational (Neuro-) Poetics”</article-title>
          .
          <source>InF:rontiers in Robotics and AI</source>
          <volume>6</volume>
          (
          <year>2019</year>
          ), p.
          <fpage>53</fpage>
          . doi:
          <volume>10</volume>
          .3389/frobt.
          <year>2019</year>
          .
          <volume>00053</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Kinder</surname>
          </string-name>
          . “
          <article-title>Computing the Afective-Aesthetic Potential of Literary Texts”</article-title>
          .
          <source>In: Ai 1</source>
          .1 (
          <issue>2019</issue>
          ), pp.
          <fpage>11</fpage>
          -
          <lpage>27</lpage>
          . doi:
          <volume>10</volume>
          .3390/ai1010002.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Kaakinen</surname>
          </string-name>
          , E. Werlen,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kammerer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Acartürk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Aparicio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baccino</surname>
          </string-name>
          , U. Ballenghein,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bergamin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Castells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Costa</surname>
          </string-name>
          , et al. “
          <article-title>IDEST: International database of emotional short texts”</article-title>
          .
          <source>InP:LOS one 17.10</source>
          (
          <year>2022</year>
          ),
          <year>e0274480</year>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0</volume>
          <fpage>274480</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kim</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Klinger</surname>
          </string-name>
          .
          <article-title>“A Survey on Sentiment and Emotion Analysis for Computational Literary Studies”</article-title>
          .
          <source>In:Zeitschrift für digitale Geisteswissenschaften</source>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .17175/2 019\_008\_v2.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>B. Liu. Sentiment</given-names>
            <surname>Analysis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Opinion</given-names>
            <surname>Mining</surname>
          </string-name>
          . Morgan Claypool,
          <year>2012</year>
          . doi1:
          <fpage>0</fpage>
          .1007/97 8-
          <lpage>3</lpage>
          -
          <fpage>031</fpage>
          -02145-9.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>N.</given-names>
            <surname>Ljubešić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Markov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fišer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Daelemans</surname>
          </string-name>
          . “
          <article-title>The LiLaH emotion lexicon of Croatian, Dutch and Slovene”</article-title>
          .
          <source>In:Proceedings of the Third Workshop on Computational Modeling of People's Opinions</source>
          , Personality, and
          <article-title>Emotion's in Social Media</article-title>
          . Barcelona,
          <string-name>
            <surname>Spain</surname>
          </string-name>
          (Online),
          <source>ACL</source>
          , pp.
          <fpage>153</fpage>
          -
          <lpage>157</lpage>
          , December,
          <year>2020</year>
          . Barcelona,
          <source>Spain (Online)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          . “
          <article-title>Distributed Representations of Words and Phrases and their Compositionality”</article-title>
          .
          <source>IPnr:oceedings of the 26th International Conference on Neural Information Processing Systems - Volume 2. Nips'13</source>
          .
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA: Curran Associates Inc.,
          <year>2013</year>
          , pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mohammad</surname>
          </string-name>
          . “
          <article-title>Obtaining Reliable Human Ratings of Valence, Arousal,</article-title>
          and Dominance for
          <volume>20</volume>
          ,000 English Words”.
          <article-title>InP: roceedings of the 56th annual meeting of the association for computational linguistics (volume 1: Long papers)</article-title>
          .
          <source>Melbourne, Australia</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>174</fpage>
          -
          <lpage>184</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P18</fpage>
          -1017.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Mohammad</surname>
          </string-name>
          and
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Turney</surname>
          </string-name>
          . “
          <article-title>Crowdsourcing a Word-emotion Association Lexicon”</article-title>
          .
          <source>In: Computational intelligence 29.3</source>
          (
          <issue>2013</issue>
          ), pp.
          <fpage>436</fpage>
          -
          <lpage>465</lpage>
          . doi:
          <volume>10</volume>
          .1111/j.1467-
          <fpage>8640</fpage>
          .
          <fpage>20</fpage>
          <lpage>12</lpage>
          .00460.x.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moors</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. De Houwer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Hermans</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Wanmaker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Van Schie</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.-L. Van Harmelen</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. De Schryver</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. De Winne</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Brysbaert</surname>
          </string-name>
          . “Norms of Valence, Arousal, Dominance, and
          <article-title>Age of Acquisition for 4,300 Dutch Words”</article-title>
          .
          <source>InB:ehavior research methods 45</source>
          (
          <year>2013</year>
          ), pp.
          <fpage>169</fpage>
          -
          <lpage>177</lpage>
          . doi:
          <volume>10</volume>
          .3758/s13428-012-0243-8.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>E. T.</given-names>
            <surname>Nalisnick</surname>
          </string-name>
          and
          <string-name>
            <given-names>H. S.</given-names>
            <surname>Baird</surname>
          </string-name>
          . “
          <article-title>Character-to-character Sentiment Analysis in Shakespeare's Plays”</article-title>
          .
          <source>In:Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers).</given-names>
          </string-name>
          <year>2013</year>
          , pp.
          <fpage>479</fpage>
          -
          <lpage>483</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>E. Öhman. “</surname>
          </string-name>
          <article-title>The validity of lexicon-based sentiment analysis in interdisciplinary research”</article-title>
          .
          <source>In: Proceedings of the workshop on natural language processing for digital humanities. NIT Silchar</source>
          , India,
          <year>2021</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Osgood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Suci</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Tannenbaum</surname>
          </string-name>
          .
          <source>The Measurement of Meaning. 47</source>
          . University of Illinois press,
          <year>1957</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Reagan</surname>
          </string-name>
          , L. Mitchell,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kiley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Danforth</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Dodds</surname>
          </string-name>
          . “
          <article-title>The Emotional Arcs of Stories are Dominated by Six Basic Shapes”</article-title>
          .
          <source>InE: PJ data science 5</source>
          .1 (
          <issue>2016</issue>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . doi:
          <volume>10</volume>
          .1140/epjds/s13688-016-0093-1.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rebora</surname>
          </string-name>
          . “
          <article-title>Sentiment Analysis in Literary Studies. A Critical Survey”</article-title>
          .
          <source>DInH: Q: Digital Humanities Quarterly 17.3</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rebora</surname>
          </string-name>
          et al. “
          <article-title>Shared Emotions in Reading Pirandello. An Experiment with Sentiment Analysis”</article-title>
          . In:Marras,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Passarotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Franzini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            , and
            <surname>Litta</surname>
          </string-name>
          , E.(eds),
          <article-title>Atti del IX Convegno Annuale AIUCD</article-title>
          .
          <article-title>La svolta inevitabile: sfide e prospettive'per l'Informatica Umanistica</article-title>
          .
          <source>Università Cattolica del Sacro Cuore</source>
          ,
          <string-name>
            <surname>Milano</surname>
          </string-name>
          (
          <year>2020</year>
          ) (
          <year>2020</year>
          ), pp.
          <fpage>216</fpage>
          -
          <lpage>221</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rehurek</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Sojka</surname>
          </string-name>
          . “
          <article-title>Gensim-Python Framework for Vector Space Modelling”</article-title>
          . In: NLP Centre, Faculty of Informatics, Masaryk University, Brno,
          <source>Czech Republic 3.2</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          . “SemEval
          <article-title>-2017 task 4: Sentiment Analysis in Twitter”</article-title>
          . In: arXiv preprint arXiv:
          <year>1912</year>
          .
          <volume>00741</volume>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>S17</fpage>
          -2088.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>S. W.</given-names>
            <surname>Stirman</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Pennebaker</surname>
          </string-name>
          . “
          <article-title>Word Use in the Poetry of Suicidal and Nonsuicidal Poets”</article-title>
          .
          <source>In: Psychosomatic medicine 63.4</source>
          (
          <issue>2001</issue>
          ), pp.
          <fpage>517</fpage>
          -
          <lpage>522</lpage>
          . doi:
          <volume>10</volume>
          .1097/
          <fpage>00006842</fpage>
          -2001
          <fpage>07000</fpage>
          -
          <lpage>00001</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Stone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Dunphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <surname>and D.</surname>
          </string-name>
          <article-title>M. OgilvieT</article-title>
          .he
          <article-title>General Inquirer: A Computer Approach to Content Analysis</article-title>
          . MIT press,
          <year>1966</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>D.</given-names>
            <surname>Teodorescu</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Mohammad</surname>
          </string-name>
          . “
          <article-title>Evaluating Emotion Arcs across Languages: Bridging the Global Divide in Sentiment Analysis”</article-title>
          .
          <article-title>InF:indings of the Association for Computational Linguistics: EMNLP 2023</article-title>
          . Singapore,
          <year>2023</year>
          , pp.
          <fpage>4124</fpage>
          -
          <lpage>4137</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .findings-emnlp.
          <volume>271</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>W.</given-names>
            <surname>Van Atteveldt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Van der Velden</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Boukes</surname>
          </string-name>
          . “
          <article-title>The Validity of Sentiment Analysis: Comparing Manual Annotation, Crowd-coding, Dictionary approaches, and Machine Learning Algorithms”</article-title>
          .
          <source>InC:ommunication Methods and Measures 15.2</source>
          (
          <issue>2021</issue>
          ), pp.
          <fpage>121</fpage>
          -
          <lpage>140</lpage>
          . doi:
          <volume>10</volume>
          .1080/19312458.
          <year>2020</year>
          .
          <volume>1869198</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>L. Van</given-names>
            <surname>Wissen</surname>
          </string-name>
          and
          <string-name>
            <surname>P. Boot. “</surname>
          </string-name>
          <article-title>An Electronic Translation of the LIWC Dictionary into Dutch”</article-title>
          .
          <source>In: Electronic lexicography in the 21st century: Proceedings of eLex 2017 conference. Lexical Computing. Leiden</source>
          , The Netherlands,
          <year>2017</year>
          , pp.
          <fpage>703</fpage>
          -
          <lpage>715</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <surname>A. B. Warriner</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Kuperman</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Brysbaert</surname>
          </string-name>
          . “Norms of Valence, Arousal, and Dominance for
          <volume>13</volume>
          ,915 English Lemmas”.
          <source>In:Behavior research methods 45</source>
          (
          <year>2013</year>
          ), pp.
          <fpage>1191</fpage>
          -
          <lpage>1207</lpage>
          . doi:
          <volume>10</volume>
          .3758/s13428-012-0314-x.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bing</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Lam</surname>
          </string-name>
          .
          <article-title>“A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and Challenges”</article-title>
          .
          <source>InI:EEE Transactions on Knowledge and Data Engineering</source>
          <volume>35</volume>
          .11 (
          <year>2022</year>
          ), pp.
          <fpage>11019</fpage>
          -
          <lpage>11038</lpage>
          . doi: https://doi.ieeecomputersociety.
          <source>org/10</source>
          .1109/TKDE.
          <year>2022</year>
          .
          <volume>3230975</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>