<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Sentiment Annotation for Lessing's Plays: Towards a Language Resource for Sentiment Analysis on German Literary Texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christian</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Burghardt, Manuel Computational Humanities, University of Leipzig, Germany https://ch.uni-leipzig.de/</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dennerlein, Katrin University of Würzburg</institution>
          ,
          <addr-line>Germany https://</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Media Informatics Group, University of Regensburg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Schmidt, Thomas Media Informatics Group, University of Regensburg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Schmidt</institution>
          ,
          <addr-line>T., Burghardt, M., Dennerlein, K. and Wolff, C.; licensed under Creative Commons License CC-BY LDK 2019 - Posters Track. Editors: Thierry Declerck and John P. McCrae</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present first results of an ongoing research project on sentiment annotation of historical plays by German playwright G. E. Lessing (1729-1781). For a subset of speeches from six of his most famous plays, we gathered sentiment annotations by two independent annotators for each play. The annotators were nine students from a Master's program of German Literature. Overall, we gathered annotations for 1,183 speeches. We report sentiment distributions and agreement metrics and put the results in the context of current research. A preliminary version of the annotated corpus of speeches is publicly available online and can be used for further investigations, evaluations and computational sentiment analysis approaches. 2012 ACM Subject Classification Document preparation → Annoation; Retrieval tasks and goals → Sentiment analysis</p>
      </abstract>
      <kwd-group>
        <kwd>and phrases Sentiment Annotation</kwd>
        <kwd>Sentiment Analysis</kwd>
        <kwd>Corpus</kwd>
        <kwd>Annotation</kwd>
        <kwd>Annotation Behavior</kwd>
        <kwd>Computational Literary Studies</kwd>
        <kwd>Lessing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Sentiment analysis typically makes use of computational methods to detect and analyze
sentiment (e.g. positive or negative) in written text [7, p.1]. Emotion analysis, a research area
very closely related to sentiment analysis, deals with the analysis of more complex emotion
categories like anger, sadness or joy in written text. More recently, sentiment and emotion
analysis have also gained attention in computational literary studies [15] and are used to
investigate novels [3, 4], fairy tales [
        <xref ref-type="bibr" rid="ref1">1, 8</xref>
        ] and plays [8, 9, 11, 12, 13]. However, there are only
few resources available for sentiment analysis on literary texts, since sentiment annotation of
literary texts has turned out to be a tedious, time-consuming, and generally challenging task
[5, 14, 15], especially for historical literary texts, with their archaic language and complex
plots [14]. Previous studies have shown that agreement levels among annotators of sentiment
annotation are rather low, since literary texts can be understood and interpreted in a number
of ways, i.e. annotations are very much dependent on the subjective understanding of the
annotator [14, 15]. On the technical side, most computational approaches for sentiment
analysis on literary texts currently employ heuristic and rule-based methods [4, 8, 9, 11, 12, 13].
For more advanced machine learning tasks as well as for standardized evaluations,
wellcurated sentiment-annotated corpora for literary texts would be an important desideratum,
but in practice they are largely missing.
      </p>
      <p>We want to close this gap for the genre of German historical plays, more precisely plays
of the playwright G. E. Lessing (1729-1781), by creating a corpus annotated with sentiment
information. Since this is a very tedious task, we first want to investigate the suitability of
different user groups with different levels of knowledge for literary works. We want to find
out how important domain knowledge (experts, semi-experts and non-experts) is for this kind
of task, since the acquisition of domain experts (literary scholars) is obviously more difficult
than the acquisition of annotators without any formal literary training, who could possibly
be hired on a large scale via a crowdsourcing platform. In this article we present first results
from an annotation study with a group of semi-experts.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Design of the Annotation Study</title>
      <p>We conducted the annotation study for semi-experts as part of a course in the Master’s
program of German Literature at the Würzburg University. The course "Sentiment Analysis
vs. Affektlehre" was focused on the plays of Lessing and revolved around the comparison
of more traditional sentiment and emotion analysis approaches with recent computational
methods [11, 12]. Each student had to write an essay and a presentation for a specific play.
In addition, students were asked to participate in the sentiment annotation study. For this
study, we chose six of the most famous plays by Lessing. Every student had to annotate
200 randomly selected speeches except for the students assigned with Damon, since this
play only consists of 183 speeches. A speech is a single utterance of a character, typically
separated by utterances of other characters beforehand and afterwards. A speech can consist
of one or multiple sentences. For each of the six plays we asked two independent students to
provide sentiment annotations. Overall, nine students participated in the annotation project.
One student, who was an advanced tutor, conducted annotations for multiple plays. The
entire corpus comprises 1,183 speeches and we acquired 2,366 annotations (2 annotations
per speech). The corpus consists of 3,738 sentences and 44,192 tokens. A speech consists on
average of 3.16 sentences and 37.36 tokens.</p>
      <p>The overall process was very similar to [14]: Students received an annotation instruction
and guidelines of three pages length, explaining the entire annotation process with examples.
The annotation material was provided via Microsoft Word, which was also used as a basic
annotation tool during the annotation process. Conducting annotation studies with well
known products like Word or Excel is not uncommon in Digital Humanities projects, as
humanities scholars are typically familiar with standard software tools [2, 14]. Every student
received a file with the play-specific speeches. A speech was presented with the following
information: (1) speaker name, (2) text of the speech and (3) position of the speech in the
entire play. Furthermore, the predecessor and successor speech were presented to give some
context information for the interpretation of the actual speech. Sentiment annotations were
documented in a predefined table structure. Annotators were asked to write down the overall
sentiment of an entire speech rather than the sentiment for specific parts of that speech. In
case there were multiple, possibly conflicting sentiment indicators in one speech, annotators
documented the most dominant sentiment for that speech.</p>
      <p>Students were asked to annotate one of the following classes in the first table: negative,
positive, neutral, mixed, uncertain, and other (we refer to this scheme as differentiated
polarity). If annotators did not choose negative or positive, they were asked to provide a
tendency for one of those two classes (binary polarity). Figure 1 illustrates the annotation
process via an example annotation. Here the annotator chose neutral for the differentiated
polarity and negative for the binary polarity.</p>
      <p>We also gathered annotations about the reference of the sentiment and the certainty of the
annotations to get more sophisticated insights about the annotation behavior. However, we
will only focus on the polarity in the upcoming sections. Students had a month to complete
the annotations, but other than that could freely organize the time to complete the task by
themselves. In the next section, we report some of the major findings from our annotation
study.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>First, we present the distributions concerning the differentiated polarity for the entire corpus
and all 2,366 annotations (table 1):</p>
      <p>Negative
734 (31%)</p>
      <p>Positive
442 (19%)</p>
      <p>Neutral
540 (23%)</p>
      <p>Mixed
405 (17%)</p>
      <p>Uncertain
235 (9%)</p>
      <p>Other
10 (0.4%)</p>
      <p>
        Most of the annotations are negative, and a significant number of annotations are mixed
and uncertain. These results are in line with current research about sentiment annotation of
literary texts [
        <xref ref-type="bibr" rid="ref1">1, 14</xref>
        ], which seem to have a general tendency toward more negative annotations.
To analyze the agreement among annotators we used Cohen’s Kappa and the average observed
agreement(AOA, number of agreements divided by the total number of annotated speeches)
as basic metrics. According to [6], Kappa metrics can be interpreted like this (table 2):
Kappa value
&lt;0.20
0.21-0.40
0.41-0.60
0.61-0.80
0.81-1.00
      </p>
      <p>Interpretation
Poor agreement
Fair agreement
Moderate agreement
Substantial agreement</p>
      <p>Very good agreement</p>
      <p>As described above, we had the annotators document differentiated (6 values) as well as
binary (2 values) polarity values for each of the speeches. In addition, we also inferred a third
variant from those annotations, which we call threefold polarity. We basically extend the
binary polarity metric by the value "neutral" in the following manner: if the value "neutral"
is annotated at the differentiated polarity, this value is taken over in this variant. For any
other case the value "positive" or "negative" is taken over from the binary polarity annotation.
We present these metrics per play along with average values for the entire corpus (table 3):</p>
      <p>Play</p>
      <p>Damon
Emilia Galotti</p>
      <p>Der Freigeist
Minna von Barnhelm</p>
      <p>Nathan der Weise
Miß Sara Sampson</p>
      <p>Overall average</p>
      <p>Annotation type
Differentiated polarity</p>
      <p>Threefold polarity</p>
      <p>Binary polarity
Differentiated polarity</p>
      <p>Threefold polarity</p>
      <p>Binary polarity
Differentiated polarity</p>
      <p>Threefold</p>
      <p>Binary polarity
Differentiated polarity</p>
      <p>Threefold polarity</p>
      <p>Binary polarity
Differentiated polarity</p>
      <p>Threefold polarity</p>
      <p>Binary polarity
Differentiated polarity</p>
      <p>Threefold polarity</p>
      <p>Binary polarity
Differentiated polarity</p>
      <p>Threefold polarity</p>
      <p>Binary polarity</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>
        The agreement levels are poor to fair for many annotation types and plays (0.00-0.40).
For annotation types with fewer classes and for several plays we identified Kappa values
with moderate agreement (0.41-0.60). The highest agreement levels are achieved for Emilia
Galotti and Miss Sara Sampson, in the latter case with substantial agreement for the binary
annotation (0.61-0.80). Overall, these results are in line with current research [
        <xref ref-type="bibr" rid="ref1">1, 5, 14</xref>
        ],
proving that sentiment annotation on literary texts is very subjective, difficult and dependent
on the interpretation of the annotator, thus leading to rather low agreement levels compared
to sentiment annotations on other text sorts like movie reviews [17] and social media content
[10]. However, for several plays and annotation types higher agreements were achieved than
compared to similar annotations with non-experts [14]. Therefore, we propose to explore
sentiment annotations with advanced experts concerning the literary text to acquire more
stable annotations. Although the agreement levels are too low to use the full corpus for
evaluation or machine learning purposes, we still think that the corpus can be useful for the
computational literary studies community and therefore provide a first alpha version of the
corpus with all annotations1.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Future Directions</title>
      <p>As a next step, we want to gather more annotations by other user groups and compare them
to each other. Although non-experts currently seem to produce lower levels of agreement, we
want to explore if we can produce valuable corpora with non-experts by using majority decision
with a high number of annotations per speech. To gain insights for possible improvements
concerning the annotation scheme and process, we also conducted a focus group with the
annotators and acquired feedback via a questionnaire. Note that we focused on the overall
sentiment of an entire speech. In the future we want to use more sophisticated models
including more precise annotations of the sentiment direction and object. Taking into account
existing studies on multi-modal sentiment analysis on dramatic texts [16], we currently also
explore the possibility to extend the speech annotation to modalities other than text, e.g.
audio and video material of the corresponding theater plays to improve agreement levels.
1 A preliminary version of this resource is available online:
https://www.dropbox.com/sh/8mu29ny8fhrpgg2/AABFXw7qYHLoJ-4yx8CBlXX9a?dl=0
5
13
14
15
16
17</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>1 2 3 4 Cecilia Ovesdotter Alm and Richard Sproat. Emotional sequencing and development in fairy tales</article-title>
          .
          <source>In International Conference on Affective Computing and Intelligent Interaction</source>
          , pages
          <fpage>668</fpage>
          -
          <lpage>674</lpage>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Lars</given-names>
            <surname>Döhling</surname>
          </string-name>
          and
          <string-name>
            <given-names>Manuel</given-names>
            <surname>Burghardt</surname>
          </string-name>
          .
          <article-title>PaLaFra - Entwicklung einer Annotationsumgebung für ein diachrones Korpus spätlateinischer und altfranzösischer Texte</article-title>
          . In Book of Abstracts,
          <source>DHd</source>
          <year>2017</year>
          , Bern, Switzerland,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Fotis</given-names>
            <surname>Jannidis</surname>
          </string-name>
          , Isabella Reger, Albin Zehe, Martin Becker, Lena Hettinger, and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Hotho</surname>
          </string-name>
          .
          <article-title>Analyzing features for the detection of happy endings in german novels</article-title>
          .
          <source>arXiv preprint arXiv:1611.09028</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Tuomo</given-names>
            <surname>Kakkonen</surname>
          </string-name>
          and
          <article-title>Gordana Galic Kakkonen</article-title>
          . Sentiprofiler:
          <article-title>Creating comparable visual profiles of sentimental content in texts</article-title>
          .
          <source>In Proceedings of the Workshop on Language Technologies for Digital Humanities and Cultural Heritage</source>
          , pages
          <fpage>62</fpage>
          -
          <lpage>69</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Evgeny</given-names>
            <surname>Kim</surname>
          </string-name>
          and
          <string-name>
            <given-names>Roman</given-names>
            <surname>Klinger</surname>
          </string-name>
          .
          <article-title>Who feels what and why? annotation of a literature corpus with semantic roles of emotions</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics</source>
          , pages
          <fpage>1345</fpage>
          -
          <lpage>1359</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Rainer</given-names>
            <surname>Leonhart</surname>
          </string-name>
          . Lehrbuch Statistik:
          <article-title>Einstieg und Vertiefung</article-title>
          .
          <source>Verlag Hans Huber</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Bing</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <article-title>Sentiment analysis: Mining opinions, sentiments, and emotions</article-title>
          . Cambridge University Press,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Saif</given-names>
            <surname>Mohammad</surname>
          </string-name>
          .
          <article-title>From once upon a time to happily ever after: Tracking emotions in novels and fairy tales</article-title>
          .
          <source>In Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities</source>
          , pages
          <fpage>105</fpage>
          -
          <lpage>114</lpage>
          . Association for Computational Linguistics,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Eric T Nalisnick</surname>
          </string-name>
          and
          <article-title>Henry S Baird. Character-to-character sentiment analysis in shakespeare's plays</article-title>
          .
          <source>In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          , volume
          <volume>2</volume>
          , pages
          <fpage>479</fpage>
          -
          <lpage>483</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Rudy</given-names>
            <surname>Prabowo</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mike</given-names>
            <surname>Thelwall</surname>
          </string-name>
          .
          <article-title>Sentiment analysis: A combined approach</article-title>
          .
          <source>Journal of Informetrics</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ):
          <fpage>143</fpage>
          -
          <lpage>157</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Manuel</given-names>
            <surname>Burghardt</surname>
          </string-name>
          .
          <article-title>An evaluation of lexicon-based sentiment analysis techniques for the plays of gotthold ephraim lessing</article-title>
          .
          <source>In Proceedings of the Second Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage</source>
          ,
          <source>Social Sciences, Humanities and Literature</source>
          , pages
          <fpage>139</fpage>
          -
          <lpage>149</lpage>
          . Association for Computational Linguistics,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>URL: http://aclweb.org/anthology/W18-4516.</mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Manuel</given-names>
            <surname>Burghardt</surname>
          </string-name>
          .
          <article-title>Toward a tool for sentiment analysis for german historic plays</article-title>
          . In Michael Piotrowski, editor,
          <source>COMHUM 2018: Book of Abstracts for the Workshop on Computational Methods in the Humanities</source>
          <year>2018</year>
          , pages
          <fpage>46</fpage>
          -
          <lpage>48</lpage>
          , Lausanne, Switzerland,
          <year>June 2018</year>
          .
          <article-title>Laboratoire laussannois d'informatique et statistique textuelle</article-title>
          . URL: https://www.researchgate.net/publication/331907713_Toward_
          <article-title>a_Tool_ for_Sentiment_Analysis_for_German_Historic_Plays.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , Manuel Burghardt, and Katrin Dennerlein. „
          <article-title>Kann man denn auch nicht lachend sehr ernsthaft sein?“ - Zum Einsatz von Sentiment Analyse-Verfahren für die quantitative Untersuchung von Lessings Dramen</article-title>
          . In Georg Vogeler, editor,
          <source>Book of Abstracts, DHd</source>
          <year>2018</year>
          , pages
          <fpage>244</fpage>
          -
          <lpage>249</lpage>
          , Cologne, Germany,
          <year>2018</year>
          . URL: https://www.researchgate.net/publication/331908052_Kann_
          <article-title>man_denn_auch_nicht_ lachend_sehr_ernsthaft_sein-Zum_Einsatz_von_Sentiment_Analyse-Verfahren_fur_ die_quantitative_Untersuchung_von_Lessings_Dramen.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , Manuel Burghardt, and
          <string-name>
            <given-names>Katrin</given-names>
            <surname>Dennerlein</surname>
          </string-name>
          .
          <article-title>Sentiment annotation of historic german plays: An empirical study on annotation behavior</article-title>
          .
          <source>In Sandra Kübler and Heike Zinsmeister</source>
          , editors,
          <source>Proceedings of the Workshop for Annotation in Digital Humantities (annDH)</source>
          , pages
          <fpage>47</fpage>
          -
          <lpage>52</lpage>
          , Sofia, Bulgaria,
          <year>August 2018</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2155</volume>
          / schmidt.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , Manuel Burghardt, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Wolff</surname>
          </string-name>
          .
          <article-title>Herausforderungen für Sentiment Analysis-Verfahren bei literarischen Texten</article-title>
          .
          <source>In Manuel Burghardt and Claudia</source>
          Müller-Birn, editors,
          <source>INF-DH-2018</source>
          , Berlin, Germany,
          <year>September 2018</year>
          .
          <article-title>Gesellschaft für Informatik e</article-title>
          .V.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>URL: https://dl.gi.de/handle/20.500.12116/16996.</mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , Manuel Burghardt, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Wolff</surname>
          </string-name>
          .
          <article-title>Toward multimodal sentiment analysis of historic plays: A case study with text and audio for lessing's emilia galotti</article-title>
          .
          <source>In 4th Conference of the Association of Digital Humanities in the Nordic Countries (DHN</source>
          <year>2019</year>
          ),
          <year>2019</year>
          . URL: https://www.researchgate.net/publication/331907721_Toward_ Multimodal_
          <article-title>Sentiment_Analysis_of_Historic_Plays_A_Case_Study_with_Text_and_ Audio_for_Lessing's_Emilia_Galotti.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Tun</given-names>
            <surname>Thura</surname>
          </string-name>
          <string-name>
            <given-names>Thet</given-names>
            ,
            <surname>Jin-Cheon Na</surname>
          </string-name>
          ,
          <article-title>and Christopher SG Khoo</article-title>
          .
          <article-title>Aspect-based sentiment analysis of movie reviews on discussion boards</article-title>
          .
          <source>Journal of information science</source>
          ,
          <volume>36</volume>
          (
          <issue>6</issue>
          ):
          <fpage>823</fpage>
          -
          <lpage>848</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>