<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Native sentiment analysis tools vs. translation services - Comparing GerVADER and VADER</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Karsten Michael Tymann</string-name>
          <email>ktymann@fh-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Louis Steinkamp</string-name>
          <email>louis.steinkamp@fh-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oxana Zhurakovskaya</string-name>
          <email>oxana.zhurakovskaya@fh-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carsten Gips</string-name>
          <email>carsten.gips@fh-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>FH Bielefeld University of Applied Sciences</institution>
          ,
          <addr-line>Minden</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>VADER is a rule-based sentiment analysis tool for English texts with a social media focus. GerVADER is a German adaptation of VADER, which was developed following the steps of VADER's development process. VADER showed high F1 scores especially for the social media domain, whereas the German adaptation achieved much lower results within the same domain, although on other test data. In this work we examine the question of whether these di erences are language-speci c. Therefore we apply an improved version of GerVADER to German texts and compare the results with the application of VADER to the same texts that are automatically translated into English. The benchmarking showed, that the translation combined with VADER achieves up to 5% higher F1 scores in all test cases, which can be explained by the translation tools automatic xing of awed sentences. However, native language tools can still be viable, since it saves time and costs and does not need another dependency to a third party service.</p>
      </abstract>
      <kwd-group>
        <kwd>VADER GerVADER sentiment analysis translation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Sentiment analysis describes the process of automatically rating texts or
sentences with a sentiment value. The sentiment value ranges from negative, to
neutral to positive and can be expressed as a numeric value or a classi cation
in one of the three sentiment categories. Compared to machine learning based
approaches classi cation can also be done by rule-based algorithms which have
the advantage that they do not require any training data. However, developing
a rule-based tool requires linguistic knowledge and is signi cantly more
dependent on the target language. Hence one can not simply transfer one language's
features, e.g. German, to another language, e.g. English. The languages can
differ in their grammar and overall sentence structure, making it very error prone
to simply transfer the algorithm to another language. How can negation be
detected, or what is the meaning of a speci c punctuation? Are there words that
can have di erent meanings in di erent contexts and how can you derive those?
Machine learning approaches simply train on lots of annotated data and extract
the features with techniques like word embedding, but for a rule-based approach
the developer has to design the process of detecting word patterns themselves.
While a rule-based tool does not need bootstrap data, it will still need some form
of lexicon with individual words or phrases to derive sentiments from. These
lexicons are usually created by humans and are often rated by di erent people to
o er an average sentiment value.</p>
      <p>
        In this work, which was part of a student project at Bielefeld University of
Applied Sciences, we will rst discuss some improvements to our rule-based
sentiment analysis tool GerVADER [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Second, we examine the question whether the
e ort to develop adaptations in the native language like GerVADER is
worthwhile at all or whether one could achieve similarly good results by translating
the corpora to be examined into English and using VADER [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] subsequently for
the analysis. In order to investigate on this, we translated our test corpora to
the English language with 3rd party tools and benchmarked the data with the
English tool VADER.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>2.1</p>
      <sec id="sec-2-1">
        <title>VADER &amp; GerVADER</title>
        <p>
          VADER is an abbreviation for \Valence Aware Dictionary and sEntiment
Reasoner" and is free to use. It is a rule based tool for analyzing sentiments for
English sentences. VADER managed to outperform other lexicon-based approaches
as well as machine learning models. Especially in the domain of social media
VADER achieved high F1 scores of up to 96%. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
        </p>
        <p>
          GerVADER is an adaptation of the VADER tool for the German language.
The process of VADER has been replicated in some steps, such as the crowd
rating for the lexicon which is based on the SentiWS lexicon [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], while others
have been simply transferred to the German language, such as the heuristics.
GerVADER is as well free to use. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Benchmarking corpora</title>
        <p>
          SCARE is a corpus consisting of Google Play Store reviews. The reviews are
categorized by their star ratings (1 to 5) and are split into 11 app categories. In
total there are over 800.000 user reviews. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
        </p>
        <p>
          The SB10k corpus consists of German tweets that are humanly labeled into
the three sentiment categories: positive, neutral, negative. It consists of almost
10.000 tweets. Both corpora will be used for benchmarking purposes. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Translation tools</title>
        <p>For translating our test data, we have mostly relied on Googles translation
service. One of the datasets has been additionally translated with MyMemory.</p>
        <p>
          Googles translation service is based on the Google Neural Machine
Translation (GNMT) system. Its hybrid model consists of a Transformer [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] encoder
and RNN decoder. The learning is based on sequence-to-sequence neural network
learning and is a mix between character and word-delimited models. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
        </p>
        <p>MyMemory1 is a large collection of Translation Memories that are collected
and provided by humans and organizations. The translations are saved as words
or sequences in databases which can then be matched by the users input. As of
now there are over 4 billion human contributions.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Process</title>
      <p>The process section is divided into two subsections. In the rst we analyze the
aws of GerVADER and how we improved the algorithm. In the second
subsection we will give insight on the benchmarking itself.
3.1</p>
      <sec id="sec-3-1">
        <title>Flaws in GerVADER</title>
        <p>
          The analysis of the initial version of GerVADER [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] showed three aws that
promised room for improvement. Firstly the negation detection is inaccurate
and can not be converted to the German language by simply translating the
negation keywords (e.g. 'not'). Secondly booster words (e.g. 'super', 'very') are
sometimes the only words with a sentiment meaning in a sentence, but they do
not get noticed, since they simply serve as booster for following words valences.
Thus the sentences receive a neutral rating, although the booster word itself
might carry sentiment meaning (e.g. 'super'). Thirdly misspelled words do not
get noticed, since the words have to be written exactly like in the lexicon. For
every problem case we developed test corpora so that we were able to tell whether
our changes improved the overall rating. Details on the changes can be found on
GitHub2.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>VADER vs. GerVADER</title>
        <p>Both VADER tools cover their own languages. A question however arises whether
it even makes sense to translate a sentiment analysis tool to one's native
language. To investigate on this we translated our test corpora with Google and
MyMemory.</p>
        <p>For MyMemory only the SB10k corpus was used, whereas for Google
Translate we tested multiple corpora (SB10k, SCARE). Additionally we constructed a
SCARE Balanced corpus, consisting of 400 positive, 400 negative and 400
neutral reviews of each SCARE corpus le. Thus a balanced le of all 3 sentiments
is built, consisting of 13.200 entries.
1 MyMemory by translated LABS https://mymemory.translated.net/
2 GitHub - GerVADER https://github.com/KarstenAMF/GerVADER
GerVADER improved in all three areas for our test corpora by adjusting the
original algorithm rules as well as adding new features such as fuzzy-matching.
Thus the overall classi cation score of German texts has increased.</p>
        <p>When comparing VADER with GerVADER Table 1 shows that VADER
outperforms GerVADER of up to 5%. Even with the improvements in GerVADER,
it is still outmatched. This is due to the fact, that translation tools do not just
translate sentences word by word, but also consist of features such as
fuzzymatching and entity or POS (part-of-speech) tagging. Those are features, that
we have partly integrated into GerVADER as well. Googles API is pre-trained
with methods of Deep Learning on millions of data and adjusted to word and
phrase sequences and not a simple word-to-word translation. Therefore it goes
beyond a simple translation mechanism. As a result, spelling errors are corrected
and the overall structure of the sentences is adapted to the desired output
language. Thus it is not just a simple translation but also a text correction, which
may explain the better results for VADER in the benchmark.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion &amp; Future work</title>
      <p>While GerVADER has been improved, looking at the translation comparison
the question arises whether GerVADER serves any purpose. Developing a native
language adopted tool is challenging and has lots of potential for creating new
aws. It requires linguistic knowledge in the target language, but this allows
one to address language speci c characteristics more appropriately than with a
translation. Also, the translation service is an additional dependency, which can
be problematic in factors of costs, time and complexity. With the current version
it might not be worth it to trade GerVADER for VADER for a maximum of 5%
F1 score improvement. However, translation tools handle the linguistic features
for the developer and are therefore an interesting research topic for VADER and
similar tools that are available in several languages.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Cieliebak</surname>
          </string-name>
          , Jan Deriu, Dominic Egger, and Fatih Uzdilli.
          <article-title>\A Twitter Corpus and Benchmark Resources for German Sentiment Analysis."</article-title>
          <source>In: Proceedings of the 4th International Workshop on Natural Language Processing for Social Media (SocialNLP</source>
          <year>2017</year>
          )
          <article-title>"</article-title>
          , Valencia, Spain,
          <year>2017</year>
          .
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W17</fpage>
          -1106.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hutto</surname>
          </string-name>
          and
          <string-name>
            <given-names>Eric</given-names>
            <surname>Gilbert</surname>
          </string-name>
          .
          <article-title>VADER: A Parsimonious Rule-Based Model for Sentiment Analysis of Social Media Text</article-title>
          .
          <year>2014</year>
          . url: https://www.aaai. org/ocs/index.php/ICWSM/ICWSM14/paper/view/8109.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Remus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Quastho</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Heyer</surname>
          </string-name>
          . \
          <article-title>SentiWS - a Publicly Available German-language Resource for Sentiment Analysis."</article-title>
          <source>In: Proceedings of the 7th International Language Ressources and Evaluation (LREC'10)</source>
          , pp.
          <fpage>1168</fpage>
          -
          <lpage>1171</lpage>
          . (
          <year>2010</year>
          ).
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Mario</surname>
            <given-names>S</given-names>
          </string-name>
          anger, Ulf Leser, Ste en Kemmerer,
          <string-name>
            <surname>Peter Adolphs</surname>
          </string-name>
          , and Roman Klinger.
          <article-title>\SCARE - The Sentiment Corpus of App Reviews with Finegrained Annotations in German"</article-title>
          .
          <source>In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ). Portoroz, Slovenia,
          <year>2016</year>
          . isbn:
          <fpage>978</fpage>
          -2-9517408-9-1. url: https://www.aclweb.org/ anthology/L16-1178/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Karsten</given-names>
            <surname>Michael</surname>
          </string-name>
          <string-name>
            <surname>Tymann</surname>
          </string-name>
          , Matthias Lutz, Patrick Palsbroker, and Carsten Gips. \
          <article-title>GerVADER - A German adaptation of the VADER sentiment analysis tool for social media texts."</article-title>
          <source>In: In Proceedings of the Conference "Lernen</source>
          , Wissen, Daten,
          <source>Analysen" (LWDA</source>
          <year>2019</year>
          ), Berlin, Germany,
          <source>September 30 - October 2</source>
          ,
          <year>2019</year>
          .
          <year>2019</year>
          . url: http://ceur- ws.org/Vol-
          <volume>2454</volume>
          / paper_14.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
          <string-name>
            <given-names>Aidan N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Lukasz Kaiser, and Illia Polosukhin. \
          <article-title>Attention Is All You Need"</article-title>
          .
          <source>In: CoRR abs/1706</source>
          .03762 (
          <year>2017</year>
          ). arXiv:
          <volume>1706</volume>
          .03762. url: http: //arxiv.org/abs/1706.03762.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Yonghui</given-names>
            <surname>Wu</surname>
          </string-name>
          , Mike Schuster, Zhifeng Chen,
          <string-name>
            <surname>Quoc</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
            , Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao,
            <given-names>Qin</given-names>
          </string-name>
          <string-name>
            <surname>Gao</surname>
          </string-name>
          , Klaus Macherey, Je Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil,
          <string-name>
            <surname>Wei</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Cli Young,
          <string-name>
            <given-names>Jason</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jason</given-names>
            <surname>Riesa</surname>
          </string-name>
          , Alex Rudnick, Oriol Vinyals, Greg Corrado, Macdu Hughes, and
          <article-title>Je rey Dean. \Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation"</article-title>
          .
          <source>In: CoRR abs/1609</source>
          .08144 (
          <year>2016</year>
          ). arXiv:
          <volume>1609</volume>
          .08144. url: http://arxiv.org/abs/1609.08144.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>