<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>In Search for Lost Emotions: Deep Learning for Opinion Taxonomy Induction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elena Melnikova</string-name>
          <email>elena.melnikova@innoradiant.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emmanuelle Dusserre</string-name>
          <email>emmanuelle.dusserre@eloquant.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muntsa Padró</string-name>
          <email>muntsa.padro@eloquant.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Eloquant</institution>
          ,
          <addr-line>Gières 38610</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Innoradiant</institution>
          ,
          <addr-line>Meylan 38240</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <fpage>57</fpage>
      <lpage>62</lpage>
      <abstract>
        <p>In this article, we present an approach for using word2vec to automatically enrich the opinions' taxonomy used by a sentiment analysis system. More specifically, we worked on emotion lexicon for the field of customer relationship management. The proposed method consists of searching for the nearest distributional neighbors of each source word, and add them to the lexicon of emotions. The hypothesis is that the contextual neighbors of the emotions will also carry an emotional coloring. The results of this experiment show that the neighborhood lexicon is not sufficiently representative. Nevertheless, most of the collected items seem to express another type of opinion, namely judgments. This unexpected result allows us to broaden our taxonomy of opinions with a new informational field, richer and more expressive.</p>
      </abstract>
      <kwd-group>
        <kwd>emotions</kwd>
        <kwd>judgments</kwd>
        <kwd>opinions</kwd>
        <kwd>taxonomy</kwd>
        <kwd>word2vec</kwd>
        <kwd>deep learning</kwd>
        <kwd>sentiment analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The automatic analysis of customer opinions is becoming one of the most pervasive
concerns of companies studying customer reviews. The sentiments or opinions
expressed in these reviews are important indicators for the company's decision-making
strategy. Thus, sentiment analysis (SA) systems need to be reliable and constantly
updated.</p>
      <p>Many SA systems are often based on social networks, tweets or SMS corpora [6, 1,
3]. The analysis of opinions essentially focuses on polarity detection in customers
feedbacks (positive, negative or neutral), [4]. Our SA system for French [7] performs
the extraction of different kind of fine-grained opinions, including emotions, to
extract more detailed information than just positive vs negative polarity. This
finegrained detection is performed using a taxonomy associating words or expressions to
the classes to be detected [11]. Though, building a complete taxonomy can be very
time consuming, and the relevant terms might depend on the domain.</p>
      <p>The present work is motivated by the wish to semi-automatically enrich the
taxonomy, in order to achieve a greater accuracy and easily adapt the system to different
sub-domains. To do so, we propose the use of word2vec [10] to add entries to an
existing taxonomy. This method has been widely and successfully used in semantic
analysis and other NLP tasks [2, 9]. We assume that the matrix constructed by
traversing our domain specific corpus (Customer Relationship Management or CRM)
would locate close to each other emotional words belonging to the same class (from
the distributional point of view). By adopting this method based on deep learning, we
intend to check if it is relevant for the enrichment of our emotions taxonomy.</p>
      <p>The rest of this paper is structured as follows. In Section 2, we describe the
word2vec method and the procedure to enrich the emotion taxonomy. Section 3
reports experiments and result discussion. Finally, Section 4 concludes our work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The word2vec method</title>
      <p>Word2vec is a statistical language model based on neural networks developed by a
team of researchers under the direction of T. Mikolov [10] at Google1. This method is
used to produce word embeddings: words in a corpus are represented as vectors in a
multidimensional space, their position in this space corresponding to their semantic
representation [8, 10].</p>
      <p>The word2vec technique is based on a distributional hypothesis: words that appear
in the same context share semantic values, thus, words that are close in the space are
semantically close. Our goal is to use this method to find sets of closest words (or
related words) and assign them to the same semantic class to enrich other kind of
taxonomies, as it was shown in [2, 5].</p>
      <p>We followed a procedure which contains two major steps. The first one is to create
a word2vec model from a given corpus. The model creation includes the processing
of the corpus (tokenization, lemmatization, etc.) and the induction of the model from
it. The second step is to use the model to calculate the distance between a selected
word (seed word) and the other words in the corpus. In our work, the words already
present in the taxonomy (that already have an assigned semantic tag) serve as seed
words. The objective is to assign to the closest words of the seed words the same
semantic class as them. The distance between two words is calculated with the cosine
of the angle between the vectors that represent them. The more this cosine is close to
1, the closest the neighbor is to the source word.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Experiments and results</title>
      <p>We first focused on the improvement of the emotion detection module, by reviewing
the emotions taxonomy used by our system. This taxonomy contains lexical items
from different linguistic genres2 (literature, psychology, familiar) with 41 classes and
more than 1100 words. This seemed too large and it was not adapted to our CRM
domain. Thus, the taxonomy was considerably reduced to constitute a specific
subtaxonomy (10 classes, 360 words). Among these classes, we find: ANGER, SADNESS,
DISSATISFACTION, LIKING, SATISFACTION, DISLIKE, DOUBT, TRUST, CALMNESS. The
lexicon for this reduced classification seemed quantitatively "poor" and not adequate
to the CRM domain. Thereby, we found it necessary to increase the number of words
1 https://code.google.com/archive/p/word2vec/
2 WordNet based taxonomy :
questions/database/ (Princeton University 2018)
https://wordnet.princeton.edu/wordnet/frequently-askedfor certain classes that contained less than 10 items (SATISFACTION, TRUST,
DISTRUST, DOUBT, DISLIKE, CALMNESS).</p>
      <p>The corpus used for the model has 15 million words and it is very specific to the
CRM domain. Despite word2vec models are expected to work better with bigger
corpus, [5] showed that for specific domains it is preferable to have a domain specific
corpus than huge amount of data.</p>
      <p>Table 1 shows a sample of results obtained when applying the word2vec method
for a selection of words of emotion classes poorly endowed. The headline shows the
seed words with their emotion classes and the following lines, the neighbors proposed
with their cosine.
This extract is obtained by using a threshold of 0.3, meaning that only words with a
cosine bigger than 0.3 are suggested as candidates to be added to the taxonomy. The
threshold is set heuristically, looking for a compromise where we retrieve enough
candidates without too much noise. In the obtained results we noticed different types
of distributional neighborhoods:
1. collocations (perdre confiance (lose confidence), abus de confiance (breach of
trust), gagner en (la) confiance (de qqn) (gain someone's trust), satisfait de
l’accueil (satisfied with the reception));
2. synonyms, antonyms and derived forms (satisfaire (to satisfy) for satisfait
(satisfied));
3. neighbors of a different semantic tag (code (code), oeuvre (work), cordialement
(cordially), fixer (to fix), supprimer (to delete), etc.).</p>
      <p>The lexica that we expected to find should come from the second type of
neighborhood, according to our goal and hypothesis. The obtained results show that these
lexica are rather infrequent and, most often, they contain antonyms. It means that we
cannot increase the classes of our taxonomy by using the closest neighbors. Thus, our
hypothesis of enriching the emotions taxonomy, and especially poorly endowed
emotions classes by using word2vec is not confirmed. Nevertheless, the idea of resorting
to a new method which provides a rich lexical and statistical information, seems to us
very attractive and less exploited.
3.1</p>
      <p>The unexpected result
To better understand the results, we study more closely the lexica coming from the
most numerous type of neighborhood, the neighbors with different semantic tag.
Some of those neighbors can be considered just noise, but we have distinguished in
this lexical layer a category of words that could characterize non-emotional opinions,
the judgments. For example, the words like beau (beautiful) and accord (agreement)
are positive polarity judgments. The words proche (close adj.), rapide (fast) are
positive or negative depending on the domain3. This type of polar judgment lexicon is as
important as the emotion lexicon for the detection and analysis of customers opinions.</p>
      <p>At the sight of these results, we extended the experiment by using word2vec with
the whole taxonomy of emotions as seed words. This allowed us to identify more
words that are likely to express judgments (see Table 2).
The judgments lexica extracted from CRM corpus better characterizes the
CRMspecific classes (see Fig. 1). This idea is corroborated by a recent research study [5].
The augmented judgment taxonomy contains 18 classes and 261 items compared to
13 classes and 197 words from the initial taxonomy of judgments. Table 3 shows an
extract from this ranking that should be expanded and completed.</p>
      <p>3 Their polarity reveals in context (le personnel est proche du client (the staff is close to the client)
[positive judgment in commercial context vs negative judgment in familiar context]).</p>
      <p>4 In this second experiment we use 0.2 as threshold in order to obtain more candidates to be added to the
judgment taxonomy
In this work, we have tested the applicability of word2vec to enrich an existent
emotion taxonomy by finding semantically close words. The obtained results are not very
satisfactory, since very few new emotional words are extracted. Nevertheless, the
method allowed us to extract judgment words which are also very important for the
SA system and much more frequent in our domain. Thus, we can use this new
taxonomy as a resource in our system and we conclude that word2vec is a useful method to
enrich existing taxonomies and even to discover new classes. But to guarantee the
quality of final resources human intervention is needed. Table 4 summarizes the most
important positive and negative points we spotted with our experiments with
word2vec.
11.</p>
      <p>In the future, we plan to perform an extrinsic evaluation of the developed taxonomies.
The convenience of using word2vec to enrich the taxonomies seems clear to us, when
considering as an alternative the manual development of the taxonomies.
Nevertheless, a final evaluation of our SA system before and after enriching the taxonomy
needs to be done. For that, we are currently working on a CRM gold-standard. Also,
we plan to perform the same experiments with word2vec trained on another CRM
corpus to adapt the taxonomy of its sub-domain.</p>
      <p>Furthermore, the taxonomy enriched with word2vec can also serve as input for a
new run of the taxonomy enrichment system. Thus, we could perform a bootstrap to
iteratively enrich the taxonomy of emotions and judgments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Abdaoui</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nzali</surname>
          </string-name>
          , M.D.T.,
          <string-name>
            <surname>Azé</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bringay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavergne</surname>
          </string-name>
          , Ch., et al.
          <article-title>ADVANSE: Sentiment, Opinion and Emotion Analysis in French Tweets</article-title>
          . DEFT: Défi Fouille de Texte,
          <year>Jun 2015</year>
          , Caen, France. Actes de la 11e Défi Fouille de Texte (
          <year>2015</year>
          ) &lt;
          <fpage>hal</fpage>
          -
          <lpage>01222629</lpage>
          &gt; Baroni,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Dinu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            &amp;
            <surname>Kruszewski</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Don't count, predict! A systematic comparison of context-counting vs. contextpredicting semantic vectors</article-title>
          .
          <source>In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pp.
          <fpage>238</fpage>
          -
          <lpage>247</lpage>
          , Baltimore, Maryland, June. Association for Computational Linguistics (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Dini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bittar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Segond</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montaner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>SOMA:</surname>
          </string-name>
          <article-title>The Smart Social CRM</article-title>
          .
          <article-title>Handling Semantic Variability of Emotion Analysis with Hybrid Technologies</article-title>
          .
          <source>In Sentiment Analysis in Social Network Elsevier</source>
          , chapter
          <volume>13</volume>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Dridi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reforgiato Recupero</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Leveraging semantics for sentiment polarity detection in social media</article-title>
          .
          <source>International Journal of Machine Learning and Cybernetics</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Dusserre</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Padró</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Bigger does not mean better ! We prefer specificity</article-title>
          .
          <source>In 12th International Conference on Computational Semantics (IWCS)</source>
          .
          <volume>19</volume>
          -
          <issue>22</issue>
          <year>September 2017</year>
          Montpellier (France) (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Hamon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fraisse</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paroubek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C..</given-names>
          </string-name>
          <article-title>Analyse des émotions, sentiments et opinions exprimés dans les tweets: présentation et résultats de l'édition 2015 du défi fouille de texte (DEFT)</article-title>
          .
          <source>In 22ème Traitement Automatique des Langues Naturelles</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Maurel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Curtoni</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Dini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>A hybrid method for sentiment analysis</article-title>
          .
          <source>In INFORSID. Présenté à Défi Fouille de Texte 2007 (DEFT'07)</source>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Dagan</surname>
            ,
            <given-names>I:</given-names>
          </string-name>
          <article-title>Improving Distributional Similarity with Lessons Learned from Word Embeddings. Transactions of the Association for Computational Linguistics (</article-title>
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Maître</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiron</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouju</surname>
            ,
            <given-names>A .</given-names>
          </string-name>
          <article-title>Utilisation conjointe LDA et Word2Vec dans un contexte d'investigation numérique</article-title>
          .
          <source>Extraction et Gestion des Connaissances</source>
          <year>2017</year>
          ,
          <year>Jan 2017</year>
          , Grenoble, France (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Mikolov</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Distributed Representations of Words and Phrases and their Compositionality</article-title>
          .
          <source>NIPS'13 Proceedings of the 26th International Conference on Neural Information Processing Systems</source>
          . Lake Tahoe,
          <string-name>
            <surname>Nevada</surname>
          </string-name>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Whitelaw</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garg</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Using Appraisal Taxonomies for Sentiment Analysis</article-title>
          .
          <source>Conference: The Second Midwest Computational Linguistic Colloquium</source>
          ,
          <string-name>
            <surname>MCLC</surname>
          </string-name>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>