<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Aspect-based Sentiment Analysis: X2Check at ABSITA 2018</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emanuele Di Rosa Chief Technology Officer App</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Check s.r.l. emanuele.dirosa @app</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>check.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Alberto Durante Research Scientist App2Check s.r.l. alberto.durante @app2check.com</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2001</year>
      </pub-date>
      <volume>13</volume>
      <abstract>
        <p>English. In this paper we describe and present the results of the two systems, called here X2C-A and X2C-B, that we specifically developed and submitted for our participation to ABSITA 2018, for the Aspect Category Detection (ACD) and Aspect Category Polarity (ACP) tasks. The results show that X2C-A is top ranker in the official results of the ACD task, at a distance of just 0.0073 from the best system; moreover, its post deadline improved version, called X2C-A-s, scores first in the official ACD results. About the ACP results, our X2C-A-s system, which takes advantage of our ready-to-use industrial Sentiment API, scores at a distance of just 0.0577 from the best system, even though it has not been specifically trained on the training set of the evaluation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. In questo articolo
descriviamo e presentiamo i risultati dei due
sistemi, chiamati qui X2C-A e X2C-B, che
abbiamo specificatamente sviluppato per
partecipare ad ABSITA 2018, per i task
Aspect Category Detection (ACD) e
Aspect Category Polarity (ACP). I risultati
mostrano che X2C-A si posiziona ad una
distanza di soli 0.0073 dal miglior
sistema del task ACD; inoltre, la sua versione
migliorata, chiamata X2C-A-s, realizzata
successivamente alla scadenza, mostra un
punteggio che lo posiziona al primo posto
nella classifica ufficiale del task ACD.
Riguardo al task ACP, il sistema
X2CA-s che utilizza il nostro standard
Sentiment API, consente di ottenere un
punteggio che dista solo 0.0577 dal miglior
sistema, nonostante il classificatore di
sentiment non sia stato specificamente
addestrato sul training set della evaluation.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        The traditional task of sentiment analysis is the
classification of a sentence according to the
positive, negative, or neutral classes. However, such
task in this simple version is not enough to detect
when a sentence contains a mixed sentiment, in
which a positive sentiment is referred to one
aspect and a negative sentiment to another aspect.
Aspect-based sentiment analysis is focused on the
sentiment classification (negative, neutral,
positive) for a given aspect/category in a sentence.
In nowadays world, reviews became an important
tool widely used by consumers to evaluate
services and products. Given the large amount of
reviews available online, systems allowing to
automatically classify reviews according to
different categories, and assign a sentiment to each of
those categories, are gaining more and more
interest in the market. The former task is called Aspect
Category Detection (ACD) since detects whether
a review speaks about one of the categories
under evaluation; the latter task, called Aspect
Category Polarity (ACP) tries to assign a sentiment
independently for each aspect. In this paper, we
present X2C-A and X2C-B, two different
implementations for dealing with the ACD and ACP
tasks, specifically developed for the ABSITA
evaluation
        <xref ref-type="bibr" rid="ref1">(Basile et al., 2018)</xref>
        . In particular, we
describe the models used to participate to the ACD
competition together with some post deadline
results, in which we had the opportunity to improve
our ACD results and evaluate our systems also on
the ACP task. The resuls show that our X2C-A
system is top ranking in the official ACD
competition and scores first, in its X2C-A-s version.
Moreover, by testing our ACD models on the ACP
tasks, with the help of our standard X2Check
sentiment API, the X2C-A-s system scores fifth at a
distance of just 0.057 from the best system, even
if the other systems have a sentiment classifier
specifically trained on the training set of the
competition. This paper is structured as follow: after
the introduction we present the descriptions of our
two systems submitted to ABSITA and the results
on the development set; then we show and discuss
the results on the official testset of the
competition for both ACD and ACP, finally we provide
our conclusions.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Systems description</title>
      <p>
        The official training dataset has been split into our
internal training set (80% of the documents) and
development set (the remaining 20%). We
randomly sampled the examples for each category,
thus obtaining different sets for training/test set,
by keeping the per category distribution of the
samples through the three sets. We submitted
two runs, as the results of the two different
systems we developed for each category, called
X2CA and X2C-B. The former has been developed
on top of the Scikit-learn library in Python
language
        <xref ref-type="bibr" rid="ref3">(Pedregosa et al., 2011)</xref>
        , and the latter on
top of the WEKA library
        <xref ref-type="bibr" rid="ref2">(Frank et al., 2016)</xref>
        in
JAVA language. In both cases, the input text has
been cleaned with a typical NLP pipeline,
involving punctuation, numbers and stopwords removal.
The two systems have been developed separately,
but the best algorithms obtained by both the model
selections are different implementations of
Support Vector Machine. More details in the
following sections.
2.1
      </p>
      <p>X2C-A
The X2C-A system has been created by
applying an NLP pipeline including a vectorization of
the collection of reviews to a matrix of token
counts of the bi-grams; then, the count matrix has
been transformed to a normalized tf-idf
representation (term-frequency times inverse
documentfrequency). As machine learning algorithm, an
implementation of the Support Vector Machine
has been used, specifically the LinearSVC. Such
algorithm has been selected as the best performer
on such dataset compared to other common
implementations available in the sklearn library.</p>
      <p>Table 1 shows the F1 score on the positive
label in the development set for each category,
where the average value on all of the categories
is 84.92%. X2C-A shows the lowest performance
on the Value category, while shows the best
performance on Location, and high score on Wifi and
Staff.
2.2</p>
      <p>X2C-B
In the model selection process, the two best
algorithms have been Naive Bayes and SMO. We built
a model with both algorithms for each category.
We took into account the F1 score on the
positive labels in order to select the best algorithm.
In this implementation, SMO (Sequential Minimal
Optimization) (Platt, 1998) (Keerthi et al., 2001)
(Hastie et al., 1998) has been the best performing
algorithm on all of the categories, and showed an
average F1 score across all categories of 85.08%.
Its scores are reported in Table 1, where we also
compare its performance with the X2C-A one on
the development set.</p>
      <p>The two systems are built on different
implementation of support vector machines, as
previously pointed out, and differ on the features
extraction process. In fact, X2C-B takes into
account a vocabulary of the 1000 most mentioned
words in the training set, according to the size
limit parameter available in the
StringToWordVector Weka function. Moreover, it uses unigrams
instead of the bi-grams extraction performed in
X2C-A. The two systems reach similar results, i.e.
high scores on Location, Wifi and Staff, and low
scores on the Value category. However, the overall
weighted performance is very close, around 85%
of F1 on the positive labels, and since for some
categories is better X2C-A and for others X2C-B,
we decided to submit both implementations, in
order to understand which is the best one on the test
set of the ABSITA evaluation.</p>
      <sec id="sec-3-1">
        <title>Category</title>
        <p>Cleanliness
Comfort
Amenities
Staff
Value
Wifi
Location
3.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results on the ABSITA testset</title>
      <sec id="sec-4-1">
        <title>Aspect Category Detection</title>
        <p>is about Location and Staff, but since only a part
of it is about Location, the location model of
this category would receive a document
containing ”noise” from its point of view. In the post
deadline runs, we reduce the ”noise” by splitting
this example review in The sight is beautiful which
is only about Location, and but the staff is rude
which is only about Staff. As we can see in
Table 2, the performance of X2C-A increased
significantly and reached a performance score that
is better even than the first classified. However,
the performance of X2C-B slighted decreased in
its X2C-B-s version. This means that the model
of this latter system is not helped by this kind of
”noise” removal technique. This last result shows
that such approach does not have a general
applicability but it depends on the model; however, it
shows to work very well on X2C-A.</p>
        <p>In order to identify the categories where we
perform better, we calculated the score of our systems
on each category1, as shown in Table 3 and Table
4. In Table 3 X2C-A is the best of our systems
on all the categories except Cleanliness and Wifi,
where X2C-B has reached the higher score. In
Table 4, X2C-A-s shows the best performance on all
of the categories. By comparing the results across
1To obtain these scores, we modified the ABSITA
evaluation script so that only one category is taken into account.
In Table 5 we show the results of the Aspect-based
Category Polarity task to which X2Check did not
formally participate. In fact, after the evaluation
deadline we had time to work on the ACP task.</p>
        <p>In order to deal with the ACP task, we decided
to take advantage of our ready-to-use, standard</p>
        <p>
          X2Check sentiment API
          <xref ref-type="bibr" rid="ref5">(Di Rosa and Durante,
2017)</xref>
          . In fact, since we do have an industrial
perspective, we realized that in a real world setting,
the fact of training an Aspect-based sentiment
system through a specific training set has a high
effort associated and cannot have a general purpose
application. In fact, a very common case is the
one in which new categories to predict have to be
quickly added into the system. In this setting, a
high effort activity of labeling examples for the
training set would be required. Moreover,
labeling a review according to the aspects mentioned
and additionally assign a sentiment to each aspect
requires a higher human effort than just labeling
the category. For this reason, we decided to not
specifically train a sentiment predictor specialized
on the given categories/aspects in the evaluation.
Thus, we performed an experimental evaluation in
which after the prediction of the category in the
review, our standard X2Check sentiment API has
been called to predict the sentiment. Since we are
aware that a review may, in general, speak about
multiple aspects and having different sentiment
associated, we decided to apply the X2C-A-s and
X2C-B-s versions which use the splitting method
described in section 3.1. More specifically:
1. each review document has been split into
sentences
2. both the X2Check sentiment API and the
X2C-A/X2C-B category classifiers were run
on each sentence. The former gives as output
the polarity of each sentence; our assumption
is that each portion of the review has a high
probability to have just one sentiment
associated. The latter gives as output all of the
detected categories in each sentence
3. the overall result of a review is given by
the collection of all of the category-sentiment
pairs found in the sentences
        </p>
        <p>
          The results shown in Table 5 show that our
assumption is valid. In fact, despite being a single
sentiment model for all of the categories, we reach
the fifth place in the official ranking with our
X2CA-s system, at a distance of just 0.057 from the
best system specifically trained on such training
set. Furthermore, the ACP performance depends
on the ACD results, in fact the former task
cannot reach a performance higher than the other. For
this reason, we decided to evaluate the sentiment
performance reached on the reviews whose
categories have been correctly predicted. Thus, we
created a score capturing the relationship between
the two results: it is the ratio between the micro
F1 score obtained in the ACP task and the one
obtained in the ACD task. This hand crafted score
shows the quality of the sentiment model, by
removing the influence of the performance on the
ACD task. The overall sentiment score obtained is
88.0% for X2C-B and 87.1% for X2C-A, showing
that even if a specific train has not been made, the
general purpose X2Check sentiment API shows
very good results (recall that, according to
          <xref ref-type="bibr" rid="ref7">(Wilson et al., 2009)</xref>
          humans agree in the sentiment
classification in the 82% of cases).
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Team</title>
        <p>1
2
3
4
X2C-A-s
5
X2C-B-s
6
7
8
baseline</p>
        <p>Tables 6 and 7 show for each category the
micro-F1 and the sentiment score of the ACP task,
calculated like in Table 4, and the relationship
between ACP and ACD scores per category. We can
see that the sentiment model has reached a very
good performance on Cleanliness, Comfort, Staff
and Location since it is close or over the 90%.
However, like noticed for the ACD results, it is
difficult to handle reviews about the Value
category.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper we presented a description of two
different implementations for dealing with the ACD
and ACP tasks at ABSITA 2018. In particular,
we described the models used to participate to the
ACD competition together with some post
deadline results, in which we had the opportunity to
improve our ACD results and evaluate our systems
also on the ACP task. The resuls show that our
X2C-A system is top ranking in the official ACD
competition and scores first, in its X2C-A-s
version. Moreover, by testing our ACD models on the
ACP tasks, with the help of our standard X2Check
sentiment API, the X2C-A-s system scores fifth
at a distance of just 0.057 from the best system,
even if the other systems have a sentiment
classifier specifically trained on the training set of the
competition.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Basile</surname>
          </string-name>
          , Valerio Basile, Danilo Croce and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Polignano</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 Aspect-based Sentiment Analysis task (ABSITA) in Tommaso Caselli</article-title>
          , Nicole Novielli, Viviana Patti, and Paolo Rosso, editors,
          <source>Proceedings of the 6th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA'18), CEUR</article-title>
          .org, Turin
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Eibe</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mark A.</given-names>
            <surname>Hall</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ian</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Witten</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>The WEKA Workbench</article-title>
          .
          <article-title>Online Appendix for ”Data Mining: Practical Machine Learning Tools</article-title>
          and Techniques”, Morgan Kaufmann, Fourth Edition,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Weiss</surname>
          </string-name>
          , R. and
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Cournapeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Duchesnay</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Scikit-learn: Machine Learning in</article-title>
          <source>Python in Journal of Machine Learning Research</source>
          , pp.
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Emanuele</given-names>
            <surname>Di</surname>
          </string-name>
          Rosa and
          <string-name>
            <given-names>Alberto</given-names>
            <surname>Durante</surname>
          </string-name>
          .
          <article-title>LREC 2016 App2Check: a Machine Learning-based system for Sentiment Analysis of App Reviews in Italian Language in Proc</article-title>
          .
          <source>of the 2nd International Workshop on Social Media World Sensors</source>
          , pp.
          <fpage>8</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Emanuele</given-names>
            <surname>Di</surname>
          </string-name>
          Rosa and
          <string-name>
            <given-names>Alberto</given-names>
            <surname>Durante</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <source>Evaluating Industrial and Research Sentiment Analysis Engines on Multiple Sources in Proc. of AI*IA 2017 Advances in Artificial Intelligence - International Conference of the Italian Association for Artificial Intelligence</source>
          , Bari, Italy,
          <source>November 14-17</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>141</fpage>
          -
          <lpage>155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Sophie de Kok</surname>
            , Linda Punt, Rosita van den Puttelaar, Karoliina Ranta, Kim Schouten and
            <given-names>Flavius</given-names>
          </string-name>
          <string-name>
            <surname>Frasincar</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Review-aggregated aspect-based sentiment analysis with ontology features in Prog Artif Intell (</article-title>
          <year>2018</year>
          )
          <volume>7</volume>
          :
          <fpage>295</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Theresa</surname>
            <given-names>Wilson</given-names>
          </string-name>
          , Janyce Wiebe and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Hoffmann</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Recognizing Contextual Polarity: An Exploration of Features for Phrase-Level Sentiment Analysis in Computational Linguistic</article-title>
          , pp.
          <fpage>399</fpage>
          -
          <lpage>433</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Seth</given-names>
            <surname>Grimes</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Expert Analysis: Is Sentiment Analysis an 80% Solution</article-title>
          ? http://www.informationweek.com/software/informationmanagement/expert
          <article-title>-analysis-is-sentiment-</article-title>
          <string-name>
            <surname>analysisan-</surname>
          </string-name>
          80-solution/d/d-id/
          <fpage>1087919</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>