<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ensemble of LSTMs for EVALITA 2018 Aspect-based Sentiment Analysis task (ABSITA) (Short Paper)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mauro Bennici</string-name>
          <email>mauro@youaremyguide.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xileny Seijas Portocarrero</string-name>
          <email>xileny@youaremyguide.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>You Are My GUide</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. In identifying the different emotions present in a review, it is necessary to distinguish the single entities present and the specific semantic relations. The number of reviews needed to have a complete dataset for every single possible option is not predictable. The approach described starts from the possibility to study the aspect and later the polarity and to create an ensemble of the two models to provide a better understanding of the dataset. Italiano. Nell'identificazione delle diverse emozioni presenti in una recensione è necessario distinguere le singole entità presenti e le singole relazioni semantiche. Il numero di recensioni necessarie per avere un dataset completo per ogni singola opzione possibile non è predicibile.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>L'approccio descritto parte dalla
possibilità di creare due modelli diversi, uno per
la parte di categorizzazione, e l'altro per
la parte di polarità. E di unire i due
modelli per ottenere una maggiore
comprensione del dataset.</p>
      <p>
        Automating the correct recognition of the various
problems can lead to the timely addressing of the
same to the persons appointed to solve them.
The research was carried out with the dataset
provided within the task called ABSITA,
Aspectbased Sentiment Analysis at EVALITA 20181
        <xref ref-type="bibr" rid="ref1">(Basile et al., 2018)</xref>
        . The task was a combination
of two tasks, Aspect Category Detection (ACD)
and Aspect Category Polarity (ACP).
      </p>
      <p>The dataset is a selection of hotel reviews taken
in Italian from the portal Booking.com.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Description of the system</title>
      <p>Each review has been cleaned up by special
characters, lemmatized and brought to lowercase
with the SpaCy2 framework.</p>
      <p>Generic Italian texts have been used, instead of
reviews in the accommodation context to be sure
that the model will be suitable for more business
models, to generate vectors in fastText3. The best
one has a dimension of 200, with character
ngrams of length 5, a window of size 5 and 10
negatives.</p>
      <p>
        The system is the ensemble of two different
models to improve the ability to discover hidden
properties
        <xref ref-type="bibr" rid="ref2">(Akhtar et al., 2018)</xref>
        .
      </p>
      <p>The first model is a bi-directional Long
ShortTerm Memory (BI-LSTM).</p>
      <p>This model is used for the discernment of the
ASPECT.
Layer (type) Output Shape Param #
===================================
e (Embedding) (None, 100, 200) 1420400
_______________________________________
b (Bidirection) (None, 512) 935936
_______________________________________
d (Dense) (None, 7) 3591
===================================
A second BI-LSTM model is used for the
discernment of POLARITY.
_______________________________________
Layer (type) Output Shape Param #
===================================
e (Embedding) (None, 100, 200) 1420400
_______________________________________
b (Bidirection) (None, 512) 935936
_______________________________________
d (Dense) (None, 14) 7182
===================================
A dropout and a recurrent_dropout of 0.1.
The optimizer for both is the RMSProp.
The loaded embedding is trainable.</p>
      <p>Both the systems use Keras4 to create the RNN
models.</p>
      <p>The models were trained and tested with a 5-fold
cross-validation with a ratio of 80% training and
20% testing. The best model was automatically
saved at each iteration.</p>
      <p>A threshold of 0.5 was used on the first model to
activate the result of the last layer. In the second
model, the threshold was of 0.43.</p>
      <p>Aspect Category Detection (ACD)</p>
      <p>The results show that the models are useful to
understand the category of a review better than
its polarity.</p>
      <p>
        After that we ensemble the two models
        <xref ref-type="bibr" rid="ref3">(Choi et
al., 2018)</xref>
        to obtain a system able to overcome
the results of every single model in the ACP task
reducing the result on the ACD task (table 3).
The ensemble has been created in cascade
making sure that a system acts as Attention to the
underlying system.
      </p>
      <p>The threshold of activation was a range between
0.45 and 0.55.</p>
      <p>
        A third model, a LightGBM5
        <xref ref-type="bibr" rid="ref3">(Bennici and
Portocarrero, 2018)</xref>
        was also tested, where the
following properties are extracted from the reviews
text:
•
•
•
•
•
•
•
•
•
•
length of the review
percentage of special characters
the number of exclamation points
the number of question marks
the number of words
the number of characters
the number of spaces
the number of stop words
the ratio between words and stop words
the ratio between words and spaces
and they are joined to the vector created by the
bigram and trigram of the text itself at word and
character level.
      </p>
      <p>The number of leaves is 250, the learner set as
‘Feature’, and a the learning rate at 0.04.
The result of the union between the three models
could not be submitted to the final evaluation,
due to the limit of 2 possible submissions, but
reported results higher than 83% in the tests
carried out after the release of the complete dataset
for ASPECT and 75% for POLARITY.</p>
      <p>Also, the inference is faster than the RNN
models.
https://keras.io
https://github.com/Microsoft/LightGBM</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>In the evaluation phase, we can see how the
results have given reason to the ensemble of the
two results.</p>
      <p>It is clear that the ACP task (table 4) is the
beneficiary of this process, instead of the ACD one
(table 3) that lost more than one point.</p>
      <p>The study of the dataset is influenced by the little
extension of the training dataset and by the
specificity of some terms that could refer to different
categories such as the comfort of the room and
the quality/price ratio.</p>
      <p>Various types of data preparation have also been
used, including the preservation of special
characters, the shape of words (to better identify
cities or places written in capital letters), and some
SMOTE functions to increase the number of
entries but with poor results and noticeable
overfitting.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>Creating an ensemble of models to bring out
various properties of a review gave better results
than using a single model in the polarity
identification.</p>
      <p>The terms used in the review are sometimes
misleading and can be used both positively or
negatively, and to identify different categories of the
hotel.</p>
      <p>In the near future, we are ready to create a
system to split the text of the review to categorize
only a single sentence, or less a single subject or
object. In this way, we will be ready to evaluate
also the polarity of the single object or subject,
and only the terms single related to it to improve
the result of the ACP task.</p>
      <p>The performance of the system will also be
evaluated by replacing all the possible entities
with variables known as:
l
l
l
l
l</p>
      <p>City
Museum
Panoramic Point
Railway station</p>
      <p>Street
and with a pre-category knew a priori as
Breakfast for words like Coffee, Cornetto, and Jam.
The expected result is to reduce the variance of
the dataset, to improve the ACD result, and to be
able to use the system in production.</p>
      <p>Finally, we will evaluate the speed and
effectiveness of a CNN model in which the tasks,
ASPECT, and POLARITY, can be studied
separately and then merged.</p>
      <p>Bennici, M. and Seijas Portocarrero, X. (2018). The
validity of dictionaries over the time in Emoji
prediction. In Tommaso Caselli, Nicole Novielli, Viviana
Patti, and Paolo Rosso, editors, Proceedings of the 6th
evaluation campaign of Natural Language Processing
and Speech tools for Italian (EVALITA’18), Turin,
Italy. CEUR.org.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Polignano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Overview of the EVALITA 2018 Aspectbased Sentiment Analysis task (ABSITA)</article-title>
          .
          <article-title>Proceedings of the 6th evaluation campaign of Natural Language Processing and Speech tools for Italian (EVALITA'18)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Akhtar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghosal</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhattacharyya</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kurohashi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2018</year>
          ,
          <article-title>October 15). A Multi-task Ensemble Framework for Emotion, Sentiment</article-title>
          and
          <string-name>
            <given-names>Intensity</given-names>
            <surname>Prediction</surname>
          </string-name>
          . Retrieved from https://arxiv.org/abs/
          <year>1808</year>
          .01216
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>J. Y.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Bumshik</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>“Combining LSTM Network Ensemble via Adaptive Weighting for Improved Time Series Forecasting</article-title>
          ,” Mathematical Problems in Engineering, vol.
          <year>2018</year>
          ,
          <string-name>
            <surname>Article</surname>
            <given-names>ID</given-names>
          </string-name>
          2470171, 8 pages. doi: https://doi.org/10.1155/
          <year>2018</year>
          /2470171.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>