<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>dataset⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Loris Di Quilio</string-name>
          <email>loris.diquilio@studenti.unich.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aspect Based Sentiment Analysis</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Aspect-based Sentiment Analysis, Aspect Category Opinion Sentiment</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DEc, University 'G. d'Annunzio'</institution>
          ,
          <addr-line>Chieti-Pescara</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this work, we report the results of some experiments with Aspect Based Sentiment Analysis (ABSA) on a dataset consisting of user reviews of products of a manufacturing company operating in the beauty industry. We focus on one of the more challenging ABSA tasks, the Aspect Category Opinion Sentiment task, and compare the results obtained by using three diferent tools.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>sentence ”it’s very reasonably priced”, when the subject is not explicitly named, the value
of the aspect term is ”NULL”.
• opinion term (o): is the word, or the words, used by opinion users to convey their
sentiments or feelings about the target entity or aspect. For example, in the sentence
”The pizza is delicious but the service is terrible”, ”delicious” and ”terrible” are opinion
terms, expressing a positive and negative sentiment toward the pizza and the service,
respectively.
• polarity (p): characterizes the sentiment orientation expressed towards an aspect
category or an aspect term. Sentiment polarity can be positive, negative, or neutral indicating
that the sentiment is favorable, unfavorable, or neither, respectively.</p>
      <p>Among the tasks of Aspect-based Sentiment Analysis that aim to predict a single sentiment
element, there are:
• Aspect Term Extraction (ATE);
• Aspect Category Detection (ACD);
• Opinion Term Extraction (OTE);
• Aspect opinion co-extraction (AOCE);
• Aspect Sentiment Classification (ASC).
• Aspect-Opinion Pair Extraction (AOPE);
• End-to-End ABSA (E2E-ABSA);
• Aspect Category Sentiment Analysis (ACSA);
• Aspect Sentiment Triplet Extraction (ASTE);
• Aspect Category Sentiment Detection (ACSD);
• Aspect Category Opinion Sentiment (ACOS).</p>
      <sec id="sec-2-1">
        <title>The tasks where multiple sentiment elements are predicted include: Following we show a summary of the tasks using the input sentence: “The pizza is delicious but the service is terrible”.</title>
      </sec>
      <sec id="sec-2-2">
        <title>Task ATE ACD OTE</title>
        <p>ASC</p>
      </sec>
      <sec id="sec-2-3">
        <title>AOPE</title>
      </sec>
      <sec id="sec-2-4">
        <title>E2E ABSA</title>
      </sec>
      <sec id="sec-2-5">
        <title>ACSA</title>
      </sec>
      <sec id="sec-2-6">
        <title>ASTE</title>
      </sec>
      <sec id="sec-2-7">
        <title>ACSD</title>
        <p>Input
sentence
sentence
sentence
sentence, pizza
sentence, service
sentence
sentence
sentence
sentence
sentence
sentence</p>
        <p>Output
pizza (a), service (a)
food (c), service(c)
delicious (o), terrible (o)
positive(p)
negative (p)
{pizza (a), delicious (o)}, {service (a), terrible (o)}
{pizza (a), positive p)}, {service (a), negative (p)}
{food (c), positive (p)}, {service (c), negative (p)}
{pizza (a), positive (p), delicious (o)},
{service (a), negative (p), terrible (o)}</p>
        <p>{food (c), pizza (a), positive (p)},
{service (c), service (a), negative (p)}
{pizza (a), food (c), delicious (o), positive (p)},
{service (a), service (c), terrible (o), negative (p)}</p>
        <p>In this paper, we will focus our attention on the ACOS task which aims at predicting all
the sentiment information at once, namely category (c), aspect term (a), opinion term (o), and
polarity (p). For the ACOS task, a relatively limited body of research and literature exists. Our
primary objective is to establish an integrated framework that leverages multiple tools for
eficient ACOS task execution.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Dataset and annotation</title>
      <p>Concerning the annotations, there are diferences with the datasets available in the literature
due to the many implicit aspects referring to packaging and opinion terms often composed
from multiple words.</p>
      <p>The dataset is composed of a training and test set, and each sentence can have multiple
annotations.</p>
      <sec id="sec-3-1">
        <title>Sentences</title>
      </sec>
      <sec id="sec-3-2">
        <title>Annotations</title>
      </sec>
      <sec id="sec-3-3">
        <title>Train</title>
        <p>623
881</p>
      </sec>
      <sec id="sec-3-4">
        <title>Test</title>
        <p>133
157</p>
      </sec>
      <sec id="sec-3-5">
        <title>Total</title>
        <p>756
1038</p>
        <p>The composition appears balanced in terms of positive and negative polarity (p), with neutral
sentiment not being of interest. As regards the categories, thirteen classes were identified, mostly
balanced, except for the category belonging to ”general satisfaction of the final consumer”.</p>
        <p>For this work, a custom template in Label Studio was built, which allows all elements of
interest to be annotated for each review. In Figure 1 we show an example of a sentence annotated
with this annotation tool: the explicitly mentioned aspect and opinion elements can be directly
selected in the text, while the polarity and the category, which is not shown, can be chosen from
the predefined ones. A translation module has been developed to convert the JSON encoding of
the dataset exported from Label Studio to other formats.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Experimental evaluation</title>
      <p>
        In this section, we present the details of the experimental evaluation we performed on our
dataset using some tools that have been specifically built for the ACOS task. We have selected
three tools that have stemmed from significant studies in this field and for which the source
code is publicly available online. All the selected tools leverage the fine-tuning of pre-trained
models, specifically T5 [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] and BERT[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], as a crucial component of their functionality:
• Paraphrase modeling [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]: the model’s objective is to generate a sequence of words,
denoted as  , from an input sentence  . The sequence  should contain all the desired
sentiment elements. Once the sequence  is generated, it’s possible to recover the so-called
”sentiment quads”  = (, , , ) . This approach aims to fully leverage the semantics of
the sentiment elements represented by  by generating them in natural language form
within the sequence  . The pre-trained language model used is T5-base. This is the only
tool among those we have considered that does not support implicit opinion terms;
• Extract Classify-ACOS [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]: This tool first performs aspect-opinion co-extraction, then
predicts category-sentiment given the extracted aspect-opinion pairs. The tool uses
the BERT model with AdamW optimizer2 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], so the data is transformed into a format
suitable for it by inserting the token CLS3 at the beginning and at the end of each sentence;
• PyABSA [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]: this tool is a variation of the original one, made for aspect-opinions pair
extraction. There is no documentation about quadruple extraction because this feature
is still experimental. The format of this tool was taken as a reference to transform the
data once exported from the annotation tool. Also in this case T5-base is used as the
pre-trained model.
      </p>
      <sec id="sec-4-1">
        <title>In the table, we show the settings we used for the experiments for each tool.</title>
        <sec id="sec-4-1-1">
          <title>Tool</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Paraphrase modeling</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>Extract Classify-ACOS</title>
          <p>PyABSA
batch-size learning rate</p>
          <p>16 3e-4
32{a, o}, 16(p), 8(c) 2e-5{a, o}, 3e5(p),(c)
16 5e-5
epochs
20
20
20</p>
          <p>The second experiment is motivated by the fact that sentences in our domain often contain
implicit opinions, frequently composed of multiple words rather than single terms. So we
established a relaxed correctness criterion for considering a prediction correct when it matches
the gold standard in terms of aspect, category, and polarity, and when the similarity between
the predicted opinion term and the real one is at least 70%. For computing string similarity we
used the SequenceMatcher4 Python function, which compares pairs of sequences by finding the
longest common subsequence while excluding uninteresting elements, with a quadratic time
2AdamW optimizer: is a stochastic gradient descent method that is based on adaptive estimation of first-order and
second-order moments with an added method to decay weights.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3CLS: this token utilized for BERT stands for classification</title>
      </sec>
      <sec id="sec-4-3">
        <title>4https://docs.python.org/3/library/difflib.html</title>
        <p>complexity for the worst case. In this way, for instance, the prediction of the opinion “super
practical to slip into my bag” can be considered correct even if the real opinion is “practical to
slip into my bag”.
3.1. Results
The performance of the models was evaluated using precision, recall, and F1-Score. Precision is
the ratio of relevant instances retrieved to all instances retrieved. Recall is the ratio of relevant
instances retrieved to all relevant instances. F1-Score is the harmonic mean of precision and
recall. The results are shown in Table 3. A quadruple prediction is deemed correct only if it
matches the gold standard in all four components, except for the last tool in the table.</p>
        <sec id="sec-4-3-1">
          <title>Tool</title>
        </sec>
        <sec id="sec-4-3-2">
          <title>Paraphrase modeling (T5-base)</title>
        </sec>
        <sec id="sec-4-3-3">
          <title>Extract Classify-ACOS (BERT)</title>
        </sec>
        <sec id="sec-4-3-4">
          <title>PyABSA (T5-base)</title>
        </sec>
        <sec id="sec-4-3-5">
          <title>PyABSA (T5-large)</title>
        </sec>
        <sec id="sec-4-3-6">
          <title>PyABSA (T5-large with similarity) Precision</title>
          <p>Among the tools with base pre-trained models (T5-base and BERT), the Paraphrase modeling
tools seem to be the overall best, but the support for the implicit opinion, lacking from this tool,
could be important for some application domains. The Extract Classify-ACOS tool seems to
be slightly better than Paraphrase modeling in terms of precision but has a significantly lower
value for recall. The last tool we considered, PyABSA, is not the best in terms of performance
but it turned out to be very well designed, allowing us to customize it for performing further
experiments using a larger pre-trained model (T5-Large) and employing a similarity criterion
for one of the components. By using the larger model the precision increased from about 32% to
41% using the standard correctness criterion, and to 54% using the relaxed correctness criterion
based on similarity.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion and future work</title>
      <p>We benchmarked three ACOS systems from existing literature on a new domain using our
custom dataset.</p>
      <p>The research aims to create a unified framework for executing various ABSA tasks using
diferent tools on the same dataset. Adapters would handle data translation into the correct
format. The framework should allow defining various experiments and exploring diferent
scenarios through automatic and controlled selection of test and train data, based on data
categories and polarities.</p>
      <p>We foresee an integrated framework where these tools’ predictions are used to automatically
or semi-automatically enhance and expand the training data, thereby improving the eficiency
and overall quality of the sentiment analysis models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lam</surname>
          </string-name>
          ,
          <article-title>A survey on aspect-based sentiment analysis: Tasks, methods, and challenges</article-title>
          ,
          <source>CoRR abs/2203</source>
          .01054 (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2203. 01054.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>M. P.</surname>
          </string-name>
          et al.,
          <article-title>Semeval-2014 task 4: Aspect based sentiment analysis</article-title>
          , in: P. Nakov, T. Zesch (Eds.),
          <source>Proc. 8th International Workshop on Semantic Evaluation, SemEval@COLING</source>
          <year>2014</year>
          , Dublin, Ireland,
          <source>August 23-24</source>
          ,
          <year>2014</year>
          , The Association for Computer Linguistics,
          <year>2014</year>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>35</lpage>
          . doi:
          <volume>10</volume>
          .3115/v1/s14-
          <fpage>2004</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>M. P.</surname>
          </string-name>
          et al.,
          <article-title>Semeval-2016 task 5: Aspect based sentiment analysis</article-title>
          ,
          <source>in: S. B</source>
          . et al. (Ed.),
          <source>Proc. 10th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT</source>
          <year>2016</year>
          , San Diego, CA, USA, June 16-17,
          <year>2016</year>
          , The Association for Computer Linguistics,
          <year>2016</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>30</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/s16-
          <fpage>1002</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>M. M. Trusca</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Frasincar</surname>
          </string-name>
          ,
          <article-title>Survey on aspect detection for aspect-based sentiment analysis</article-title>
          ,
          <source>Artificial Intelligence Review</source>
          <volume>56</volume>
          (
          <year>2023</year>
          )
          <fpage>3797</fpage>
          -
          <lpage>3846</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10462-022-10252-y.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Brauwers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Frasincar</surname>
          </string-name>
          ,
          <article-title>A survey on aspect-based sentiment classification</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <volume>65</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>65</lpage>
          :
          <fpage>37</fpage>
          . doi:
          <volume>10</volume>
          .1145/3503044.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <volume>140</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>140</lpage>
          :
          <fpage>67</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>S. V.</surname>
          </string-name>
          et al.,
          <article-title>Instruction tuning for few-shot aspect-based sentiment analysis</article-title>
          , in: J.
          <string-name>
            <surname>Barnes</surname>
            ,
            <given-names>O. D.</given-names>
          </string-name>
          <string-name>
            <surname>Clercq</surname>
          </string-name>
          , R. Klinger (Eds.),
          <source>Proc. 13th Workshop on Computational Approaches</source>
          to Subjectivity, Sentiment, &amp;
          <article-title>Social Media Analysis</article-title>
          ,
          <source>WASSA@ACL</source>
          <year>2023</year>
          , Toronto, Canada, July
          <volume>14</volume>
          ,
          <year>2023</year>
          , Association for Computational Linguistics,
          <year>2023</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proc</source>
          .
          <year>2019</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</article-title>
          , NAACL-HLT
          <year>2019</year>
          ,
          <string-name>
            <surname>Minneapolis</surname>
            <given-names>MN</given-names>
          </string-name>
          , USA, June 2-7,
          <year>2019</year>
          , Vol
          <volume>1</volume>
          , Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>W. Z.</surname>
          </string-name>
          et al.,
          <article-title>Aspect sentiment quad prediction as paraphrase generation</article-title>
          , in: M. M. et al. (Ed.),
          <source>Proc. of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP</source>
          <year>2021</year>
          ,
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <fpage>7</fpage>
          -
          <issue>11</issue>
          <year>November</year>
          ,
          <year>2021</year>
          , Association for Computational Linguistics,
          <year>2021</year>
          , pp.
          <fpage>9209</fpage>
          -
          <lpage>9219</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .emnlp-main.
          <volume>726</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Aspect-category-opinion-sentiment quadruple extraction with implicit aspects and opinions</article-title>
          , in: C. Z. et al. (Ed.),
          <source>Proc. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021</source>
          , Vol
          <volume>1</volume>
          ,
          <string-name>
            <surname>August</surname>
          </string-name>
          1-
          <issue>6</issue>
          ,
          <year>2021</year>
          , Association for Computational Linguistics,
          <year>2021</year>
          , pp.
          <fpage>340</fpage>
          -
          <lpage>350</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>29</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          , in: Y. Bengio, Y. LeCun (Eds.),
          <source>3rd International Conference on Learning Representations, ICLR</source>
          <year>2015</year>
          , San Diego, CA, USA, May 7-
          <issue>9</issue>
          ,
          <year>2015</year>
          , Conference Track Proceedings,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>A modularized framework for reproducible aspect-based sentiment analysis</article-title>
          ,
          <source>CoRR abs/2208</source>
          .01368 (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2208.01368.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>