<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Machine Learning approach for Sentiment Analysis for Italian Reviews in Healthcare</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luca Bacco</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Cimino</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Paulon</string-name>
          <email>paulon@webmonks.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mario Merone</string-name>
          <email>m.meroneg@unicampus.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
          <email>felice.dellorlettag@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Istituto di Linguistica Computazionale “Antonio Zampolli” (ILC-CNR), ItaliaNLP Lab</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universita` Campus Bio-Medico</institution>
          ,
          <addr-line>UCBM</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Webmonks s.r.l</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present our approach to the task of binary sentiment classification for Italian reviews in healthcare domain. We first collected a new dataset for such domain. Then, we compared the results obtained by two different systems, one including a Support Vector Machine and one with BERT. For the first one, we linguistic pre-processed the dataset to extract hand-crafted features exploited by the classifier. For the second one, we oversampled the dataset to achieve better results. Our results show that the SVMbased system, without the worry of having to oversample, has better performance than the BERT-based one, achieving an F1-score of 91.21%.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Nowadays, when people want to buy a product
or service, they often rely on online reviews of
other buyers/users (think of online sales giants
like Amazon). Likewise, patients are
increasingly relying on reviews on social media, blogs
and forums to choose a hospital where to be
cured. This behaviour is occurring not only abroad
        <xref ref-type="bibr" rid="ref8 ref8 ref9 ref9">(Greaves et al., 2012; Gao et al., 2012)</xref>
        , but also in
Italy. This is also demonstrated by the
increasing amount of reviews in QSalute1, one of the
most popular Italian ranking websites in
healthcare. These reviews are often ignored by hospital
companies, which do not exploit the potential of
such data to understand patients’ experiences and
consequently improve their services. Due to the
large amount of data, there is a need for automatic
analysis techniques. To meet these needs, we
decided to introduce a sentiment analysis system
based on machine learning techniques, in order
to classify whether a review has positive or
negative sentiment. Since such systems require
annotated data, the first step was to build a brand-new
dataset. We present it in the next section. Then,
we developed two systems based on two
different classifiers described in Section 3 together with
the features extracted from the text. In Sections 4
and 5 we show the experiments conducted during
this study, the obtained results and their
discussion. Finally, the last section provides concluding
remarks and some possible future developments.
While there exist several works on affective
computing in several domains for the Italian language
        <xref ref-type="bibr" rid="ref1 ref2 ref2 ref3 ref3 ref4 ref4 ref6 ref6">(Basile et al., 2018; Cignarella et al., 2018;
Barbieri et al., 2016)</xref>
        , at the time we are writing there
are no references in literature that address this
particular domain in Italian. Thus, for the best of our
knowledge, this is the first work of sentiment
analysis on Italian reviews in healthcare.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Dataset</title>
      <p>QSalute is an Italian portal where users share their
experiences about hospitals, nursing homes and
doctors. We have collected a total of 47,224
documents (i.e. reviews). Each document consists
of the free text of the review and other metadata
such as the document id, the disease area to which
the document belongs and the title. In addition,
among the provided metadata there is the average
grade, i.e. the mean over the votes in four
categories: Competence, Assistance, Cleaning and
Services.</p>
      <p>In this work, documents with an average grade
less than or equal to 2 were assigned to the
negative class (-1), while documents with an average
grade greater than or equal to 4 were assigned
to the positive class (1). The remaining
documents were labelled with the neutral class (0). The
dataset is strongly unbalanced towards the
positive class: 40641 reviews for the positive class,
3898 for the neutral class and 2685 for the negative
class. However, in this work, neutral reviews were
discarded thus resulting in a dataset composed by
43326 reviews. The following analyses are then
referred to this subset: in Table 1 we report some
features of the dataset for each site (i.e. the
disease area), while the distribution of tokens over
their length is reported in Figure 1.</p>
      <p>Site
Nervous System</p>
      <p>Hearth
Haematology
Endocrinology</p>
      <p>Endoscopy</p>
      <p>Facial</p>
      <p>Genital
Gynaecology</p>
      <p>Infections
Ophthalmology</p>
      <p>Oncology
Otorhinology</p>
      <p>Skin
Plastic Surgery
Pneumology
Rheumatology</p>
      <p>
        Senology
Thoracic Surgery
Vascular Surgery
In order to build the first system, we followed the
approach proposed by
        <xref ref-type="bibr" rid="ref10">(Mohammad et al., 2013)</xref>
        for the sentiment analysis of English tweets and
we adapted it for Italian reviews in healthcare.
More precisely, we implemented a Support
Vector Machine (SVM) classifier with linear kernel,
in terms of liblinear
        <xref ref-type="bibr" rid="ref7">(Fan et al., 2008)</xref>
        rather than
libsvm in order to scale better to large numbers of
samples, as also reported in the documentation2 of
the LinearSVC model.
      </p>
      <p>Firstly, all documents pass through a
preprocessing pipeline, consisting of a sentence
splitter, a tokenizer and a Part-Of-Speech (POS)
tagger (all of these tools have been previously
developed by the ItaliaNLP3 laboratory). Then,
documents pass through a step of feature extraction,
illustrated in the next section.</p>
      <sec id="sec-2-1">
        <title>3.1.1 Feature Extraction</title>
        <p>
          All features were chosen due to their effectiveness
shown in several tasks for sentiment classification
for Italian
          <xref ref-type="bibr" rid="ref5">(Cimino and Dell’Orletta, 2016)</xref>
          . We
refer to these features under the name of
handcrafted features and embedding features.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Raw and Lexical Text Features</title>
        <p>(Uncased) Word n-grams: presence or
absence of contiguous sequences of n tokens in
the document text, with n=f1, 2, 3g.
Lemma n-grams: presence or absence of
contiguous sequences of n lemmas occurring
in the document text, with n=f1, 2, 3g.
Character n-grams: presence or absence of
contiguous sequences of n characters
occurring in the document text, with n=f2, 3, 4,
5g.</p>
        <p>Number of tokens: total number of tokens
of the document.</p>
        <p>Number of sentences: total number of
sentences of the document.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Morpho-syntactic Features</title>
      </sec>
      <sec id="sec-2-4">
        <title>Coarse-grained Part-Of-Speech n-grams:</title>
        <p>presence or absence of contiguous sequences
of n grammatical categories, with n=f1, 2,
3g.</p>
        <p>Fine-grained Part-Of-Speech n-grams:
presence or absence of contiguous sequences
of n (fine-grained) grammatical categories,
with n=f1, 2, 3g.</p>
        <p>2www.scikit-learn.org/stable/modules/
generated/sklearn.svm.LinearSVC.html
3www.italianlp.it</p>
        <p>
          Word Embeddings Combination: this feature
is composed of three vectors. Each vector was
calculated by the mean over word embeddings
belonging to a specific fine-grained grammatical
category: adjectives (excluding possessive
adjectives), nouns (excluding abbreviations), and verbs
(excluding modal and auxiliary verbs). Word
embeddings used in this work are vectors of 128
dimensions, and they were extracted from a corpus
of more than 46 million tweets. Such embeddings
were already used in
          <xref ref-type="bibr" rid="ref2 ref3 ref4 ref6">(Cimino et al., 2018)</xref>
          and
they are available for download at the website of
ItaliaNLP4. Furthermore, three features have been
added to indicate the absence of word embeddings
belonging to such categories, for a total of 387
(128 3 + 3) features.
3.2
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>System 2 based on BERT</title>
        <p>
          We also implemented Bidirectional Encoder
Representations from Transformers, or as better
known, BERT, to classify the sentiment of the
reviews. BERT is a pre-trained language model
developed by
          <xref ref-type="bibr" rid="ref2 ref3 ref4 ref6">(Devlin et al., 2018)</xref>
          at Google AI
Language. Pre-trained BERT (available at its GitHub
4www.italianlp.it/resources/italianword-embeddings
page5) may be fine-tuned on a specific NLP task in
a specific domain, such as the sentiment analysis
for reviews in the healthcare domain. To do that,
the original text must be tokenized with its own
tokenizer.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>We conducted two types of experiments. In the
first one, we wanted to evaluate which of the
systems was the best. For each configuration, we have
trained and tested the system using a stratified
kfold cross-validation (with k = 5). In the second
part, we wanted to evaluate the robustness of the
best system in a context out-domain, dividing the
folders by disease sites. The software has been
entirely developed in Python.
4.1</p>
      <sec id="sec-3-1">
        <title>System 1</title>
        <p>We tested three different configurations of our
SVM-based system, depending on the sets of
features used in the experiment: only hand-crafted
features (more than 626 thousands features), only
embeddings (387 features), and a combination of
both. For such experiments, the features that have
shown to not bring improvements to the
performance (numbers of tokens and sentences), or even
5www.github.com/google-research/bert
to lower it (fg-POS n-grams, Lemmas n-grams
with n=f2, 3g) during a preliminary
experimental phase were excluded from the hand-crafted
features set. Thus, it turns out that such set is
composed only of Uncased Word and cg-POS n-grams
with n=f1, 2, 3g, and Lemmas. In order to reduce
the dimensionality of the set, but also to improve
the performance of our system, the features pass
through a step of filtering after being extracted for
the training set. Each feature that appears less than
a certain threshold th within the training set can
be assumed to be not so relevant and is therefore
discarded. Such threshold has been set equal to 1
(th=1) after a search of the optimal value during
the preliminary experimental phase.
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>System 2</title>
        <p>The experiments with BERT were conducted
using the same partition into the 5 folds used
during the experiments with the SVM-based
classifier. This division allowed us to compare the
results achieved by the two classifiers. The BERT
model used in our experiments is the multilingual
cased pre-trained one.</p>
        <p>We tested two different approaches. These
experiments have followed two pipelines. In the first
one, the model was fine-tuned with folds from the
original dataset described in section 2. In the
second one, each fold was obtained by oversampling
the minority class (i.e. the negative one) in the
original fold. The oversampling was obtained by
multiplying each negative sample in the fold by 4.
These results in the ratio of negative to positive
samples being increased from about 1:16 to about
1:4. Other experiments were conducted further
increasing the ratio to about 1:2, but this has not led
to significant improvements in performance at the
expense of computational time. For both the
approaches, the model was fine-tuned for 5 epochs
on a 12 GB NVIDIA GPU with Cuda 9.0 with the
following hyperparameters:
maximum sequence length of 128 tokens (it
seems reasonable since this number is very
close to the average length of the documents
in the dataset, as reported in Figure1),
batch size of 24 samples,
and a learning rate of 5 10-5.</p>
        <p>F1(1) (%)</p>
        <p>SVM
Hand-crafted 98.90 0.07 82.73
aEmbeddings 96.16 0.15 62.37</p>
        <p>Both 98.94 0.04 83.47</p>
        <p>BERT
w/o oversampling / /
w/ oversampling 98.60 0.04 77.56</p>
        <p>Baseline 96.80 0.00
Table 2 resumes the results of the experiments
in stratified 5-fold cross-validation. The
performances are reported in terms of the macro average
of F1-score.</p>
        <p>After analyzing these results, we took the best
model and we used it in the leave-one-site-out
cross-validation context to test the reliability of the
system in an out-domain (site) problem. These
results are resumed in Table 3.</p>
        <p>First of all, we can notice that such
performances are much higher of the baseline system,
i.e. the performance achieved by a hypothetical
model that classifies all the samples as belonging
to the majority class (that is, the positive class).</p>
        <p>Due to the strong dataset imbalance and the low
batch size, training BERT without oversampling
the dataset leads the system to classify all samples
as belonging to the majority class, i.e. the
positive class. This leads to often obtain very bad
performance, i.e. the baseline performance. Anyway,
when this problem does not come up, the classifier
shows the lowest value of the F1-score. These
results clearly show the difficulties of BERT to deal
with unbalanced datasets. Oversampling the
minority class has shown to partially cope with such
problems, leading to an improvement in terms of
repeatability and performance.</p>
        <p>For what concerns the experiments with the
SVM-based system, they have shown that
handcrafted features have greater relevance for the task
than the embedding features. This suggests that
the (Italian) healthcare reviews domain may be
particularly lexical. Thus, sets of lexical features
show better performance than those
similaritybased features. However, the resulting best model
is the one with both sets of features,
outperforming the BERT-based system best configuration by
about three percentage points.</p>
        <p>Furthermore, the leave-one-site-out
experiments with this model result in a very good
performance, showing the system to be reliable also
in an out-domain (site) context. This last result
can be due to two factors: 1) the high degree of
overlap of the lexicon found in one domain on the
lexicon of all other domains; 2) a larger size of the
set used for training.</p>
        <p>
          In addition to the two main phases of
experiments, we further investigated the confidence of
the best model developed in making decisions.
The motivation behind this study is that it may
have application in real-world cases, where an
automated system is required to filter the documents
on which it is highly confident (i.e., above a certain
threshold) and then passes the most complex
documents to a human operator. To do so, we applied
the Platt scaling
          <xref ref-type="bibr" rid="ref12">(Platt, 1999)</xref>
          method on top of the
trained SVM model. This step is needed to
convert the output of the model from a decision score
d 2 ( 1; +1), i.e. the distance of the test
sample from the trained boundary, to a probabilistic
score p 2 [0; 1], representative of the system
confidence in making the decision. Figure 2 resumes
the results of this analysis. As expected, the
number of documents on which the system makes a
decision falls as the confidence threshold required of
the system increases. However, this trend does not
have such a negative slope and still classify more
than 91% of the documents with 99% confidence.
At the same time, the performance advantage is
clear, leading to an increase of F1-score on
negative samples by more than ten percentage points.
6
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>
        In this paper, we have introduced a novel
system for sentiment analysis for Italian reviews in
Healthcare. For the best of our knowledge, this is
the first work of this kind in such domain. To do
so, we have collected the first dataset for this
domain from the web. Then, we have implemented
and compared two types of classifiers of the state
of the art for such task, the SVM and BERT.
Despite the strong dataset imbalance, we have
obtained very good results, especially with the
SVMbased system, which outperformed the
BERTbased one, while maintaining a low computational
burden during training. However, there is a chance
that increasing the maximum sequence length of
BERT it may outperform our best-developed
system. Also, recent work
        <xref ref-type="bibr" rid="ref11">(Nozza et al., 2020)</xref>
        has analyzed the contribution of language-specific
models, showing in general improvements over
BERT multilingual for a wide variety of NLP
tasks. For this reason, it might be worth
including in future works the use of specific models for
Italian, such as GilBERTo6, UmBERTo7, and
AlBERTo8. The latter was already used for a
sentiment classification task
        <xref ref-type="bibr" rid="ref13">(Polignano et al., 2019)</xref>
        .
Future works on this dataset may also tackle the
task of sentiment classification including the
neutral class or sentiment regression of the average
scores. Moreover, future research may tackle the
task of cataloguing reviews to the area of disease
they belong, maybe including other features from
metadata such as titles.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>Our work is possible thanks to a general R&amp;D
agreement between the National Research
Council of Italy (CNR) and Confindustria (the main
association representing manufacturing and service
companies in Italy) and a specific R&amp;D agreement
between Webmonks s.r.l , CNR and the Campus
Bio-Medico University of Rome (UCBM).
6www.github.com/idb-ita/GilBERTo
7www.github.com/musixmatchresearch/
umberto
8www.github.com/marcopoli/AlBERTo-it</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Barbieri et al.2016]
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nissim</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Novielli</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Overview of the Evalita 2016 SENTIment POLarity Classification Task</article-title>
          ..
          <source>Proceedings of Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ) &amp;
          <article-title>Fifth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>Final Workshop (EVALITA</source>
          <year>2016</year>
          ), Napoli, Italy, December 5-
          <issue>7</issue>
          ,
          <fpage>2016</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Basile et al.2018] Basile,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Croce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            , &amp;
            <surname>Polignano</surname>
          </string-name>
          .
          <string-name>
            <surname>M</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 Aspect-based Sentiment Analysis Task (ABSITA)</article-title>
          .
          <source>Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          )
          <article-title>co-located with the Fifth Italian Conference on Computational Linguistics (CLiC-it</article-title>
          <year>2018</year>
          ), Turin, Italy,
          <source>December 12-13</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Cignarella et al.2018]
          <string-name>
            <surname>Cignarella</surname>
            ,
            <given-names>A. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frenda</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosco</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Overview of the EVALITA 2018 Task on Irony Detection in Italian Tweets (IronITA)</article-title>
          .
          <source>Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          )
          <article-title>co-located with the Fifth Italian Conference on Computational Linguistics (CLiC-it</article-title>
          <year>2018</year>
          ), Turin, Italy,
          <source>December 12-13</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Cimino et al.2018]
          <article-title>Cimino A</article-title>
          .,
          <string-name>
            <surname>De Mattei L</surname>
          </string-name>
          . &amp;
          <string-name>
            <surname>Dell'Orletta F</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Multi-task Learning in Deep Neural Networks at EVALITA 2018</article-title>
          .
          <article-title>Proceedings of the 6th evaluation campaign of Natural Language Processing and Speech tools for Italian (</article-title>
          <source>EVALITA'18)</source>
          ,
          <fpage>86</fpage>
          -
          <lpage>95</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Cimino and
          <string-name>
            <surname>Dell'Orletta2016] Cimino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Dell'Orletta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Tandem LSTM-SVM approach for sentiment analysis</article-title>
          .
          <source>In of the Final Workshop 7 December</source>
          <year>2016</year>
          , Naples (p.
          <fpage>172</fpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Devlin et al.2018]
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Fan et al.2008] Fan,
          <string-name>
            <given-names>R. E.</given-names>
            ,
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. W.</given-names>
            ,
            <surname>Hsieh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            ,
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. R.</given-names>
            , &amp;
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. J.</surname>
          </string-name>
          <year>2008</year>
          .
          <article-title>LIBLINEAR: A library for large linear classification</article-title>
          .
          <source>Journal of machine learning research 9</source>
          .
          <string-name>
            <surname>Aug</surname>
          </string-name>
          (
          <year>2008</year>
          ):
          <fpage>1871</fpage>
          -
          <lpage>1874</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>[Gao</surname>
          </string-name>
          et al.
          <year>2012</year>
          ]
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>G. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCullough</surname>
            ,
            <given-names>J. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>A changing landscape of physician quality reporting: analysis of patients' online ratings of their physicians over a 5-year period</article-title>
          .
          <source>Journal of medical Internet research</source>
          ,
          <volume>14</volume>
          (
          <issue>1</issue>
          ),
          <year>e38</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Greaves et al.2012]
          <string-name>
            <surname>Greaves</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Millett</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Consistently increasing numbers of online ratings of healthcare in England</article-title>
          .
          <source>J Med Internet Res</source>
          ,
          <volume>14</volume>
          (
          <issue>3</issue>
          ),
          <year>e94</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Mohammad et al.2013]
          <string-name>
            <surname>Mohammad</surname>
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiritchenko</surname>
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zhu</surname>
            <given-names>X.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>NRC-Canada: Building the state-of-the-art in sentiment analysis of tweets</article-title>
          .
          <source>In Proceedings of the Seventh international workshop on Semantic Evaluation Exercises, SemEval-2013</source>
          .
          <fpage>321</fpage>
          -
          <lpage>327</lpage>
          , Atlanta, Georgia, USA
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Nozza et al.2020]
          <string-name>
            <surname>Nozza</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bianchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>What the [mask]? making sense of language-specific BERT models</article-title>
          . arXiv preprint arXiv:
          <year>2003</year>
          .02912.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Platt1999]
          <string-name>
            <surname>Platt</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>1999</year>
          .
          <article-title>Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods</article-title>
          .
          <source>Advances in large margin classifiers</source>
          ,
          <volume>10</volume>
          (
          <issue>3</issue>
          ),
          <fpage>61</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Polignano et al.2019] Polignano,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Basile</surname>
          </string-name>
          , P., de Gemmis,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Semeraro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            , &amp;
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>AlBERTo: Italian BERT Language Understanding Model for NLP Challenging Tasks Based on Tweets</article-title>
          . In CLiC-it.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>