<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DeepReading @ SardiStance: Combining Textual, Social and Emotional Features</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mar´ıa S. Espinosa NLP</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IR Group UNED</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Spain mespinosa@lsi.uned.es</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alvaro Rodrigo NLP</string-name>
          <email>rodrigo.agerri@ehu.eus</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IR Group UNED</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Spain alvarory@lsi.uned.es</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Roberto Centeno NLP &amp; IR Group UNED</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Rodrigo Agerri HiTZ Center - Ixa University of the Basque Country UPV/EHU</institution>
        </aff>
      </contrib-group>
      <fpage>12</fpage>
      <lpage>13</lpage>
      <abstract>
        <p>In this paper we describe our participation to the SardiStance shared task held at EVALITA 2020. We developed a set of classifiers that combined text features, such as the best performing systems based on large pre-trained language models, together with user profile features, such as psychological traits and social media user interactions. The classification algorithms chosen for our models were various monolingual and multilingual Transformer models for text only classification, and XGBoost for the non-textual features. The combination of the textual and contextual models was performed by a weighted voting ensemble learning system. Our approach obtained the best score for Task B, on Contextual Stance Detection.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>One of the most important research topics in the
field of Natural Language Processing (NLP) is
automatic information extraction from textual data.
The recent rise of social media has completely
changed the way in which people communicate
their ideas and has thus led to the emergence of
new research problems regarding the automatic
analysis of online contents, such as sentiment
analysis, emotion recognition, or fake news
detection. Stance detection (usually considered as
a subproblem of sentiment analysis) is part of
the aforementioned family of research problems</p>
      <p>
        Copyright © 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
        <xref ref-type="bibr" rid="ref10 ref11">(Ku¨ c¸u¨k and Can, 2020)</xref>
        . While there are
various formulation of the stance detection task, for
SardiStance 2020 the aim is to detect the stance
(AGAINST, FAVOR or NEUTRAL) conveyed by
a given tweet with respect to a specific, previously
given topic
        <xref ref-type="bibr" rid="ref14">(Mohammad et al., 2016)</xref>
        , namely,
about the Sardines movement in Italy.
      </p>
      <p>
        Thus, we address the problem of
automatic stance detection in tweets written in
Italian language for the SardiStance 2020 shared
task
        <xref ref-type="bibr" rid="ref1 ref10 ref2 ref5">(Cignarella et al., 2020)</xref>
        , organized within
EVALITA 2020
        <xref ref-type="bibr" rid="ref1 ref10 ref2 ref5">(Basile et al., 2020)</xref>
        . In this paper
we include the participation of three teams within
the framework of the DeepReading project 1: (1)
Ixa Group, (2) UNED group, and (3)
DeepReading Group. While Ixa focused on developing text
classifiers based on textual information only (Task
A), UNED was more interested in exploring how
to use contextual information available (Task B).
Likewise, DeepReading is the product of
combining both Ixa and UNED systems into one.
      </p>
      <p>In this sense, the main idea behind our model
is to exploit textual information, based on
finetuning large pre-trained language models for text
classification, together with contextual
information using several feature categories, such as
psychological traits of the user, social media data, and
network based features. As a result of our joint
effort, we submitted 4 and 5 runs, respectively, to
tasks A and B. The official results show that our
systems obtained the 3rd position among the
constrained runs submitted to Task A, which
considered only textual information for prediction, and
1st position from 13 participants for Task B, which
considered textual and contextual information.</p>
      <sec id="sec-1-1">
        <title>1http://ixa2.si.ehu.es/deepreading/</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Systems Description</title>
      <p>In this section we first describe the text
classification systems developed for Task A and then the
contextual features used to train XGBoost
classifiers for Task B. We also include a description of
the strategies used to combine the classifiers from
both tasks, which resulted in the winner system for
Task B.
2.1</p>
      <sec id="sec-2-1">
        <title>Task A: Textual Stance Detection</title>
        <p>
          The main objective of our participation in Task A
was to benchmark the performance, on the stance
detection task for Italian, of large pre-trained
language models based on the transformer
architecture
          <xref ref-type="bibr" rid="ref19 ref3">(Vaswani et al., 2017)</xref>
          . This would help us to
identify the best performing models which will be
leveraged to generate features for Task B
(Contextual Stance Detection).
        </p>
        <p>
          As for many other Natural Language Processing
(NLP) tasks, current best performing systems for
text classification are based on large pre-trained
language models which allow to build rich
representations of text based on contextual word
embeddings. Deep learning methods in NLP
represent words as continuous vectors on a low
dimensional space, called word embeddings. The
first approaches generated static word embeddings
          <xref ref-type="bibr" rid="ref13 ref19 ref3">(Mikolov et al., 2013; Bojanowski et al., 2017)</xref>
          ,
namely, they provided a unique vector-based
representation for a given word independently of the
context in which the word occurs. This means that
polysemy cannot be represented.
        </p>
        <p>In order to address this problem, contextual
word embeddings were proposed. The idea is to
be able to generate word representations
according to the context in which the word occurs.
Currently there are many approaches to generate such
contextual word representations, but we will
focus on publicly available multilingual and
monolingual pre-trained models for Italian.</p>
        <p>
          There are several multilingual versions of these
models. Thus, the multilingual version of BERT
          <xref ref-type="bibr" rid="ref12 ref16 ref17 ref6 ref7">(Devlin et al., 2019)</xref>
          was trained for the top 100
languages with the largest Wikipedias. More
recently, XLM-RoBERTa
          <xref ref-type="bibr" rid="ref12 ref16 ref17 ref6 ref7">(Conneau et al., 2019)</xref>
          distributes a multilingual model which contains
104 languages trained on 2.5 TB of Common
Crawl data. Italian is included in both multilingual
models.
        </p>
        <p>
          These multilingual models perform very well in
tasks involving high-resourced languages such as
English or Spanish, but their performance drops
when applied to languages not so well represented
in the language model
          <xref ref-type="bibr" rid="ref1 ref10 ref2 ref5">(Agerri et al., 2020)</xref>
          .
Although this is still an open issue, a number of
reasons can be found in the literature. First, each
language has to share the quota of substrings and
parameters with the rest of the languages
represented in the pre-trained multilingual model. As
the quota of substrings partially depends on corpus
size, this means that larger languages such as
English or Spanish are better represented than other
languages such as Italian. Moreover, multilingual
models also seem to behave better for structurally
similar languages
          <xref ref-type="bibr" rid="ref1 ref10 ref2 ref5">(Karthikeyan et al., 2020)</xref>
          .
        </p>
        <p>We have benchmarked four monolingual
pretrained language models for Italian: AlBERTo,
GilBERTo, UmBERTo and Italian BERT XXL
with the aim of comparing them with respect to the
multilingual pre-trained models previosly
mentioned, namely, mBERT and XLM-RoBERTa.</p>
        <p>
          AlBERTo is a BERT base pre-trained
lowercased model containing a vocabulary of 128k
terms from 200M of Italian tweets
          <xref ref-type="bibr" rid="ref12 ref16 ref17 ref6 ref7">(Polignano et
al., 2019)</xref>
          .
        </p>
        <p>
          The Italian BERT XXL models 2 are also based
on the BERT base architecture. The training data
contains the Italian Wikipedia, various parts of the
OPUS corpus and the OSCAR corpus for Italian
          <xref ref-type="bibr" rid="ref12 ref16 ref17 ref6 ref7">(Ortiz Sua´rez et al., 2019)</xref>
          , for a total of 81GB of
Italian text.
        </p>
        <p>
          GilBERTo3 is based on the RoBERTa base
          <xref ref-type="bibr" rid="ref12 ref16 ref17 ref6 ref7">(Liu
et al., 2019)</xref>
          architecture, an improved, optimized
version of BERT which discards the next sentence
prediction task. The model was trained using the
Italian Oscar
          <xref ref-type="bibr" rid="ref12 ref16 ref17 ref6 ref7">(Ortiz Sua´rez et al., 2019)</xref>
          , which
contains 71GB of text. The vocabulary used
consisted of 32k BPE subwords tokenized by the
SentencePiece tokenizer4.
        </p>
        <p>UmBERTo5 also leverages the RoBERTa base
architecture, the OSCAR corpus for Italian and the
SentencePiece tokenizer, but it adds Whole Word
Masking to the training process. The idea is to
mask an entire word, instead of subwords, if at
least one of all (sub-)tokens generated by
SentencePiece was originally selected as mask.</p>
        <sec id="sec-2-1-1">
          <title>2https://github.com/dbmdz/berts 3https://github.com/idb-ita/GilBERTo 4https://github.com/google/sentencepiece 5https://github.com/musixmatchresearch/umberto</title>
          <p>In this task, we use several sets of features with the
purpose of trying to model user’s behaviour when
writing a tweet. We obtain such features from both
the text and the social network. Our hypothesis
is that the stance of a user regarding a particular
tweet is highly correlated with the way of writing
of the own user extracted in terms of
psychological and emotional features. On the other hand, we
focus on exploring how the concept of
“homopohily”, namely, the tendency of individuals to
associate and bond with similar individuals, previously
studied in DellaPosta et al. (2015). In order to test
this hypothesis, we have tested different models
that are explained below.</p>
          <p>In this task, we use several sets of features with
the purpose of trying to model user’s behaviour
when writing a tweet. We obtain such features
from text and the network.</p>
          <p>The complete set of features extracted from the
data is depicted in Table 1. The set of features used
in the model can be divided into five main types:
psychological, emotional, Twitter-based,
networkbased, and language model features.</p>
          <p>Psychological features. These features were
extracted using a third-party API developed by
Symanto6. Each tweet was sent to the API in order
to retrieve the personality traits and
communication styles obtained from the analysis of the tweet
contents.</p>
          <p>The personality traits value would be either
“emotional” or “rational” depending on the
analysis of the user’s text. The value returned by
the API when the communication styles are
re6https://symanto-research.github.io/symanto-docs/
quested is a collection of traits, such as
selfrevealing, which means sharing one’s own
experience and opinion; fact-oriented, which implies
focusing on factual information, objective
observations or statements; information-seeking, that is,
posing questions; and action-seeking or aiming to
trigger someone’s action by giving
recommendation, requests or advice.</p>
          <p>
            Emotional features. In order to retrieve the
emotion values from the tweets, we used Russell’s
circumplex model of affect
            <xref ref-type="bibr" rid="ref18">(Russell, 1980)</xref>
            .
Russell argues that emotions can be conceptualized
in a two-dimensional continuous space where the
axes correspond to the degree of arousal and
valence (or pleasure). These two dimensions form
a Cartesian space that can be configured in a
circular order in which the different combinations of
valence and arousal correspond to one of four
discrete emotion regions: tired, tense, excited, and
pleased.
          </p>
          <p>
            The values for the degree of arousal and valence
of the tweets were obtained using an adaptation to
Italian language of the Affective Norms for
English Words (ANEW)
            <xref ref-type="bibr" rid="ref4">(Bradley and Lang, 1999)</xref>
            .
This database was developed from translations of
the 1,034 English words present in the ANEW
dictionary and from words taken from Italian
semantic norms
            <xref ref-type="bibr" rid="ref15">(Montefinese et al., 2014)</xref>
            .
          </p>
          <p>Twitter features. Exploring how the users
behave in the social network could offer some
insights on the stance tendency of the users. The
collection of Twitter data of each user contained
four features: the number of statuses published by
the user, the number of users followed by the user,
the number of users following the user, and the
creation date of the Twitter account of the user.</p>
          <p>Network features. Using the FRIEND.csv
data provided, we built a network consisting of
669817 nodes (or users) and 2847197 edges (or
relationships) in order to represent the following
network of the users. From that network, we
extracted a sub-graph containing the users of known
stance from the training data and the users
involved in testing in order to calculate the mean
distances of each user to the rest of known stance
users using the following formula:
dT (n) =</p>
          <p>PjiT=j1 d2n!i</p>
          <p>1
jT j
where jT j is the total number of users of a
determined stance (AGAINST, FAVOR, NONE) and</p>
          <p>2
dn!i corresponds to the square distance in users
from node n to node i. From this calculation we
obtained 3 values per user: mean distance to users
against (dagainst), mean distance to users in
favor (dfavor), and mean distance to neutral users
(dnone)</p>
          <p>Language model features. In order to
incorporate the language model results into the rest of
the features of the system we choose the best
performing, at the development phase, of the models
described in Section 2.1, which was UmBERTo.
Since this kind of language models use a great
amount of features for learning and training, the
strategy used in order to incorporate the language
model without having a great imbalance in the
number of features representing each category,
consisted in extracting the probabilities assigned
by the model to each class for each tweet. In this
way, the language model would be present in 3 of
the 18 features of the model, and it would
therefore have a balanced size with regards to the rest
of features of the model.
3
3.1</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <sec id="sec-3-1">
        <title>Task A</title>
        <p>As we use the base version of every transformer
model we can fine-tune them in a basic GPU of
12GB RAM. Hyperparameter tuning (batch size,
maximum sequence length, learning rate and
number of epochs) was performed on the development
set. For mBERT, AlBERTo, Italian BERT XXL
and UmBERTo the best configuration was:
maximum sequence length 256, batch 32, learning rate
5e-5, and 5 epochs. For GilBERTo we used the
same values except the number of epochs, which
was increased to 10. Finally, the best performing
hyperparameters for XLM-RoBERTa was the
following: maximum sequence length 256, batch 16,
learning rate 2e-5, and 10 epochs.</p>
        <p>While the monolingual models clearly
outperformed both mBERT and XLM-RoBERTa on the
development data, we decided to submit the three
best monolingual runs and the best multilingual
one. Table 2 reports the official results obtained
by each of the models and their position with
respect to the ranking of constrained runs for Task
A released by the task organizers. Our
submission based on Italian BERT XXL was clearly the
best of our four runs, although its performance was
around 1.5 scores in F1 lower than the winner
system for Task A. Furthermore, the ranking obtained
in the test does not correspond with the results
obtained during the development phase, where
UmBERTo outperformed the other monolingual
models by more than 3 points in F1 score.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Task B</title>
        <p>We presented a total of five models to Task B,
which consisted of different combinations of the
features listed in Table 1.</p>
        <p>
          Models 1, 2, and 3. During the training and
development phases of the models, several
configurations were tested on models 1, 2, and 3,
including training with different classifiers, such as
Random Forest Classifier, Decision Tree Classifier and
XGBoost Classifier. The best performing
classifier was XGBoost configured for multi-class
classification and taking into account class weights in
order to deal with the imbalance present in the
data. XGBoost is an efficient and scalable
implementation of gradient boosting framework by
          <xref ref-type="bibr" rid="ref9">(Friedman, 2001)</xref>
          . With regards to the set of
features, the first approach to the task considered only
psychological, emotion, and Twitter features. For
the second model, network features were added
to the feature set. Finally, model 3 considered
the probabilities of each class (AGAINST,
FAVOR, NONE) predicted by the UmBERTo
language model as three additional features for
training.
        </p>
        <p>Models 4 and 5. These two models were
constructed using voting based ensemble learning.
The voting system for model 4 considered
predictions of models 1, 2, and 3 as well as
predictions by the best performing language models on</p>
      </sec>
      <sec id="sec-3-3">
        <title>Team</title>
        <p>Ixa
DeepReading
DeepReading</p>
        <p>UNED</p>
        <p>UNED
the development data: UmBERTo, GilBERTo, and
Italian BERT XXL, described in Section 2.1. The
most common predicted value among the 6
systems was chosen as the final prediction of model
4. In case of having two or more values with the
same counts, the final value is randomly selected.
On the other hand, model 5 used a weighted voting
ensemble learning in which each of the systems
considered had as weight the F1 value obtained on
the development data. Therefore, the model
considered the weighted predictions of each system in
order to choose the final prediction.</p>
        <p>Table 3 shows the official results obtained by
each model and their position with respect to the
ranking for Task B on Contextual Stance
Detection. As it can be noted, model 5 ranked first in this
task, obtaining an average F1 of 0.7445. Models
3 and 4 also had promising results in the official
test set, ranking third and fourth, respectively, and
just 0.0079 below the system which obtained the
second best result. Model 2 had a slightly worse
performance, ranking seventh from a total of 13,
but still 0.0604 above the baseline. Finally, model
1 had the lowest performance, ranking last for the
task.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>Furthermore, we can see that predictions from
model 3 also experimented a great increase in true
positives of each of the classes. This increase
is related to the inclusion of the language model
into the features of model 2, which demonstrates
the importance of textual data in stance detection
tasks.</p>
      <p>Finally, models 4 and 5 shows the adequacy of
combining several complementary systems in
order to improve results. Since each single model
can detect the stance for different instances, a
proper combination of them could outperform
single models.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>In this paper we have shown the benefits of
exploiting information from different and
heterogeneous sources. For our participation to the
SardiStance 2020 shared task we have experimented with
classifiers trained with the textual content of the
tweets as well as with features based on social
networks. This combination of features has allowed
us to obtain the best overall results in the task.</p>
      <p>As future work, we plan to further explore the
contribution of network information. Besides, we
want to develop new divergent models and study
how to combine them.
This work has been partially funded by the
Spanish Ministry of Science, Innovation and
Universities (DeepReading RTI2018-096846-B-C21,
MCIU/AEI/FEDER, UE), and DeepText
(KK2020/00088), funded by the Basque Government.
Rodrigo Agerri is additionally funded by the
RYC2017-23647 fellowship and acknowledges the
donation of a Titan V GPU by the NVIDIA
Corporation. Maria S. Espinosa is also funded by the
European Social Fund through the Youth
Employment Initiative (YEI 2019).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Agerri et al.2020]
          <article-title>Rodrigo Agerri</article-title>
          , In˜aki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, and
          <string-name>
            <given-names>Eneko</given-names>
            <surname>Agirre</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Give your text representation models some love: the case for basque</article-title>
          .
          <source>In LREC 2020</source>
          , pages
          <fpage>4781</fpage>
          -
          <lpage>4788</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Basile et al.2020]
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>EVALITA 2020: Overview of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>EVALITA</source>
          <year>2020</year>
          .
          <article-title>CEUR-WS.org</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Bojanowski et al.2017]
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>TACL</source>
          ,
          <volume>5</volume>
          :
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Bradley and Lang1999]
          <string-name>
            <surname>Margaret M Bradley and Peter J Lang</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Affective norms for english words (anew): Instruction manual and affective ratings</article-title>
          .
          <source>Technical Report 1, Technical report C-1</source>
          , the center for research in psychophysiology, University of Florida.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Cignarella et al.2020]
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Teresa</surname>
          </string-name>
          <string-name>
            <surname>Cignarella</surname>
          </string-name>
          , Mirko Lai, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>SardiStance@EVALITA2020: Overview of the Task on Stance Detection in Italian Tweets</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ).
          <article-title>CEUR-WS.org</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Conneau et al.2019]
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Conneau</surname>
          </string-name>
          , Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzma´n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Unsupervised cross-lingual representation learning at scale</article-title>
          . arXiv:
          <year>1911</year>
          .02116.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Devlin et al.2019]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Friedman2001]
          <string-name>
            <surname>Jerome H Friedman</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Greedy function approximation: a gradient boosting machine</article-title>
          .
          <source>Annals of statistics</source>
          , pages
          <fpage>1189</fpage>
          -
          <lpage>1232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Karthikeyan et al.
          <year>2020</year>
          ]
          <string-name>
            <given-names>K</given-names>
            <surname>Karthikeyan</surname>
          </string-name>
          ,
          <string-name>
            <surname>Zihan</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stephen Mayhew</surname>
            , and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Roth</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Crosslingual ability of multilingual bert: An empirical study</article-title>
          .
          <source>In International Conference on Learning Representations (ICLR).</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Ku¨
          <article-title>c¸u¨k and Can2020] Dilek Ku¨ c¸u¨k and</article-title>
          <string-name>
            <given-names>Fazli</given-names>
            <surname>Can</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Stance detection: A survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>53</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Liu et al.2019] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>RoBERTa: A robustly optimized bert pretraining approach</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Mikolov et al.2013]
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Mohammad et al.2016]
          <string-name>
            <given-names>Saif</given-names>
            <surname>Mohammad</surname>
          </string-name>
          , Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and
          <string-name>
            <given-names>Colin</given-names>
            <surname>Cherry</surname>
          </string-name>
          .
          <year>2016</year>
          . SemEval
          <article-title>-2016 task 6: Detecting stance in tweets</article-title>
          .
          <source>In SemEval-2016)</source>
          , pages
          <fpage>31</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Montefinese et al.2014]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Montefinese</surname>
          </string-name>
          , Ettore Ambrosini, Beth Fairfield, and
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Mammarella</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The adaptation of the affective norms for english words (anew) for italian</article-title>
          .
          <source>Behavior research methods</source>
          ,
          <volume>46</volume>
          (
          <issue>3</issue>
          ):
          <fpage>887</fpage>
          -
          <lpage>903</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Ortiz Sua´rez et al.2019]
          <article-title>Pedro Javier Ortiz Sua´rez, Benoˆıt Sagot,</article-title>
          and
          <string-name>
            <given-names>Laurent</given-names>
            <surname>Romary</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures</article-title>
          .
          <source>Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-7) 2019. Cardiff, 22nd July</source>
          <year>2019</year>
          , pages
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Polignano et al.2019]
          <string-name>
            <given-names>Marco</given-names>
            <surname>Polignano</surname>
          </string-name>
          , Pierpaolo Basile, Marco de Gemmis, Giovanni Semeraro, and
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>AlBERTo: Italian BERT Language Understanding Model for NLP Challenging Tasks Based on Tweets</article-title>
          .
          <source>In Proceedings of the Sixth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2019</year>
          ), volume
          <volume>2481</volume>
          . CEUR.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>[Russell1980] James A Russell</source>
          .
          <year>1980</year>
          .
          <article-title>A circumplex model of affect</article-title>
          .
          <source>Journal of personality and social psychology</source>
          ,
          <volume>39</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1161</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Vaswani et al.2017]
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
          <string-name>
            <surname>Łukasz Kaiser</surname>
            , and
            <given-names>Illia</given-names>
          </string-name>
          <string-name>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>