<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>QMUL-SDS @ SardiStance2020: Leveraging Network Interactions to Boost Performance on Stance Detection using Knowledge Graphs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rabab Alkhalifa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arkaitz Zubiaga</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Imam Abdulrahman bin Faisal University</institution>
          ,
          <country country="SA">Saudi Arabia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Queen Mary University of London</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents our submission to the SardiStance 2020 shared task, describing the architecture used for Task A and Task B. While our submission for Task A did not exceed the baseline, retraining our model using all the training tweets, showed promising results leading to (favg 0.601) using bidirectional LSTM with BERT multilingual embedding for Task A. For our submission for Task B, we ranked 6th (f-avg 0.709). With further investigation, our best experimented settings increased performance from (f-avg 0.573) to (f-avg 0.733) with same architecture and parameter settings and after only incorporating social interaction features- highlighting the impact of social interaction on the model's performance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Framed as a classification task, the stance
detection consists in determining if a textual utterance
expresses a supportive, opposing or neutral
viewpoint with respect to a target or topic
        <xref ref-type="bibr" rid="ref11">(Küçük
and Can, 2020)</xref>
        . Research in stance detection has
largely been limited to analysis of single
utterances in social media. Furthering this research, the
SardiStance 2020 shared task
        <xref ref-type="bibr" rid="ref13 ref6">(Cignarella et al.,
2020)</xref>
        focuses on incorporating contextual
knowledge around utterances, including metadata from
author profiles and network interactions. The task
included two subtasks, one solely focused on the
textual content of social media posts for
automatically determining their stance, whereas the other
allowed incorporating additional features
available through profiles and interactions. This
pa0Copyright © 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
per describes and analyses our participation in the
SardiStance 2020 shared task, which was held as
part of the EVALITA
        <xref ref-type="bibr" rid="ref3">(Basile et al., 2020)</xref>
        campaign and focused on detecting stance expressed
in tweets associated with the Sardines movement.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        In social media, classical features can be
extracted by using stylistic signals from text such as
bag of n-grams, char-grams, part-of-speech labels,
and lemmas
        <xref ref-type="bibr" rid="ref16">(Sobhani et al., 2019)</xref>
        , structural
signals such as hashtags, mentions, uppercase
characters, punctuation marks, and the length of the
tweet
        <xref ref-type="bibr" rid="ref17 ref18">(Wojatzki et al., 2018; Sun et al., 2016)</xref>
        ,
and pragmatic signals related to author’s profile
        <xref ref-type="bibr" rid="ref9">(Graells-Garrido et al., 2020)</xref>
        . With modern deep
learning models, there is shift towards
contextualised representations using word vector
representation algorithms, either by having
personalised language models trained on task specific
language or as a pre-trained language model
offered after training using complex architecture and
billions of documents. Using deep learning
layers as automated feature engineering methods can
be implemented to train the model afterwards. In
        <xref ref-type="bibr" rid="ref1">(Augenstein et al., 2016)</xref>
        , they utilized
Bidirectional Conditional Encoding using LSTM
achieving state-of-the-art results on stance detection task.
Recently, there is a resurgence of research in
incorporating network homophily
        <xref ref-type="bibr" rid="ref12">(Lai et al., 2017)</xref>
        to represent social interactions within a network.
Moreover, Knowledge graphs
        <xref ref-type="bibr" rid="ref19">(Xu et al., 2019)</xref>
        can in turn represent these complex network
relationships (e.g. authors friendships) as simple
embedded vectors sampled considering the nodes and
weighted edges within the network complexity
structure.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Definition of the Tasks</title>
      <p>
        The stance detection task has been defined in
previous work as consisting in determining the
viewpoint of an utterance with respect to a
target topic
        <xref ref-type="bibr" rid="ref11">(Küçük and Can, 2020)</xref>
        , while others
define it as that consisting in determining an
author’s viewpoint with respect to the veracity of a
rumour, usually referred to as rumour stance
classification
        <xref ref-type="bibr" rid="ref21">(Zubiaga et al., 2018)</xref>
        . SardiStance
focuses on the former, and is split into two subtasks:
Textual Stance Detection (Task A) and
Contextual Stance Detection (Task B)
        <xref ref-type="bibr" rid="ref13 ref6">(Cignarella et al.,
2020)</xref>
        . Baselines are provided for Task A using
SVM+unigrams as (f-avg. 0:578), and for Task B
as (f-avg. 0:628)
        <xref ref-type="bibr" rid="ref13 ref6">(Lai et al., 2020)</xref>
        .
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Settings</title>
    </sec>
    <sec id="sec-5">
      <title>Frequency-based features: These represent fre</title>
      <p>
        quency vectors including unigram, punctuation
and hashtags provided by
        <xref ref-type="bibr" rid="ref13 ref6">(Cignarella et al., 2020)</xref>
        .
Further, we include TFiDF vectors.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Embedding-based features: word embedding</title>
      <p>
        Italian Wikipedia Embedding
        <xref ref-type="bibr" rid="ref4">(Berardi et al.,
2015)</xref>
        trained using GloVe 1, Fasttext with
        <xref ref-type="bibr" rid="ref5">(Bojanowski et al., 2017)</xref>
        2 trained using skip-gram
model and with 300 dimensions, and TWITA
embedding
        <xref ref-type="bibr" rid="ref2">(Basile et al., 2018)</xref>
        . For TWITA,
two versions of the same tweets were generated.
One preprocessing words where each vector has
100 dimensions, provided by
        <xref ref-type="bibr" rid="ref13 ref6">(Cignarella et al.,
2020)</xref>
        3 and referred to as TWITA100. The other
1https://github.com/MartinoMensio/it_
vectors_wiki_spacy
      </p>
      <p>
        2https://fasttext.cc/docs/en/
pretrained-vectors.html
3https://github.com/mirkolai/
one trained by us without any preprocessing and
each vector has 300 dimensions, referred to as
TWITA300. We also experimented with
multilingual BERT in Task A 4
        <xref ref-type="bibr" rid="ref7">(Devlin et al., 2019)</xref>
        .
Cosine similarity vectors which was introduced
previously in
        <xref ref-type="bibr" rid="ref1 ref10 ref17 ref8">(Eger and Mehler, 2016)</xref>
        to encode
the word meaning within the embedding space. In
our work, we used TWITA300 to train the
similarity vectors of all the words in the training set.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Network-based features: Encoding users graph.</title>
      <p>To represent user interactions as nodes and edges,
we used a counting scalar value and added one if
each of the following relationships exists:
friendships, retweets, quotes and replies, e.g. if all of
them exist then the edge weight between two
accounts is four. We calculated all the accounts
provided and generate a directed complex graph
conditioned by the existence of friendship, resulting
in 669,745 nodes, 2,871,791 edges with an
average in-degree of 4.2879 and average out-degree of
4.2879.</p>
      <p>
        Generating GNN Embeddings. Taking as input
the encoded network relationships, GNN
embeddings use different sampling techniques to
represent every node as a vector. To extract these
vectors, we experiment with different graph
neural network models, namely struct2vec
        <xref ref-type="bibr" rid="ref15">(Ribeiro
et al., 2017)</xref>
        , deepwalk
        <xref ref-type="bibr" rid="ref14">(Perozzi et al., 2014)</xref>
        and
node2vec
        <xref ref-type="bibr" rid="ref1 ref10 ref17 ref8">(Grover and Leskovec, 2016)</xref>
        .
      </p>
    </sec>
    <sec id="sec-8">
      <title>NeuralNetwork-based features As illustrated in</title>
      <p>Figure 1, we have different deep learning
models to extract features separately for both word
embedding and similarity vectors matrices. In
our work, we experiment with Convectional
Neural Network (CNN) models and Long short-term
memory (LSTM) models. Variations of CNN
models where applied to NLP downstream tasks
as feature extraction methods for text
classification. In our work, we used two variations of CNN.
In one model, we used a CNN as a one-head
1DCNN with kernel size of 5 allowing the model to
extract features with 5-grams vectors using 32
filters. Followed by a max pooling layer with pool
size of 2 then flattened layer. In another model,
we used a CNN as a multi-headed 2D-CNN with
1, 2, 3, 5 grams filter sizes, initialising the kernel
weights with a Rectified Linear Unit (ReLU)
activation function and normal distribution weights.
Followed by a max pooling layer with different
evalita-sardistance/</p>
      <p>
        4https://tfhub.dev/tensorflow/bert_
multi_cased_L-12_H-768_A-12/2
pooling sizes taken as one columns pooling
filter with the maximum text length excluding few
grams sizes. For the LSTM, we used two variants.
One is a simple bidirectional LSTM of 64 units
followed by concatenations of max pooling and
average pooling layers, and attention bidirectional
LSTM proposed by
        <xref ref-type="bibr" rid="ref20">(Yang et al., 2016)</xref>
        using 64
units followed by 128 units then attention layers5.
Feature Reduction. We experiment with different
reduction length: 50, 100 and 150. Then. we set
our PCA reduction to 100 as it showed best
performance on evolution set.
      </p>
      <p>
        Sentence Cleaning. We set the cleaning function
to match the preprocessing function by
        <xref ref-type="bibr" rid="ref13 ref6">(Cignarella
et al., 2020)</xref>
        to generate TWITA100.
      </p>
      <p>We used four final layers to receive the features
and concatenate them (see Figure 1). In all of the
experiments, our dropout layer set to 0:2, followed
by a dense layer with rule activation function and
another dropout layer of 0:2. Finally, a
probability vector of the three classes is generated. To
determine the correct class, we choose the one class
with the highest probability.
5</p>
    </sec>
    <sec id="sec-9">
      <title>Results</title>
      <p>In this section, we discuss the results of our
systems submitted to the two tasks.</p>
      <p>For Task A, we used attention Bidirectional
LSTM model performance compared to using
different word embedding models, also we
analysed impact of the preprocessing of the runs.
Since there are too many parameters to compare
with, we compared the performance of the
embedding models. Our submitted models, BERT and
TWITA300 illustrated in Table 1 with showed
most promising results using different settings.
With only %80 training data, similarity vectors
generalised better than all other embedding
models. While, when all data are trained, the best
model is the multilingual BERT embedding with
no pre-processing (f-avg 0.601), followed by
similarity vectors using cleaned text (f-avg 589).</p>
      <p>For Task B, we used different feature extraction,
frequency vectors, word embedding and social
interaction embedding models, and monitor their
performance while activating the pre-processing
step in all experiments. With a diverse range of
parameters, we experimented with a total of 3845
random runs. Then, we selected the best
mod5https://www.kaggle.com/mlwhiz/
attention-pytorch-and-keras
Tst. f-avg
els considering macro f-score for the two classes
under consideration (AGAINST and FAVOR)
(favg). Results are shown in Table 2. By
comparing our runs by adding social interaction features,
our models with different settings showed a clear
improvement on our models. In 1#M, we utilise
Conv2D (see NeuralNetwork-based features) for
embedding vectors with TfiDF unigram and tweet
length, where the model achieved an increase on
performance of (f-avg 0.16) when social
interaction vectors incorporated into the model. All other
models showed the same improvement with an
increase of (f-avg 0.115, 0.118, 0.081, 0.021) for
3#M, 5#M, 7#M and 9#M, respectively.
6</p>
    </sec>
    <sec id="sec-10">
      <title>Discussion and main findings</title>
      <p>The pipeline depicted in Figure 1 was designed
to investigate the impact of multiple features on
stance detection using variations of feature
extraction methods, which have been experimented in
previous work but we adapted them to the Italian
language in our settings. The training set contains
2132 instances with no evaluation set. In our work,
we create a stratified split of 80-20 to evaluate the
model, which leads to a training data with 1705
samples. Further, our investigation attempted to
randomise different settings, with the aim of
submitting the top two with highest f-avg score on the
remaining set (Eval. 426) for both tasks.
Consequently, we found that this methodology did not
generalise well with the testing results. However,
our main findings remain consistent across
different settings when compared with our results
using the stratified split (T%80) and when the model
was retrained using all the data (T%100). While
our submission evaluated both tasks separately, we
discuss all conclusions jointly in this section.</p>
      <p>Having different random settings over all
frequency-based features (14, in our case) would
be a bad strategy to evaluate the methods and
come up with the best approach. To verify
if we need to include all of these, we run
an experiment by including only one feature
from (unigram, Tfidf_unigram, chargrams,
network_reply_community, userinfobio). The
selection of these features where based on selecting
the best runs using only one feature from our
randomised parameters. Using all the training set
and CONV2D with (fasttext;TWEC300) and
reduced SVs with deepwalk user’s social
interaction vector, (userinfobio;chargrams) achieved
(favg 0.703 and 0.704), respectively. This is also
higher than using AttLSTM for the same
settings which achieved (f-avg 0.638 and 0.610).</p>
      <p>In general, we achieve better performance with
CONV2D than AttnLSTM for the same settings
on the test data. In another experiment, we reduced
all the 14 frequency-based parameters achieving
(f-avg 0.714) which performs worse than our best
3#M (see 2). Our main conclusion is that the
number of features available is not necessarily
correlated with the model’s performance boost.</p>
      <p>In another experiment, we attempted to
compare the performance of TWEC100 with
TWEC300 (see Section 4). From Table 1, we
observed that lower dimensionality and
preprocessing may cause the model to under perform
by around (f-avg 0.050), at least. Though, this
impact was not significant with T%100. However,
matching the processing between the embedding
vocabulary and the annotated set yields better
performance. For example, TWITA100 was
more persistent on performance between T%80
and T%100. This highlights the importance
of pre-processing and reducing the differences
between the embedding vocabularies and labelled
sentences. In general, our embedding experiment
for Task A show high sensitivity on model
performance with pre-processing settings.</p>
      <p>Inspired by previous work on encoding word
meanings, we experimented with SVs embedding.</p>
      <p>Interestingly, these vectors showed high f-avg,
Eval.</p>
      <p>f-avg</p>
      <p>Tst. f-avg</p>
      <p>T%80 T%100
%
are the ones represented with . Bold fonts show highest/above baseline results
better than BERT and TWITA300 with T%80
and social interactions. Using different random
although it showed a significant drop when the
runs, our best model achieved (f-avg 0.733)
levermodel was trained with T%100. This finding
aging deepwalk-based knowledge graphs
embedopens an investigation towards the ability of SVs
dings, FastText and similarity feature vectors
exto perform better under different settings. For that,
tracted by two multi-headed convolutional neural
we removed PCA(SVs) and run same settings of
networks from auther’s utterance. This motivates
#M1, and our model achieved (f-avg 0.678),
showour future, aiming to reduce the model complexity
ing a significant impact of SVs on model’s
perforand automate the feature selection process.
mance. Further, we investigate the robustness of
deepwalk modelling over node2vec and struct2vec
for the same best settings of #M1, resulting on
(favg 0.641 and 0.604) for node2vec and struct2vec,
respectively. Also, in terms of accuracy, the
deepwalk model produces an improved accuracy of
(% 0.725) compared to node2vec (% 0.665) and
struct2vec (% 0.658). This indicates that deepwalk
is more reliable on this testing set than other
models.
7</p>
    </sec>
    <sec id="sec-11">
      <title>Conclusion</title>
      <p>In this work, we described a state-of-the-art stance
detection system leveraging different features
including author profiling, word meaning context
8</p>
    </sec>
    <sec id="sec-12">
      <title>Acknowledgments</title>
      <p>This research utilised Queen Mary’s Apocrita
HPC facility, supported by QMUL Research-IT.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Isabelle</given-names>
            <surname>Augenstein</surname>
          </string-name>
          , Tim Rocktäschel, Andreas Vlachos, and
          <string-name>
            <given-names>Kalina</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Stance detection with bidirectional conditional encoding</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>876</fpage>
          -
          <lpage>885</lpage>
          , Austin, Texas, November.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Mirko Lai, and
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Long-term social media data collection at the university of turin</article-title>
          .
          <source>In Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . CEUR-WS.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>EVALITA 2020: Overview of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ).
          <article-title>CEUR-WS.org</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Giacomo</given-names>
            <surname>Berardi</surname>
          </string-name>
          , Andrea Esuli, and Diego Marcheggiani.
          <year>2015</year>
          .
          <article-title>Word embeddings go to italy: A comparison of models and training datasets</article-title>
          .
          <source>In IIR.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>5</volume>
          :
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Teresa</surname>
          </string-name>
          <string-name>
            <surname>Cignarella</surname>
          </string-name>
          , Mirko Lai, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>SardiStance@EVALITA2020: Overview of the Task on Stance Detection in Italian Tweets</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ). CEURWS.org.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In Proceedings of NAACL-HLT</source>
          , pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Steffen</given-names>
            <surname>Eger</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Mehler</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>On the linearity of semantic change: Investigating meaning variation via dynamic graph models</article-title>
          .
          <source>In Proceedings of ACL (Volume 2: Short Papers)</source>
          , pages
          <fpage>52</fpage>
          -
          <lpage>58</lpage>
          , Berlin, Germany,
          <year>August</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Eduardo</given-names>
            <surname>Graells-Garrido</surname>
          </string-name>
          ,
          <article-title>Ricardo Baeza-Yates, and</article-title>
          <string-name>
            <given-names>Mounia</given-names>
            <surname>Lalmas</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Every colour you are: Stance prediction and turnaround in controversial issues</article-title>
          .
          <source>In 12th ACM Conference on Web Science</source>
          ,
          <source>WebSci '20, page 174-183</source>
          , New York, NY, USA. Association for Computing Machinery.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Grover</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jure</given-names>
            <surname>Leskovec</surname>
          </string-name>
          .
          <year>2016</year>
          . node2vec:
          <article-title>Scalable feature learning for networks</article-title>
          .
          <source>In Proceedings of ACM SIGKDD</source>
          , pages
          <fpage>855</fpage>
          -
          <lpage>864</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Dilek</given-names>
            <surname>Küçük</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fazli</given-names>
            <surname>Can</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Stance detection: A survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>53</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Mirko</given-names>
            <surname>Lai</surname>
          </string-name>
          , Marcella Tambuscio, Viviana Patti, Giancarlo Ruffo, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Extracting graph topological information and users' opinion</article-title>
          .
          <source>In Lecture Notes in Computer Science</source>
          , volume
          <volume>10456</volume>
          LNCS, pages
          <fpage>112</fpage>
          -
          <lpage>118</lpage>
          . Springer Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Mirko</given-names>
            <surname>Lai</surname>
          </string-name>
          , Alessandra Teresa Cignarella, Delia Irazú Hernández Farías, Cristina Bosco, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Multilingual stance detection in social media political debates</article-title>
          .
          <source>Computer Speech &amp; Language</source>
          ,
          <volume>63</volume>
          :
          <fpage>101075</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Bryan</given-names>
            <surname>Perozzi</surname>
          </string-name>
          , Rami Al-Rfou, and
          <string-name>
            <given-names>Steven</given-names>
            <surname>Skiena</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Deepwalk: Online learning of social representations</article-title>
          .
          <source>In Proceedings of ACM SIGKDD</source>
          , pages
          <fpage>701</fpage>
          -
          <lpage>710</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Leonardo FR Ribeiro</surname>
          </string-name>
          , Pedro HP Saverese, and Daniel R Figueiredo.
          <year>2017</year>
          .
          <article-title>struc2vec: Learning node representations from structural identity</article-title>
          .
          <source>In Proceedings of ACM SIGKDD</source>
          , pages
          <fpage>385</fpage>
          -
          <lpage>394</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Parinaz</given-names>
            <surname>Sobhani</surname>
          </string-name>
          , Diana Inkpen, and
          <string-name>
            <given-names>Xiaodan</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Exploring deep neural networks for multitarget stance detection</article-title>
          .
          <source>Computational Intelligence</source>
          ,
          <volume>35</volume>
          (
          <issue>1</issue>
          ):
          <fpage>82</fpage>
          -
          <lpage>97</lpage>
          , feb.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Qingying</given-names>
            <surname>Sun</surname>
          </string-name>
          , Zhongqing Wang,
          <string-name>
            <surname>Qiaoming Zhu</surname>
            , and
            <given-names>Guodong</given-names>
          </string-name>
          <string-name>
            <surname>Zhou</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Exploring various linguistic features for stance detection</article-title>
          .
          <source>In Natural Language Understanding and Intelligent Applications</source>
          , pages
          <fpage>840</fpage>
          -
          <lpage>847</lpage>
          , Cham. Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Wojatzki</surname>
          </string-name>
          , Torsten Zesch, Saif Mohammad, and
          <string-name>
            <given-names>Svetlana</given-names>
            <surname>Kiritchenko</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Agree or Disagree: Predicting Judgments on Nuanced Assertions</article-title>
          .
          <source>In Proceedings of *SEM</source>
          , pages
          <fpage>214</fpage>
          -
          <lpage>224</lpage>
          , Stroudsburg, PA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Zhenhui</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Qiang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Wei Chen</surname>
          </string-name>
          , Yingbao Cui, Zhen Qiu, and
          <string-name>
            <given-names>Tengjiao</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Opinion-aware knowledge embedding for stance detection</article-title>
          .
          <source>In Jie Shao, Man Lung Yiu</source>
          , Masashi Toyoda, Dongxiang Zhang, Wei Wang, and Bin Cui, editors,
          <source>Web and Big Data</source>
          , pages
          <fpage>337</fpage>
          -
          <lpage>348</lpage>
          , Cham.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Zichao</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Diyi</given-names>
            <surname>Yang</surname>
          </string-name>
          , Chris Dyer, Xiaodong He,
          <string-name>
            <surname>Alex Smola</surname>
            , and
            <given-names>Eduard</given-names>
          </string-name>
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Hierarchical attention networks for document classification</article-title>
          .
          <source>In Proceedings of NAACL-HLT</source>
          , pages
          <fpage>1480</fpage>
          -
          <lpage>1489</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Arkaitz</given-names>
            <surname>Zubiaga</surname>
          </string-name>
          , Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, and
          <string-name>
            <given-names>Isabelle</given-names>
            <surname>Augenstein</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Discourseaware rumour stance classification in social media using sequential classifiers</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>54</volume>
          (
          <issue>2</issue>
          ):
          <fpage>273</fpage>
          -
          <lpage>290</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>