<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modeling the Fake News Challenge as a Cross-Level Stance Detection Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Costanza Conforti</string-name>
          <email>cc918@cam.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohammad Taher Pilehvar</string-name>
          <email>mp792@cam.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nigel Collier</string-name>
          <email>nhc30@cam.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Language Technology Lab, University of Cambridge</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The 2017 Fake News Challenge Stage 1, a shared task for stance detection of news articles and claims pairs, has received a lot of attention in recent years [ea18]. The provided dataset is highly unbalanced, with a skewed distribution towards unrelated samples - that is, randomly generated pairs of news and claims belonging to di erent topics. This imbalance favored systems which performed particularly well in classifying those noisy samples, something which does not require a deep semantic understanding. In this paper, we propose a simple architecture based on conditional encoding, carefully designed to model the internal structure of a news article and its relations with a claim. We demonstrate that our model, which only leverages information from word embeddings, can outperform a system based on a large number of hand-engineered features, which replicates one of the winning systems at the Fake News Challenge [HASC17], in the stance detection of the related samples.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Stance classi cation has been identi ed as a key
subtask in rumor resolution [ZAB+18]. Recently, a similar
approach has been proposed to address fake news
detection: as a rst step towards a comprehensive model
for news veracity classi cation, a corpus of news
articles, stance-annotated with respect to claims, has been
Copyright © CIKM 2018 for the individual papers by the papers'
authors. Copyright © CIKM 2018 for the volume as a collection
by its editors. This volume and its papers are published under
the Creative Commons License Attribution 4.0 International (CC
released for the Fake News Challenge (FNC-1)1.</p>
      <p>Characteristics of the corpus - The FNC-1
corpus is based on the Emergent dataset [FV16], a
collection of 300 claims and 2,595 articles discussing the
claims. Each article is labeled with the stance it
expresses toward the claim and summarized into a
headline by accredited journalists, in the framework of a
project for rumor debunking [Sil15].</p>
      <p>For creating the FNC-1 corpus, the headlines and
the articles were paired and labeled with the
corresponding stance, distinguishing between agreeing
(AGR), disagreeing (DSG) and discussing (DSC).
Additional 266 labeled samples were added to avoid
cheating [ea18]. Moreover, a number of unrelated
(UNR) samples were obtained by randomly matching
headlines with articles discussing a di erent claim. As
shown in Table 1, the nal class distribution was highly
skewed in favor of the UNR class, which amounted to
almost three quarters of the samples (Table 1).</p>
      <p>Characteristics of the FNC-1 winning
models - As a consequence of being randomly generated,
classi cation of UNR samples is relatively easy.
Moreover, given that the UNR samples constitute the large
majority of the corpus, most competing systems were
designed in order to perform well on this
easy-todiscriminate class. In fact, the three FNC-1
winning teams proposed relatively standard architectures
(mainly based on multilayer perceptrons, MLPs)
leveraging a large number of classic, hand-engineered NLP
features. While those systems performed very well on
the UNR class - reaching a F1 score higher than .99
they were not as e ective in the AGR, DSG and DSC
classi cation [ea18].</p>
      <p>FNC-1 as a Cross-Level Stance Detection
task - As shown in Table 3, one speci c
characteristic of the FNC-1 corpus consists in the clear
asymmetry in length between the headlines and the articles.</p>
      <p>While headlines consist of one sentence, the structure
1http://www.fakenewschallenge.org/
Article.
"Fear not arachnophobes, the story of Bunbury's \spiderman" might not be all it seemed.
[...] scientists have cast doubt over claims that a spider burrowed into a man's body [...] The story went global [...]
Earlier this month, Dylan Thomas [...] sought medical help [...] he had a spider crawl underneath his skin.
Mr Thomas said a [...] dermatologist later used tweezers to remove what was believed to be a "tropical spider".
[image via Shutterstock]
But it seems we may have all been caught in a web... of misinformation.</p>
      <p>Arachnologist Dr Framenau said [...]it was "almost impossible" [...] to have been a spider [...]
Dr Harvey said: "We hear about people going on holidays and having spiders lay eggs under the skin". [...]
Something which is true, [...] is that certain arachnids do live on humans. We all have mites living on our faces [...]
Dylan Thomas has been contacted for comment."
of an article is better described as a sequence of
paragraphs, where each paragraph plays a di erent role
in telling a story. Single paragraphs usually expresses
di erent views of a topic. Following the terminology
introduced by [JPN14], we propose to call this variant
of the classic Stance Detection task Cross-Level Stance
Detection.</p>
      <p>As shown in Table 1, an article consists in passages
presenting a news, reporting about interviews, giving
general background information and discussing similar
events happened in the past. In contemporary
newswriting prose, the most salient information is usually
condensed in the very rst paragraphs, following the
Inverted Pyramid style. This allows the reader for
rapid decision making [Sca00].</p>
      <p>For these reasons, we believe that detecting the
stance of an article with respect to a headline requires
a deep understanding not only of the position taken
in each paragraph with respect to the headline, but
also of the the complex interactions within the
article's paragraphs, as illustrated by the example in
Table 3. On the contrary, compressing both the
headline's and article's content into xed-size vectors, as
in the feature-based systems described in the previous
paragraph, fails in detecting those ne-grained
relationships and results in sub-optimal performance on
the stance detection of AGR, DSG and DSC samples.</p>
      <p>To test this assumption, we propose a simple
architecture based on conditional encoding, which is
designed in order to model the complex interactions
between headlines and articles described above, and we
compare it with one of the feature-based systems which
won the FNC-1 [HASC17].</p>
      <p>fnc-1
fnc-1-rel
instances</p>
      <p>In order to be able to assess the ability of the
systems to model the complex headline-article interplay
described above, we lter out the noisy UNR samples
and consider only the related samples (AGR, DSG and
DIS). Those samples were manually collected and
labeled by professional journalists and require deep
semantic understanding in order to be classi ed,
constituting a di cult task even for humans. This is
evident when looking at the inter-annotator agreement
of human raters, which drops from Fleiss' = :686
to :218 when including or excluding the UNR
samples, as reported in [ea18]. The nal label distribution
is reported in Table 1.
We implemented the model proposed by the team
Athene, which was ranked second at FNC-12. The
model consists of a 7-layer MLP with ReLU
activation. On the top of the architecture, a softmax layer
is used for prediction (Figure 1).</p>
      <p>Input is given in the form of a large matrix of
handengineered features. The considered set includes the
concatenation of feature vectors which separately
consider the headline and the article - like the presence of
refuting or polarity words (taken from a hand-selected
list of words as `hoax' or `debunk', and tf-idf weighted
Bag of Words vectors - and features which combine the
headline and the article (joint features in Figure 1)
- like word/ngram overlap between the headline and
the article, and cosine similarity of the embeddings of</p>
      <p>2We used Athene as the baseline as the FNC-1 winning model
was an ensemble [BSP17].
nouns and verbs between the headline and the article.</p>
      <p>Moreover, topic-based features based on non-negative
matrix factorization, latent Dirichlet allocation and
latent semantic indexing were used. For a detailed
description of the features, refer to [HASC17].
In order to model the headline-article interactions
described in Section 1, we adapt the
bidirectional conditional encoding architecture rst proposed
by [ARVB16] for stance detection of tweets.</p>
      <p>First, the article is split into n paragraphs. Both
the headline and the paragraphs are converted into
their embedding representations. The headline is then
processed by a Bi-LSTMh (Eq 1). Each paragraph is
then encoded by a further Bi-LSTMS1 (Eq 2), whose
initial cell states are initialized with the last states of
the respectively forward and backward LSTMs which
compose Bi-LSTMh (see Figure 2 for a representation
of the architecture's forward part). As pointed out
in [ARVB16], this allows Bi-LSTMS1 to read the
paragraph in a headline-speci c manner.</p>
      <p>Hh = Bi-LSTMh(Eh)
Hsi = Bi-LSTMS1(Esi )
8i 2 f1; :::; ng
(1)
(2)
where Eh 2 Re H and Esi 2 Re Si are
respectively the embedding matrix of the headline and of the
ith paragraph, H and Si are respectively the headline
and the ith paragraph length, e is the embedding size,
l is the hidden size, Hh 2 Rl H and Hs1 2 Rl Si .</p>
      <p>Then, each paragraph representation, conditionally
encoded on the headline, is processed by another
BiLSTMS2, conditioned on the previous paragraph. We
start the paragraph-conditioned reading of the
article from the bottom, as we assume the most salient
s
it
n
u
f
rsyea reobm
l
resulting in a matrix Hsi 2 Rl Si . We employ a
similar self-attention mechanism as in [ea16] in order to
soft-select the most relevant elements of the sentence.</p>
      <p>Given the sequence of vectors fh1; :::; hS g which
compose HSi , the nal representation of the ith paragraph
si is obtained as follows:
uit = tanh(Wshit + bs)</p>
      <p>ui&gt;t us
it = exp Pt ui&gt;t us
si = X</p>
      <p>thit
t
(4)
(5)
(6)
where the hidden representation of the word at
position t, uit, is obtained though a one-layer MLP (Eq 4).</p>
      <p>The normalized attention matrix t is then obtained
though a softmax operation (Eq 5). Finally, si is
computed by a weighted sum of all hidden states ht with
the weight matrix t (Eq 6). The sentence
representations fs1; :::; sng are aggregated using a backward
LSTM, as in Figure 2. The nal prediction y^ is
obtained with a softmax operation over the tagset.
For the feature-based model, we downloaded the
feature matrices used by [HASC17] for their FNC-1 best
submission3 and selected the columns corresponding to
the related samples. For the conditional model, we
initialized the embedding matrix with word2vec
embeddings4. Only words which occurred more than 7 times
were included in the embedding matrix. Words not
included in word2vec were zero-initialized. In order to
avoid over tting, we did not ne-tune the embeddings
during training. The main structures of the models
were implemented in keras, using Tensor ow for
implementing customized layers. Refer to Appendix ??
for the complete list of hyperparameters used to train
both architectures.
3.2</p>
      <p>Evaluation Metrics
In the FNC-1 context, a so-called FNC score was
proposed for evaluation: this hierarchical evaluation
metric gives 0.25 points for a correct REL/UNR classi
cations, which is incremented of 0.75 points in case of a
correct AGR/DSA/DSC classi cation5. This was
motivated by the high imbalance in favor of UNR class.</p>
      <p>However, as in our experiments we are only
considering REL samples, the FNC score does not constitute
a useful evaluation metric. Following [ea18], we use
macro-averaged precision, recall and F1 score, which
is less a ected by the high class imbalance (Table 1).</p>
    </sec>
    <sec id="sec-2">
      <title>Feature-based model</title>
    </sec>
    <sec id="sec-3">
      <title>Conditional model</title>
      <p>RECALL 6363..37%% 3609..73%% 6345..91%% 4581..00%%</p>
      <p>R
G
A</p>
      <p>DSG DSC
Target Class</p>
      <p>C
E
R
P
1304 67 532 68.6%
18.46% 0.94% 7.53% 31.4%
Results of experiments are reported in Table 4. The
proposed conditional model clearly outperforms the
feature-based baseline for all considered metrics,
despite having a considerably minor number of trainable
parameters. Interestingly, the feature-based model
seems to o er a better generalization over the test set,
while the gap between development and test set
performance in the conditional model seems to indicate
over tting.</p>
      <p>Detailed performance on single classes is shown in
Figure 3. Thanks to the presence of features speci
cally designed to target the presence of refuting words,
the baseline model is able to reach a Precision of 15.2%
in classifying the very infrequent DSG class (7.5% of
occurrences). The conditional model, which did not
receive any explicit signal of the presence of negation,
su ers more from this data imbalance, and reaches
a Precision of 7.3% on DSG samples. On the other
hand, by attening the entire article into a xed-size
vector, the feature-based system looses the nuances in
the argumentative structure of the news story. As a
consequence, this system struggles to distinguish
between AGR and DSC samples and tends to favor the
most frequent DSC class, which receives the highest
Precision and Recall scores. On the contrary, the
conditional model is able to spot the subtle di erences
between AGR and DSC samples, reaching high
Precision and satisfactory Recall in both classes despite the
large class imbalance - 27.7% AGR vs. 65.2% DSC
samples.</p>
      <p>3https://drive.google.com/open?id=0B0muIdcdTp7UWVyU0duSDRUd3c
4https://code.google.com/archive/p/word2vec/
5https://github.com/FakeNewsChallenge/fnc-1baseline/blob/master/utils/score.py</p>
      <sec id="sec-3-1">
        <title>Conclusions</title>
        <p>Given the results discussed in the previous Section,
we believe the strategy of modeling the FNC-1 as an
Asymmetric Stance Detection problem is promising.
In future work, we will carry on a detailed qualitative
analysis to test the extent to which our conditional
model is able to model the narrative structures of
articles and their interactions with the headlines. The
generalizability of such architecture to other domains
can be tested on other publicly available corpora, as
the recently released ARC dataset by [ea18].</p>
      </sec>
      <sec id="sec-3-2">
        <title>Acknowledgments</title>
        <p>The rst author (CC) would like to thank the Siemens
Machine Intelligence Group (CT RDA BAM MIC-DE,
Munich) and the NERC DREAM CDT (grant no.
1945246) for partially funding this work. The third
author (NC) is grateful for support from the UK
EPSRC (grant no. EP/MOO5089/1).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [ARVB16]
          <string-name>
            <given-names>Isabelle</given-names>
            <surname>Augenstein</surname>
          </string-name>
          , Tim Rocktaschel, Andreas Vlachos, and
          <string-name>
            <given-names>Kalina</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <article-title>Stance detection with bidirectional conditional encoding</article-title>
          .
          <source>In Proceedings of EMNLP 2016</source>
          , pages
          <fpage>876</fpage>
          {
          <fpage>885</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[BSP17] [ea16] [ea18] [FV16] Sean Baird</source>
          , Doug Sibley, and
          <string-name>
            <given-names>Yuxi</given-names>
            <surname>Pan</surname>
          </string-name>
          .
          <article-title>Talos targets disinformation with fake news challenge victory</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          https://blog.talosintelligence.com/
          <year>2017</year>
          /06/,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Zichao</given-names>
            <surname>Yang</surname>
          </string-name>
          et al.
          <article-title>Hierarchical attention networks for document classi cation</article-title>
          .
          <source>In Proceedings of NAACL-HLT</source>
          <year>2016</year>
          , pages
          <fpage>1480</fpage>
          {
          <fpage>1489</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Hanselowski</surname>
          </string-name>
          et al.
          <article-title>A retrospective analysis of the fake news challenge stancedetection task</article-title>
          .
          <source>In Proceedings of COLING 2018</source>
          , pages
          <year>1859</year>
          {
          <year>1874</year>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>Emergent: a novel data-set for stance classi cation</article-title>
          .
          <source>In Proceedings of NAACL-HLT</source>
          <year>2016</year>
          , pages
          <fpage>1163</fpage>
          {
          <fpage>1168</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [HASC17]
          <article-title>Andreas Hanselowski, PVS Avinesh, Benjamin Schiller, and Felix Caspelherr. Description of the system developed by team athene in the fnc-1</article-title>
          .
          <source>Technical report</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [JPN14]
          <string-name>
            <given-names>David</given-names>
            <surname>Jurgens</surname>
          </string-name>
          , Mohammad Taher Pilehvar, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Navigli</surname>
          </string-name>
          . Semeval-2014 [Sca00]
          <article-title>[Sil15] task 3: Cross-level semantic similarity</article-title>
          .
          <source>In Proceedings of SemEval 2014</source>
          , pages
          <fpage>17</fpage>
          {
          <fpage>26</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Scanlan</surname>
          </string-name>
          .
          <article-title>Reporting and writing: Basics for the 21st century</article-title>
          . Harcourt College Publishers,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Craig</given-names>
            <surname>Silverman</surname>
          </string-name>
          .
          <article-title>Lies, damn lies and viral content</article-title>
          . Columbia University,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [ZAB+18]
          <string-name>
            <surname>Arkaitz</surname>
            <given-names>Zubiaga</given-names>
          </string-name>
          , Ahmet Aker, Kalina Bontcheva, Maria Liakata, and
          <string-name>
            <given-names>Rob</given-names>
            <surname>Procter</surname>
          </string-name>
          .
          <article-title>Detection and resolution of rumours in social media: A survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>51</volume>
          (
          <issue>2</issue>
          ):
          <fpage>32</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>