<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SSN NLP at CheckThat! 2020: Tweet Check Worthiness Using Transformers, Convolutional Neural Networks and Support Vector Machines</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sachin Krishan Thyaharajan</string-name>
          <email>sachinkrishnan18128@cse.ssn.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kayalvizhi Sampath</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thenmozhi Durairaj</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rishivardhan Krishnamoorthy</string-name>
          <email>rishivardhan18126@cse.ssn.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>SSN College Of Engineering</institution>
          ,
          <addr-line>Chennai</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Social media has become a signi cant source of information for a large fraction of the population. One such popular social media is twitter. An excessive amount of misinformation spread has become ubiquitous. But it is immensely computationally expensive to verify every claim made in every tweet. In this paper, the authors have explored Machine learning solutions to score a tweet on its worthiness to be factchecked. In this paper, we present approaches using CNN, Transformer models and SVM for CLEF-2020 CheckThat! Check-Worthiness task.</p>
      </abstract>
      <kwd-group>
        <kwd>Check Worthiness</kwd>
        <kwd>Convolutional Neural Network</kwd>
        <kwd>Transformers</kwd>
        <kwd>Support Vector Machine</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The rampant spread of fake news on social media has become an all too
familiar plight. Fake news has become indistinguishable from real news. The hazards
caused by fake news create confusion and misunderstanding about important
social and political issues. According to a study on Twitter, false news travels
faster than true news. Research project nds humans, not bots, are primarily
responsible for the spread of misleading information [17]. With important
political gures and business tycoons active on Twitter, it has become a rather
important stage for global information. It is unrealistic to check every tweet to
verify the information it holds due to the exorbitant computational requirement
with 500 million tweets posted every day.</p>
      <p>
        This puts us in need of an algorithm that can lter or rank tweets based on their
check worthiness which is the goal of the CLEF-2020 Check That!'s [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Check
Worthiness task 1. We use CNN, Transformer models and SVM to score each
tweet based on their tweet worthiness.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Prevention of fake news collides with freedom of speech, hence detection can
be done based on objective facts to curb fake news [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Prevalent
state-of-theart fact-checking methods employ the use of feature engineering techniques to
extract features from each sentence and thereby its context. ClaimBuster[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
system arose as the rst work to check worthiness. The system extracts sentiment
features, TF-IDF word representations, POS tags and named entities. These
features were derived from sentence level, thereby no contextual information
between sentences was captured. Extending ClaimBuster's work, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] incorporated
contextual awareness into representation by the inclusion of sentence positioning
in a speaker segment, speaker mentioning the opponent, audience reactions and
sentence similarity to segments as features. [14] proposes a deep learning
framework for detecting Rumors from Microblogs with Recurrent Neural Networks for
learning hidden representations that capture the variation of contextual
information of relevant posts over time.
      </p>
      <p>
        In last year's CheckThat! [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] Team Copenhagen [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] achieved the best
performance. They made use of LSTM RNN that learned dual token embeddings,
domain-speci c embeddings and syntactic dependencies. Team TheEarthIsFlat
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] made use of a feed-forward neural network with two hidden layers which takes
as input Standard Universal Sentence Encoder (SUSE) embeddings [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for the
current sentence as well as for the two previous sentences as a context.
      </p>
      <p>Our system's approach can be considered to pivot on sentence-level features.
We do not include context-aware features into our data due to the low amount
of training data and less compute power.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Dataset description</title>
      <p>The data for the task was given in two formats: .tsv and .json. The training set
has 672 data points, the development set has 150 data points. We have used the
.tsv format which is a TAB separated text le. The text encoding is UTF-8. A
row of the .tsv le has the following format:
topic id &lt; T ab &gt; tweet id &lt; T ab &gt;tweet url &lt; T ab &gt; tweet text &lt; T ab &gt; claim
&lt; T ab &gt; check worthiness</p>
      <p>The column descriptions are as follows:
topic id unique ID for the topic of the tweet
tweet id unique ID for each tweet given by Twitter
tweet url URL of the given tweet
tweet text text content in the tweet
claim It is a binary value of 1 if the tweet contains a claim else 0
check worthiness It is a binary value of 1 if the tweet is worthy to be fact checked else 0
4
4.1</p>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <sec id="sec-4-1">
        <title>Data preparation</title>
        <p>
          The tweets given were preprocessed initially to remove stop words and
punctuation. Then, all the tweets were normalized. The whole data was lemmatized and
tokenized using the nltk library[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Training the models</title>
        <p>The data prepared is given to various models that include Convolutional Neural
Network (CNN), Transformers and Support Vector Machine.</p>
        <p>
          CNN
Convolutional neural networks (CNN) are trained on top of pre-trained word
vectors for classi cation tasks [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The training data was vectorized using a
Word2Vec vectorizer using a pre-trained Google News word vector model with
3 million 300-dimension English word vectors and then padded with zeros up to
the maximum sequence length in the given sentences. The padded sequence was
fed to the convolutional neural network made of an input layer, an embedding
layer, 5 convolutional layers each followed by max-pooling layers, a dropout layer
and 2 dense layers. The convolutional layers each use 200 lters and the layers
have kernel sizes of 2,3,4,5 and 6 sequentially with RelU activation. The following
dense layer has 128 nodes with RelU activation and the following dense layer has
2 nodes with sigmoid activation. The model was trained and saved for classifying
test data.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Transformers</title>
        <p>
          The models used are derivatives of Google's BERT [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. BERT stands for
Bidirectional Encoder Representations from Transformers. The models used in this
paper are RoBERTa [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and XLNet [18] which are derivatives of BERT that give
better performance. XLNet is developed to work seamlessly with the Auto
Regression objective, including integrating Transformer-XL and the careful design
of the two-stream attention mechanism. It manages to overcome the de
ciencies of BERT whilst requiring more compute power and memory (GPU/TPU
memory). The XLNet model was ne-tuned over the pre-trained XLNet Base
Cased language model that comprises 12 Transformer blocks, 12 self-attention
heads and 768 hidden dimensions. RoBERTa makes use of a robustly optimized
method that improves on BERT by modifying key hyperparameters in BERT. It
was ne-tuned over the RoBERTa base language model that comprises 12
Transformer blocks, 12 self-attention heads and 768 hidden dimensions with a total
parameter of 215M. For our case of experimentation, both transformer models
were trained with hyperparameters as 50 epochs with training batch size set to
128 and the learning rate set to 4e-5.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Support Vector Machine</title>
        <p>SVM or Support Vector Machine is a traditional machine learning algorithm. The
principle behind the algorithm is it creates a hyperplane which separates the data
into classes. The text is vectorized using a count vectorizer. The resulting count
matrix is transformed to a normalized TF-IDF representation then classi ed
using SVM [16] classi ers of scikit-learn [15].
4.3</p>
      </sec>
      <sec id="sec-4-5">
        <title>Choosing models</title>
        <p>Amongst these four models experimented, RoBERTa was submitted. It was
selected based on the f1 score and because of the proven robustness of transformer
models for natural language processing tasks. The performance of our models in
the development set is shown in Table 1.</p>
        <p>Model F1
RoBERTa 0.730
XLNet 0.642
CNN 0.654</p>
        <p>SVM 0.590
Accenture
Team Alex
check square
QMUL-SDS
Tobb Etu
SSN NLP
Factify
BustingMisinformation
nlpir01
ZHAW
UAICS
TheUniversityofShe eld</p>
        <p>MAP RR
We have presented our submission for Task 1 of CheckThat! @ CLEF-2020 to
predict the check-worthiness of tweets. The task was approached with various
methods that include transformers, CNN and SVM. Among these models, the
Roberta transformer model performed better than other methodologies on the
development set. Hence, the output from RoBERTa was submitted for
evaluation. From table 2, the results show that our team has a perfect recall score and
a good MAP score. The performance can further be improved with some other
transformer models such as XLNet, Electra, BERT, etc and inclusion of more
data samples for training.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgement</title>
      <p>We express our sincere thanks to DST-SERB and HPC laboratory for providing
space and materials needed for our work.
14. Ma, J., Gao, W., Mitra, P., Kwon, S., Jansen, J., Wong, K.F., Cha, M.: Detecting
rumors from microblogs with recurrent neural networks (07 2016)
15. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O.,
Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A.,
Cournapeau, D., Brucher, M., Perrot, M., Duchesnay, E.: Scikit-learn: Machine
learning in Python. Journal of Machine Learning Research 12, 2825{2830 (2011)
16. Prasetijo, A.B., Isnanto, R.R., Eridani, D., Soetrisno, Y.A.A., Arfan, M., Sofwan,
A.: Hoax detection system on indonesian news sites based on text classi cation
using svm and sgd. In: 2017 4th International Conference on Information Technology,
Computer, and Electrical Engineering (ICITACEE). pp. 45{49 (2017)
17. Vosoughi, S., Roy, D., Aral, S.: The spread of true and false news online. Science
359(6380), 1146{1151 (2018). https://doi.org/10.1126/science.aap9559, https://
science.sciencemag.org/content/359/6380/1146
18. Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., Le, Q.V.: Xlnet:
Generalized autoregressive pretraining for language understanding (2019)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Arampatzis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eickho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , N. (eds.):
          <article-title>Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <source>Interaction Proceedings of the Eleventh International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ).
          <source>LNCS (12260)</source>
          , Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Atanasova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karadzhov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohtarami</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Da San Martino, G.:
          <article-title>Overview of the clef-2019 checkthat! lab: Automatic identi cation and veri cation of claims. task 1: Check-worthiness</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Barron-Ceden~o,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Elsayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            , Da San Martino, G.,
            <surname>Hasanain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Suwaileh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Haouari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Babulkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Hamdan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Nikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Shaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sheikh Ali</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          : Overview of CheckThat! 2020:
          <article-title>Automatic identi cation and veri cation of claims in social media</article-title>
          .
          <source>In: Arampatzis et al. [1]</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>yi Kong</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hua</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Limtiaco</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>John</surname>
          </string-name>
          , R.S.,
          <string-name>
            <surname>Constant</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guajardo-Cespedes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sung</surname>
            ,
            <given-names>Y.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strope</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kurzweil</surname>
          </string-name>
          , R.:
          <article-title>Universal sentence encoder (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Favano</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>C.M.L.P.:</surname>
          </string-name>
          <article-title>TheEarthIsFlat's submission to CLEF'19 CheckThat! challenge</article-title>
          .
          <source>In: CLEF 2019 Working Notes. Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum. CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Lugano, Switzerland (
          <year>2019</year>
          ) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Figueira</surname>
          </string-name>
          , ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>The current state of fake news: challenges and opportunities</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>121</volume>
          ,
          <issue>817</issue>
          {
          <volume>825</volume>
          (12
          <year>2017</year>
          ). https://doi.org/10.1016/j.procs.
          <year>2017</year>
          .
          <volume>11</volume>
          .106
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>H.C.S.J.L.C.</surname>
          </string-name>
          <article-title>: Neural weakly supervised fact check-worthiness detection with contrastive sampling-based ranking loss</article-title>
          .
          <source>In: CLEF 2019 Working Notes. Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum. CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Lugano, Switzerland (
          <year>2019</year>
          ) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hassan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arslan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tremayne</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Toward automated fact-checking: Detecting check-worthy factual claims by claimbuster</article-title>
          . pp.
          <year>1803</year>
          {
          <volume>1812</volume>
          (08
          <year>2017</year>
          ). https://doi.org/10.1145/3097983.3098131
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hassan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tremayne</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Detecting check-worthy factual claims in presidential debates</article-title>
          .
          <source>In: Proceedings of the 24th ACM International on Conference on Information and Knowledge Management</source>
          . p.
          <year>1835</year>
          {
          <year>1838</year>
          . CIKM '
          <volume>15</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA (
          <year>2015</year>
          ). https://doi.org/10.1145/2806416.2806652, https://doi.org/10.1145/2806416. 2806652
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Convolutional neural networks for sentence classi cation (</article-title>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoyanov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Roberta: A robustly optimized bert pretraining approach (</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Loper</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Nltk: The natural language toolkit (</article-title>
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>