<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detecting COVID-19-Related Conspiracy Theories in Tweets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Youri Peskine</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giulio Alfarano</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ismail Harrando</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Papotti</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raphael Troncy EURECOM</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France ifrstName.lastName@eurecom.fr</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Misinformation in online media has become a major research topic the last few years, especially during the COVID-19 pandemic. Indeed, false or misleading news about coronavirus have been characterized as an infodemic1 by the World Health Organization, because of how fast it can spread online. A considerable vector of spreading misinformation is represented by conspiracy theories. During this challenge, we tackled the problem of detecting COVID-19-related conspiracy theories in tweets. To perform this task, we used different approaches such as a combination of TFIDF and machine learning algorithms, transformer-based neural networks or Natural Language Inference. Our best model obtains a MCC score of 0.726 for the main task on the validation set and a MCC score of 0.775 on the test set making it the best performing method among the challenge competitors.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The full description of the task is detailed in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and more
information about the dataset can be found in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Text classification
is a problem widely studied in many diferent fields for various
applications such as sentiment analysis or topic modeling. Standard
machine learning based approaches used in combination with Term
Frequency - Inverse Document Frequency (TFIDF) are considered
decent baselines [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for performing text classification tasks.
However, the recent introduction of transformer-based architectures like
BERT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], RoBERTa [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or DistilBERT [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] has allowed significant
improvement in various text-based problems [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>In order to tackle this challenge, we studied three diferent kind of
approaches. The first uses a combination of TFIDF and machine
learning algorithms. The second approach uses Natural Language
Inference (NLI) combined with metadata from Wikipedia. The third
approach aims at fine-tuning transformer-based models. In the
following sections, we discuss the experiments we pursued for each
of these approaches. In order to ease reproducibility, we release all
our code at https://github.com/D2KLab/mediaeval-fakenews.</p>
    </sec>
    <sec id="sec-3">
      <title>TFIDF-based approach</title>
      <p>TFIDF is one of the most widely used feature extraction techniques
in the field of text processing, often used in parallel with
preprocessing techniques such as tokenization, capitalization and stop
word removal, which we also applied to the dataset in question. The
derived features are fed to diferent supervised machine learning
1https://www.who.int/health-topics/infodemic
methods that allow us to obtain a first baseline. Several algorithms
have been tested: Decision Tree, Naive Bayes classifier (Gaussian
and Bernoullian), AdaBoost, Ridge and Logistic Regression. In the
case of Task 1, these were used in a multi-class asset. In the
multilabel case of Task 2, we used a multi-output classifier with the
diferent methods listed above as estimators: in this scenario the
algorithm instantiates a binary model for each conspiracy theory.
Finally, in the Task 3, only the strictly tree-based algorithms were
tested, since they are the only ones to allow a multi-label and
multiclass output.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>NLI-based approach</title>
      <p>This approach relies on leveraging pre-trained language models
that are then fine-tuned on the task of NLI. Put simply, given two
statements (a premise and a hypothesis), these models are trained to
classify the logical relationship between them: entailment
(agreement or support), contradiction (disagreement), or neutrality
(undetermined). Since these models are trained to project statements that
share similar opinions into close points in their embedding space,
our hypothesis was to identify the tweets that support/discuss the
same conspiracies using this common embedding space.</p>
      <p>For the Task 1, we need to diferentiate between the diferent
stances (agreement, discussion, neutrality) regarding conspiracy
theories. Therefore, we generate an embedding for all the tweets
using the fine-tuned model, and we then classify them using a
K-nearest Neighbor classifier, the idea being that tweets sharing
similar stances would be embedded close to each other. This
approach can also be applied as is for the second task.</p>
      <p>For the Task 2, we provided as a premise to the model a definition
of each conspiracy theory, relying mostly on the Wikipedia articles
describing them, thus classifying a tweet as pertaining to one of the
listed conspiracies if the pre-trained model predicts that there is a
entailment relationship between the definition of the conspiracy
(premise) and the tweet text (hypothesis).</p>
      <p>Finally, as a combination of both methods, we also used some
annotated tweets related to a specific conspiracy as a premise
instead of a definition of the conspiracy, and proceed to classify them
by whether an entailment relationship is found.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Transformer-based approach</title>
      <p>
        Transformer-based models have been performing remarkably well
on text classification tasks in the last few years. For our submissions,
we used RoBERTa [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] large pre-trained models and
COVID-TwitterBERT (CT-BERT) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] pre-trained models in diferent ways to tackle
each sub-task. The CT-BERT model is a BERT-large model
pretrained on COVID-related tweets. For this approach, we decided to
use some pre-processing on the input tweets. We replaced all emojis
with their textual meaning, and removed all the ’#’ characters.
      </p>
      <p>The simplest strategy when working with transformer-based
models for the first task is to approach it as a 3-class classification
Y. Peskine, G. Alfarano, I. Harrando, P. Papoti, R. Troncy
problem. Both RoBERTa and CT-BERT are fine tuned on the data to
perform classification with a weighted Cross Entropy loss function.
We approach the second task as a multi-label binary classification
problem. Both models are fine tuned to perform this objective with
a weighted Binary Cross Entropy loss function.</p>
      <p>The third and main task can be performed with diferent
strategies. We first try to combine our results of the first two tasks, by
labeling the tweets with the level detected in Task 1 for the
conspiracy theories detected in Task 2. While this approach obtains
convincing results, it is not able to deal properly with cases where
tweets discuss about one conspiracy theory but support another
one. An alternative approach is to train both transformer-based
models to perform the main task directly. These models are fine
tuned for nine diferent classification problems with nine Cross
Entropy loss functions, one for each conspiracy theory. The final
loss is the weighted sum of the nine losses. This training framework
is illustrated in Figure 1. The advantage of such approach is that a
single model is trained to perform all the tasks at once, because the
ifrst two tasks are just simplifications of the main task.</p>
      <p>
        In our experiments, the weights of all our loss functions are
proportional to the inverse frequency of each class or sub-class
they are related to, and all our models are trained using the AdamW
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] optimizer.
3
      </p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND ANALYSIS</title>
      <p>Our results for this challenge are presented in Table 1. All the models
have been first evaluated on a stratified 5-fold cross-validation set
and then evaluated on the test set.</p>
      <p>Transformer-based approaches obtained the most competitive
results. First we notice that RoBERTa models are under-performing
compared to CT-BERT models on all the tasks. The latter models
are more suited to this dataset because it contains tweets that use
plenty of COVID-related vocabulary that would not be understood
with the former models. It is also worth mentioning that models
trained on the main task (Task 3) perform better on Task 1 than
their task-specific counterpart.</p>
      <p>For the TFIDF approach, the best performing method for Task
1 is the Support Vector Machine with a MCC score of 0.461. For
Task 2, the Decision Tree gave the best result with 0.585 using the
multi output classifier. On Task 3, the best result was given by the
Decision Tree with an MCC of 0.497.</p>
      <p>For the NLI approach, the observed results of these methods on
cross-validation did not measure up to the fully-trained models.</p>
      <p>Models</p>
      <p>TFIDF (SVC)
NLI transformer</p>
      <p>RoBERTa</p>
      <p>CT-BERT
RoBERTa-task3</p>
      <p>CT-BERT-task3</p>
      <p>Ensembling Models
TFIDF (Multi output clf)</p>
      <p>NLI Wikipedia</p>
      <p>RoBERTa</p>
      <p>CT-BERT
RoBERTa-task3</p>
      <p>CT-BERT-task3
Ensembling Models</p>
      <p>TFIDF (DT)
RoBERTa-task1+task2
CT-BERT-task1+task2</p>
      <p>RoBERTa-task3</p>
      <p>CT-BERT-task3
Ensembling Models</p>
      <p>Evaluation MCC
0.461
0.426
0.624
0.676
0.667
0.700
0.716
0.585
0.310
0.731
0.780
0.734
0.743
0.781
0.497
0.675
0.717
0.690
0.706
0.726
They remain an interesting alternative in the case where annotated
data is lacking: a few tweets of each class or, minimally, just the
definition of the classes are enough to provide some decent results.</p>
      <p>Looking at conspiracy theories, our worst results on Tasks 2
and 3 are about the Intentional Pandemic theory, even though it
is the most represented class in the dataset. Instead, our best
results are obtained with the Harmful Influence and New World Order
theories. One possible explanation is that both conspiracies can
be represented with very specific keywords (’5g’ or ’NWO’ for
example).</p>
      <p>We also performed late fusion ensembling through majority
voting with diferent combination of transformer-based models to
further improve our results on all the tasks. While this was our
best results on a stratified 5-fold cross-validation set, this was less
competitive on the test set.
4</p>
    </sec>
    <sec id="sec-7">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>In this paper, we presented three diferent methods to perform
COVID-19-related conspiracy theories detection in tweets. The first
approach consist of a combination of standard machine learning
based algorithm with TF-IDF. The second approach uses NLI with
transformer-based models and Wikipedia enrichment. The last
approach aims at fine-tuning transformer-based models on the
given dataset. Our best model obtains a MCC score of 0.775 for the
main task on the test set which outperforms by a large margin all
the other competitors in this challenge.</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work has been partially supported by CHIST-ERA within the
CIMPLE project (CHIST-ERA-19-XAI-003).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . (
          <year>2019</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CL/
          <year>1810</year>
          .04805
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Kowsari</surname>
            ,
            <given-names>Jafari</given-names>
          </string-name>
          <string-name>
            <surname>Meimandi</surname>
            , Heidarysafa, Mendu, Barnes,
            <given-names>and Brown. 2019. Text</given-names>
          </string-name>
          <string-name>
            <surname>Classification Algorithms</surname>
          </string-name>
          : A Survey.
          <source>Information</source>
          <volume>10</volume>
          ,
          <issue>4</issue>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Yinhan</given-names>
            <surname>Liu</surname>
          </string-name>
          , Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>RoBERTa: A Robustly Optimized BERT Pretraining Approach</article-title>
          . (
          <year>2019</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CL/
          <year>1907</year>
          .11692
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Loshchilov</surname>
          </string-name>
          and
          <string-name>
            <given-names>Frank</given-names>
            <surname>Hutter</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Decoupled Weight Decay Regularization</article-title>
          . (
          <year>2019</year>
          ).
          <source>arXiv:cs.LG/1711.05101</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Shervin</given-names>
            <surname>Minaee</surname>
          </string-name>
          , Nal Kalchbrenner, Erik Cambria, Narjes Nikzad, Meysam Chenaghlu, and
          <string-name>
            <given-names>Jianfeng</given-names>
            <surname>Gao</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Deep Learning Based Text Classification: A Comprehensive Review</article-title>
          . (
          <year>2021</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CL/
          <year>2004</year>
          .03705
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Martin</given-names>
            <surname>Müller</surname>
          </string-name>
          , Marcel Salathé, and
          <string-name>
            <given-names>Per E</given-names>
            <surname>Kummervold</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <string-name>
            <surname>COVIDTwitter-BERT: A Natural Language Processing Model to Analyse</surname>
          </string-name>
          COVID-19 Content on Twitter. (
          <year>2020</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CL/
          <year>2005</year>
          .07503
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Daniel Thilo Schroeder, Stefan Brenner, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>FakeNews: Corona Virus and Conspiracies Multimedia Analysis Task at MediaEval 2021</article-title>
          . In Multimedia Benchmark Workshop.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Daniel Thilo Schroeder, Petra Filkukova, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>WICO Text: A Labeled Dataset of Conspiracy Theory and 5G-Corona Misinformation Tweets</article-title>
          . In Workshop on Open Challenges in
          <source>Online Social Networks (OASIS).</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Victor</given-names>
            <surname>Sanh</surname>
          </string-name>
          , Lysandre Debut, Julien Chaumond, and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Wolf</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter</article-title>
          . (
          <year>2020</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CL/
          <year>1910</year>
          .01108
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>