<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MeVer team tackling Corona virus and Conspiracies using Ensemble Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga Papadopoulou</string-name>
          <email>olgapapa@iti.gr</email>
          <email>papadop@iti.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Symeon Papadopoulos</string-name>
          <email>papadop@iti.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Technologies Institute - ITI, CERTH</institution>
          ,
          <addr-line>Thessaloniki</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents the approach developed by the Media Verification (MeVer) team to tackle the task of Corona Virus and Conspiracies Multimedia Analysis Task at the MediaEval 2021 Challenge. We utilized ensemble learning and propose a two-stage classification approach that aims to overcome the challenge of the imbalanced and relatively small training dataset. We deal with the problem as binary classification in the first stage and in the second stage we predict the multi-labels. We experimented with fine-tuning pre-trained Bidirectional Encoder Representations from Transformers (BERT) and achieved a score of 0.294 in terms of the Matthews Correlation Coeficient (MCC), which is the oficial evaluation metric of the task. Additionally, leveraging on the proposed two-stage classification approach, we extracted a set of feature representations (BoW, TfIDF, embeddings) and classify them using traditional machine learning algorithms (Support Vector Machines, Logistic Regression) achieving in the best run a score of 0.292 of MCC.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The challenge of COVID 19-related misinformation has emerged
with the COVID 19 pandemic and continues to concern the
community about the amount of misinformation being disseminated
and its implications for many areas, such as health and society
[
        <xref ref-type="bibr" rid="ref15 ref6">6, 15</xref>
        ]. The need to develop methods to combat the dissemination
of COVID-related conspiracies triggered the organization of the
last year’s task of FakeNews: Coronavirus and 5G conspiracy in the
MediaEval 2020 Challenge [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and this year’s task of FakeNews:
Corona Virus and Conspiracies Multimedia Analysis [
        <xref ref-type="bibr" rid="ref10 ref12">10, 12</xref>
        ].
      </p>
      <p>A critical role in developing accurate methods for the automatic
detection of misleading tweets (and any other text or multimedia
item) plays the amount of annotated training samples. Due to the
relatively small training dataset provided to deal with the challenge
of detecting corona virus conspiracies, our approach follows a
twostage pipeline built on ensemble classification. In the first stage the
task is converted to a binary classification problem that classifies
the tweets in COVID Conspiracy tweets (involving both promoting
and discussing a conspiracy) and non-Conspiracy tweets. In the
second stage, the COVID conspiracy tweets are further classified
in promoting conspiracy (tweets that promotes, supports, claim,
insinuate some connection between COVID-19 and various
conspiracies) and discussing conspiracy (just mentioning the existing
various conspiracies connected to COVID-19). The final output of
the methods is a three-class prediction.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Several works have been introduced dealing with the detection and
verification of COVID 19-related misinformation utilizing machine
and deep learning approaches [
        <xref ref-type="bibr" rid="ref1 ref16 ref3">1, 3, 16</xref>
        ]. An overview of
CONSTRAINT 2021 Shared Tasks: Detecting English COVID-19 Fake
News and Hindi Hostile Posts [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] shows that BERT or its variations
was used for building the most successful models.
      </p>
      <p>
        A significant contribution to combat misinformation is the
creation of large enough annotated datasets which will serve to build
more accurate models. Patwa et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] released a dataset of 10,700
social media posts and articles of real and fake news on COVID-19.
In Shahi et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], the first multilingual cross-domain dataset of
5,182 fact-checked news articles for COVID-19 was introduced.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>
        We first utilized the approach that we had developed in last year’s
task of FakeNews: Coronavirus and 5G conspiracy [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We adapted
the method by corresponding the 5G conspiracy class to the Promote
Conspiracy class of this year’s task, the Other Conspiracy to Discuss
Conspiracy and the Non-Conspiracy to Non-Conspiracy. In short,
it is a two-step classification approach that first applies an initial
classification based on ensemble learning in order to provide a
firstlevel classification of the Conspiracy and Non-conspiracy tweets
and then a second step that predicts the classified Conspiracy tweets
whether they are promoting conspiracy or discussing a conspiracy.
For further details about the approach, the reader is referred to last
year’s working notes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>In addition, we run a set of complementary experiments based
on the proposed approach, i.e. leveraging on the two-stage
classification, experimenting with diferent feature representations and
classify them using machine learning algorithms. In the following,
we first describe how we deal with the imbalanced dataset, then
we list the diferent combinations of features and models that we
used in our experiments and we conclude with the results of the
proposed runs on the provided testing set of unseen tweets.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Dealing with the imbalanced dataset</title>
      <p>
        The provided dataset consist of 1,554 tweets in total for which
516 promotes COVID-related conspiracies (Promote), 271 discusses
COVID-related conspiracies (Discuss) and 767 do not refer to
COVIDrelated conspiracies (Non-Conspiracy). Johnson el al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] published
a survey on deep learning with class imbalance showing that
machine and deep learning approaches are essentially afected in terms
of prediction accuracy when trained with imbalanced samples. To
this end, we sub-sample training tweets of the majority classes in
order to balance the training sets and build the proposed
classiifers. Specifically, the classifiers of the first stage were trained with
540 samples of Conspiracy tweets (270 random samples of Promote
class and 270 random samples of Discuss class) and 540 samples of
Non-Conspiracy tweets. In the second stage, we trained a binary
classifier with positive class the Promote class and negative class
the Discuss class and a three-class model (Promote, Discuss,
NonConspiracy) by randomly selecting 270 samples from each class for
balance.
3.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Feature representation and machine learning algorithms</title>
      <p>
        In our additional experiments, we extracted five feature
representations: i) BoW: A simple and efective model for text representation
is the Bag-of-Words (BoW) Model. The model throws away all of
the order information in the words and focuses on the occurrence
of words in a tweet. ii) TFIDF: term frequency–inverse document
frequency reflects how important a word is to a tweet in a
collection of tweets. iii) BERT: We employ the bert-base-uncased version
of BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which is a compact transformer model, trained on
lower-cased English text. iv) Distil: We employ DistilBERT [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
which is a small, fast, cheap and light Transformer model trained by
distilling BERT base. v) Roberta: We employ the RoBERTa model
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which is built on BERT and and modifies key hyperparameters,
removing the next-sentence pretraining objective and training with
much larger mini-batches and learning rates. Each feature
representation is fed in an SVM and a LR and we conclude with a set of
multiple classifiers (Bow + SVM, BERT + LR, etc.)
3.3
      </p>
    </sec>
    <sec id="sec-6">
      <title>Runs</title>
      <p>The classifiers are trained on the samples presented in Section 3.1.
In the first stage the predictions of the binary classifiers and fused
using majority voting and provided to the second stage where the
ifnal predictions are calculated. We submitted four runs based on
diferent combinations of the feature representations and machine
learning algorithms.</p>
      <p>• Run 1: In the first stage we build a ensemble of binary
models combining all feature representations and both
machine learning algorithms. The predictions of the models
are fused using majority voting and in the second stage the
tweets classified as Conspiracy are further fed in an
ensemble of three-class models and binary classifiers (Promote
vs Discuss) again trained on all combinations.
• Run 2: The first stage is the same as with Run 1 and in
the second stage we fuse the predictions of binary
classiifers trained on Promote vs Discuss classes and the
Conspiracy vs Non-Conspiracy classes. For the Conspiracy vs
Non-Conspiracy models we use all training samples.
• Run 3: In the first stage we select a combination of feature
representations and machine learning algorithms which
derived as the best combination in terms of accuracy based
on cross validation. In the second stage we follow the
combinations of Run 2.
• Run 4: In the first stage, we select only the combinations of
BoW and BERT feature representations and LR and SVM.
For each combination, we train N models with sub-samples
of the training set. Similarity, the second stage fuses the
O. Papadopoulou et al.
predictions on models trained on the same combinations
on the three classes.</p>
      <p>
        • Run 5: This run is the method proposed in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-7">
      <title>RESULTS AND ANALYSIS</title>
      <p>
        The proposed approach of Papadopoulou et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] achieved the best
score (among our runs) of 0.294 in terms of MCC on the provided
testing set of unseen tweets for the task of FakeNews: Corona Virus
and Conspiracies Multimedia Analysis Task. In Table 1, the
evaluation results in terms of MCC on the unseen tweets are presented for
the five submitted runs. We observed that the accuracy of the four
additional runs compared to the approach of Papadopoulou et al.
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is slightly worse. The fact that traditional feature representation
such as BoW and TFIDF combined with emebeddings achieve
similar results to more complex deep learning approaches highlights
the challenge of the limited training data. We assume that with a
significantly larger training set the approach of Papadopoulou et
al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] will achieve much better predictions.
5
      </p>
    </sec>
    <sec id="sec-8">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>The proposed method achieves fairly accurate results in the task
of FakeNews: Corona Virus and Conspiracies Multimedia Analysis
Task. We followed our approach introduced in the MediaEval 2020
Challenge and based on the proposed pipeline we experimented
with diferent setups by extracting several feature representations
and using them to train traditional machine learning algorithms.
We noticed that fusing the predictions of diferent feature
representations and classification models we achieved almost the same
results as with fine-tuning pre-trained BERT, one of the most
popular transformer models. We observed that the limitation or the
relatively small training set afects the prediction accuracy of the
models negatively and augmentation techniques to create more
samples of the minority classes could be a step to improve the
predictions in future implementations.</p>
    </sec>
    <sec id="sec-9">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work is supported by the WeVerify project, which is funded
by the European Commission under contract number 825297.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Mabrook</surname>
            <given-names>S</given-names>
          </string-name>
          <string-name>
            <surname>Al-Rakhami and Atif M Al-Amri</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Lies kill, facts save: detecting COVID-19 misinformation in twitter</article-title>
          .
          <source>Ieee Access</source>
          <volume>8</volume>
          (
          <year>2020</year>
          ),
          <fpage>155961</fpage>
          -
          <lpage>155970</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Mohamed</surname>
            <given-names>K Elhadad</given-names>
          </string-name>
          , Kin Fun Li,
          <string-name>
            <given-names>and Fayez</given-names>
            <surname>Gebali</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <source>Detecting misleading information on covid-19. Ieee Access</source>
          <volume>8</volume>
          (
          <year>2020</year>
          ),
          <fpage>165201</fpage>
          -
          <lpage>165215</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Justin</surname>
            <given-names>M</given-names>
          </string-name>
          <string-name>
            <surname>Johnson and Taghi M Khoshgoftaar</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Survey on deep learning with class imbalance</article-title>
          .
          <source>Journal of Big Data</source>
          <volume>6</volume>
          ,
          <issue>1</issue>
          (
          <year>2019</year>
          ),
          <fpage>27</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Yinhan</given-names>
            <surname>Liu</surname>
          </string-name>
          , Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .
          <volume>11692</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Salman</given-names>
            <surname>Bin</surname>
          </string-name>
          <string-name>
            <surname>Naeem</surname>
          </string-name>
          , Rubina Bhatti, and
          <string-name>
            <given-names>Aqsa</given-names>
            <surname>Khan</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>An exploration of how fake news is taking over social media and putting public health at risk</article-title>
          .
          <source>Health Information &amp; Libraries Journal 38</source>
          ,
          <issue>2</issue>
          (
          <year>2021</year>
          ),
          <fpage>143</fpage>
          -
          <lpage>149</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Olga</given-names>
            <surname>Papadopoulou</surname>
          </string-name>
          , Giorgos Kordopatis-Zilos, and
          <string-name>
            <given-names>Symeon</given-names>
            <surname>Papadopoulos</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>MeVer Team Tackling Corona Virus and 5G Conspiracy Using Ensemble Classification Based on BERT</article-title>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Parth</given-names>
            <surname>Patwa</surname>
          </string-name>
          , Mohit Bhardwaj, Vineeth Guptha, Gitanjali Kumari, Shivam Sharma, Srinivas Pykl,
          <string-name>
            <surname>Amitava Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>Asif Ekbal</surname>
            , Md Shad Akhtar, and
            <given-names>Tanmoy</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Overview of constraint 2021 shared tasks: Detecting english covid-19 fake news and hindi hostile posts</article-title>
          . In International Workshop on Combating On line
          <article-title>Ho st ile Posts in Regional Languages dur ing Emerge ncy Si tuation</article-title>
          . Springer,
          <fpage>42</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Parth</given-names>
            <surname>Patwa</surname>
          </string-name>
          , Shivam Sharma, Srinivas Pykl, Vineeth Guptha, Gitanjali Kumari, Md Shad Akhtar, Asif Ekbal,
          <string-name>
            <surname>Amitava Das</surname>
            , and
            <given-names>Tanmoy</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Fighting an infodemic: Covid-19 fake news dataset</article-title>
          . In International Workshop on Combating On line
          <article-title>Ho st ile Posts in Regional Languages dur ing Emerge ncy Si tuation</article-title>
          . Springer,
          <fpage>21</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Konstantin</surname>
            <given-names>Pogorelov</given-names>
          </string-name>
          , Daniel Thilo Schroeder, Stefan Brenner, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          . Online,
          <volume>13</volume>
          -
          <fpage>15</fpage>
          December
          <year>2021</year>
          .
          <article-title>FakeNews: Corona Virus and Conspiracies Multimedia Analysis Task at MediaEval 2021</article-title>
          . In MediaEval 2021 Workshop.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Konstantin</surname>
            <given-names>Pogorelov</given-names>
          </string-name>
          , Daniel Thilo Schroeder, Luk Burchard, Johannes Moe, Stefan Brenner, Petra Filkukova, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>FakeNews: Corona Virus and 5G Conspiracy Task at MediaEval 2020</article-title>
          . In MediaEval 2020 Workshop.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Konstantin</surname>
            <given-names>Pogorelov</given-names>
          </string-name>
          , Daniel Thilo Schroeder, Petra Filkuková, Stefan Brenner,
          <article-title>and year=2021 Johannes Langguth</article-title>
          ,
          <source>booktitle=Proc. of the 2021 Workshop on Open Challenges in Online Social Networks</source>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          . WICO Text:
          <article-title>A Labeled Dataset of Conspiracy Theory and 5G-Corona Misinformation Tweets</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Victor</surname>
            <given-names>Sanh</given-names>
          </string-name>
          , Lysandre Debut, Julien Chaumond, and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Wolf</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter</article-title>
          . arXiv preprint arXiv:
          <year>1910</year>
          .
          <volume>01108</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Gautam</given-names>
            <surname>Kishore</surname>
          </string-name>
          Shahi and
          <string-name>
            <given-names>Durgesh</given-names>
            <surname>Nandini</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>FakeCovid-A multilingual cross-domain fact check news dataset for COVID-19</article-title>
          . arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>11343</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Sander van Der Linden</surname>
            , Jon Roozenbeek, and
            <given-names>Josh</given-names>
          </string-name>
          <string-name>
            <surname>Compton</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Inoculating against fake news about COVID-19</article-title>
          . Frontiers in psychology 11 (
          <year>2020</year>
          ),
          <fpage>2928</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Apurva</surname>
            <given-names>Wani</given-names>
          </string-name>
          , Isha Joshi, Snehal Khandve, Vedangi Wagh, and
          <string-name>
            <given-names>Raviraj</given-names>
            <surname>Joshi</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Evaluating deep learning approaches for covid19 fake news detection</article-title>
          . In International Workshop on Combating On line
          <article-title>Ho st ile Posts in Regional Languages dur ing Emerge ncy Si tuation</article-title>
          . Springer,
          <fpage>153</fpage>
          -
          <lpage>163</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>