<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Team Sigmoid at CheckThat!2021 Task 3a: Multiclass fake news detection with Machine Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Abdullah Al Mamun Sardar</string-name>
          <email>mamun35-1930@diu.edu.bd</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shahalu Akter Salma</string-name>
          <email>shahalu35-2315@diu.edu.bd</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Md. Sanzidul Islam</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Md. Arid Hasan</string-name>
          <email>arid.cse0325.c@diu.edu.bd</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Touhid Bhuiyan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Daffodil International University</institution>
          ,
          <addr-line>Dhaka</addr-line>
          ,
          <country country="BD">Bangladesh</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fake News Detection, Support Vector Machine</institution>
          ,
          <addr-line>Multinomial Naïve Bayes, LSTM</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Fake news is affecting our lives since the internet has become popular. Particularly, in this era of social media it is very easy to spread and be affected by fake news. In this work we have developed machine learning models which can classify a news claim into four classes. This work has been done under the competition of CheckThat!2021 task-3a. We have conducted our experiment on Check that lab's dataset. Our work has been done only on linguistic features. We have experimented both with traditional Machine Learning algorithms and Deep Learning algorithms. LSTM outperformed other traditional machine learning algorithms and with Adam optimizer LSTM gave a f1-macro score of 26.07%.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The internet has become one of the biggest part of our life. We are using it every day in our day to
day life. The usage of the internet and number of users are being increased every day. According to
datareportal.com [16] the number of internet user by April, 2021 is 4.72 billion (which is 60% of the
world’s population). Every day we read news articles, blogs and news content in various forms (e.g.,
images and videos). So it is easy to get distracted and deceived by any sort false representation of an
event or event which does not exist but depicted with a verified style. Fake news is affecting our life
and bringing damage to so many people. In 2016 US presidential Election fake news played a vital role.
Alexandre Bovet et. al. 2019 [22] showed 25% of the news in the time of US presidential election was
biased. In [15], data shows that the number of blog posts appear only in WordPress is 70 million each
month. These findings only show that the number of potential fake news are not so little. And we can
easily be misled by those false news. During the COVID-19 pandemic we have seen so many fake news
spread in different communities all around the world. In India several fake news spread during this
pandemic which created confusion about COVID-19 in the community. In [17] summarized the
COVID-19 related fake news in India. Another fake news spread in India which claim the people who
are taking COVID-19 vaccine may die within two years [18]. In Bangladesh several fake news created
a huge confusion among students during this pandemic. Several fake news claimed the higher secondary
examination will be taken place soon in May-June, 2020. In 2019, A mother was killed in Bangladesh
as a result of a fake news which claims children’s head are being used in the construction of Padma
Bridge but later investigation did not find any clue of this claim and they also found that the mother
was innocent [20]. In October, 2020 a fake news claimed that France footballer Paul Pogba left France
National Football team [19]. Several studies have been done to classify and categorize fake news. B
Bharali et. al. 2017 [21] addressed six categories of fake news: “1. Disinformation, 2. Propaganda 3.</p>
      <p>2021 Copyright for this paper by its authors.
Hoaxes 4. Satire/Parody 5. Inaccurate 6. Partisanship”.</p>
      <p>To detect fake news automatically we need to use machine learning techniques because the amount
of articles appearing on the internet every single day is impossible to classify with traditional
approaches. The sub-domain of machine learning which deals with text classification is called Natural
Language Processing. Using NLP techniques, we can detect whether a given news article is fake or real
or some further classification. Researchers have done a lot of work in this field. Some have worked
only with text and some others have worked with additional entities like news sources and user opinions.
After analyzing those previous work, we have come up with two research questions to conduct our
study.</p>
      <p>RQ1: Can we identify the impact factor in linguistic analysis based fake news detection?
RQ2: Between traditional machine learning and deep learning which one gives better performance?
The rest of the paper is structured as follows: In the literature review section we have discussed how
others have conducted their research and which approach they have found as best performer. Next, the
research methodology section will describe step by step how we have conducted our experiment. And
then in the result analysis section we will compare the result of different methodologies we have used.
The rest three sections are conclusion, acknowledgement and reference.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature Review</title>
      <p>
        Fake news detection got wide attention in the machine learning research community after the US
election in 2016. T S Reshmi et. al. 2021 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] has done a work on fake news detection using source
information. They addressed some common features (Lexical, Syntactic, Visual, Statistical, Users, Post,
Network) which are being used in content based classification. Rohit Kumar et. al. 2021 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] worked on
fake news detection by using deep neural network. They have experimented on real world dataset
(Buzzfeed and Politifact). They have divided their classification process into three parts. One for text
based classification, the other two are social context based and combination of these two. They have
found their best performing model with deep neural network. Anshika Choudhary et. al. 2020 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
worked on linguistic features to detect whether a news is fake or real. They have considered four
linguistic features in their study as follows: syntax-based, sentiment-based, grammatical and
readability-based evidence. With traditional machine learning based ensemble methods they got an
accuracy of 72% but sequential neural network based model outperforms the previous one and got an
accuracy of 86%. Thomas Felber et. al. 2021 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] worked on a content based experiment. They have
considered unique word count, average word per post and average character per post. They have
experimented with different machine learning algorithms and with support vector machine (SVM) they
got the highest accuracy which is 95.70%. Elena Shushkevich et. al. 2021 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] worked on an ensemble
method which performed better than a single algorithm based model. Mohammad Hadi Goldani et. al.
2020 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] worked on fake news detection with a capsule neural network. Their dataset contains two types
of news. One which texts are small in length and the other which texts are medium or large in length.
They have implemented three types of word embedding techniques (Static, Non-static and multichannel
word embedding). Gautam Kishore Shahi et. al. 2020 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] build a dataset for fake news detection. The
specialty of their dataset is that it contains more than one language (English, Spanish, French,
Portuguese, Hindi, Turkish, Italian, Chinese, Croatian, Telugu). Typical fake news datasets are most of
the time built on only one language. They have done a benchmark study on their dataset and with the
BERT based classification model they got a F1-score of 76%. The table given below is a summary of
some published work on fake news detection tasks.
      </p>
      <sec id="sec-2-1">
        <title>Description of different types Kaggle Fake BM25, of dataset for fake news news dataset, Fake Space detection. Comparison on news challenge,</title>
      </sec>
      <sec id="sec-2-2">
        <title>Vector Model, [9] [10]</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Research Methodology</title>
      <p>
        There are various types of research methodologies in Natural Language Processing for fake news
detection. To detect fake news, researchers usually follow some particular methodologies. Some use
only a content and context based approach, some have taken into account the user opinion on social
media on the same news from various users and some other researchers considered the source of the
news [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However, in this research we have only worked with a linguistic based approach as
our dataset only contains news titles and main news content. We have conducted our research in the
following steps.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3.1. Data Analysis</title>
      <p>As this work has been done under CheckThat!2021, they have provided the dataset. The dataset
contains four columns: public_id, title, text and our rating. ‘our rating’ basically contains the classes we
will classify. A snippet of the dataset has been given below.
Our rating contains four classes: False, Partially false, True and Other. A distribution of these classes
in the figure below:
When it comes to news content based fake news detection, some researchers considered only main news
content, not the title and some researchers took account of both title and main news content. In our
research we have applied both approaches and shown the performance analysis in the result analysis
section.</p>
    </sec>
    <sec id="sec-5">
      <title>3.2. Text Preprocessing</title>
      <p>In the preprocessing step, first, we have removed number and punctuation from our dataset. Then
we removed one and two-character length words. We used python regular expressions for this task.
Machine Learning algorithms consider “Word” and “word” as separate entities and this is a problem in
experiment because they both have the same meaning. To avoid this problem, we have converted the
whole dataset into lowercase character. After doing these steps we have removed stop words. We have
used the nltk library to remove stopwords. After then we did word level tokenization on both title and
text using nltk.word_tokenize. For stemming we have used Portstemmer() and for lemmatizaiton we
used WordNetLemmatizer(). After completing the preprocessing steps we have started feature
extraction which we will discuss in the next section.</p>
    </sec>
    <sec id="sec-6">
      <title>3.3. Feature Extraction</title>
      <p>Feature extraction is mandatory for machine learning tasks. It helps to reduce the training time by
reducing dimensionality. In the feature extraction phase we have used two feature extraction techniques
which are common in natural language processing. We used TF-IDF and CountVectorizer in our work.
Term Frequency-Inverse Document Frequency (TF-IDF) is a statistical measure commonly used in
natural language processing which evaluates how relevant a word is to a document in a collection of
documents. This is done by multiplying two metrics: how many times a word appears in a document,
and the inverse document frequency of the word across a set of documents [23]. CountVectorizer is
used to transform a given text into a vector on the basis of the frequency (count) of each word that
occurs in the entire text.</p>
    </sec>
    <sec id="sec-7">
      <title>3.4. Research Model</title>
      <p>In this research we have followed a research model which has been presented below. After getting
the dataset the first thing we did is preprocessing it and then we have extracted features using some
common feature extraction technique and then used some traditional machine learning algorithm and a
deep learning algorithm (LSTM) to classify the news into one of the four classes given in the ‘our
rating’.</p>
      <sec id="sec-7-1">
        <title>News Text</title>
      </sec>
      <sec id="sec-7-2">
        <title>Preprocessing</title>
      </sec>
      <sec id="sec-7-3">
        <title>Feature Extraction</title>
      </sec>
      <sec id="sec-7-4">
        <title>ML/DL Algorithms</title>
      </sec>
      <sec id="sec-7-5">
        <title>Result</title>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>4. Result Analysis</title>
      <p>In our work we have experimented with both traditional ML algorithms and deep learning
algorithms. We have conducted our experiment in two different ways. In the first approach, we have
conducted the experiment by concatenating both news title and main news content. And in the second
approach we have experimented only with news content without taking the news title into account. For
both approach we took 80% for training and 20% for testing. In the first approach, by using count
vectorizer and tf-idf transformer in the pipeline we got a f1-macro score of 35% with logistic regression.
And by limiting the maximum feature to 1000, we were able to increase the performance by 5%.
Support Vector Machine with linear carnal gave 38% f1-macro score. In the first approach, among all
traditional machine learning algorithms which we have experimented, Multinomial Naïve Bayes
classifier gave the best f1-macro score (43%) on the training dataset.</p>
      <p>In second approach, Random Forest Classifier and XGBClassifier performed better than the first
approach and the rest four algorithm did not do well than the first approach. A performance comparison
has been given below.</p>
      <p>F1-macro
score
(One)
We have got our best performing algorithm with deep learning. We applied our second approach with
LSTM and it performed way better than traditional ML algorithms, with softmax activation function
and adam optimizer we got a validation accuracy of 99%. But we have found our model was over fitted
after the CheckThat!2021 result publication and the performance was very poor (26.07% f1-macro
score). We are working on our LSTM based model to overcome this problem. For task-3a, the best
score was 83.76% and the least score was 13.47%.</p>
    </sec>
    <sec id="sec-9">
      <title>5. Conclusion</title>
      <p>In this work we tried to build a machine learning model to classify fake news. We have
experimented with both machine learning and deep learning based models. Our Machine Learning
based model was outperformed by deep learning based model on training data. That’s why we submitted
it to the completion but for the overfitting problem it did not perform well on test data. We are working
on our best performing deep learning model to make it better and improve the performance.</p>
    </sec>
    <sec id="sec-10">
      <title>6. Acknowledgements</title>
      <p>This is unbelievable support we’ve got from some of the faculty members and seniors of DIU NLP
and ML Research Lab to continue our whole research flow from the beginning. We acknowledge Dr.
Firoj Alam for guiding us and informing the secretes of research workshops. Also, we’re thankful to
Daffodil International University for the workplace support and the academic collaboration in some
cases. Dr. Touhid Bhuiyan and Dr. Sheak Rashed Haider Noori also supported us with guidance,
motivation, and advocating in institutional supports. Lastly, for sure we’re thankful to our God always
for every fruitful work with our given knowledge.</p>
    </sec>
    <sec id="sec-11">
      <title>7. References</title>
      <p>Khan, Junaed Younus, et al. "A benchmark study on machine learning methods for
fake news detection." arXiv preprint arXiv:1905.04749 (2019).</p>
      <p>Roy, Arjun, et al. "A deep ensemble framework for fake news detection and
classification." arXiv preprint arXiv:1811.04670 (2018).</p>
      <p>Shahi, Gautam Kishore. "AMUSED: An Annotation Framework of Multi-modal Social
Media Data." arXiv preprint arXiv:2010.00502 (2020).</p>
      <p>Shahi, Gautam Kishore, Anne Dirkson, and Tim A. Majchrzak. "An exploratory study
of covid-19 misinformation on twitter." Online Social Networks and Media 22 (2021):
100104.</p>
      <p>https://hostingtribunal.com/blog/blog-posts-per-day/#gref
https://datareportal.com/global-digital-overview
https://www.dw.com/en/india-covid-misinformation/a-57414876
https://www.hindustantimes.com/india-news/will-you-die-within-2-yrs-after-takingvaccine-centre-busts-fake-news-post-101621961181280.html</p>
      <p>https://www.thedailystar.net/sports/football/news/fake-unacceptable-news-pogbaslams-rumors-him-quitting-france-1984513</p>
      <p>https://bdnews24.com/bangladesh/2019/07/24/no-one-listened-to-what-she-saidwitnesses-recount-lynching-of-a-mother-in-bangladesh</p>
      <p>Bharali, Bharati, and Anupa Lahkar Goswami. "Fake news: Credibility, cultivation
syndrome and the new age media." Media Watch 9.1 (2017): 118-130.</p>
      <p>Bovet, Alexandre, and Hernán A. Makse. "Influence of fake news in Twitter during the
2016 US presidential election." Nature communications 10.1 (2019): 1-14.
https://monkeylearn.com/blog/what-is-tf-idf/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Kaliyar</surname>
            ,
            <given-names>Rohit</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
            , Anurag Goswami, and
            <given-names>Pratik</given-names>
          </string-name>
          <string-name>
            <surname>Narang</surname>
          </string-name>
          .
          <article-title>"DeepFakE: improving fake news detection using tensor decomposition-based deep neural network."</article-title>
          <source>The Journal of Supercomputing 77.2</source>
          (
          <year>2021</year>
          ):
          <fpage>1015</fpage>
          -
          <lpage>1037</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Choudhary</surname>
            , Anshika, and
            <given-names>Anuja</given-names>
          </string-name>
          <string-name>
            <surname>Arora</surname>
          </string-name>
          .
          <article-title>"Linguistic feature based learning model for fake news detection and classification</article-title>
          .
          <source>" Expert Systems with Applications</source>
          <volume>169</volume>
          (
          <year>2021</year>
          ):
          <fpage>114171</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Felber</surname>
            ,
            <given-names>Thomas. "</given-names>
          </string-name>
          <article-title>Constraint 2021: Machine Learning Models for COVID-</article-title>
          19
          <source>Fake News Detection Shared Task." arXiv preprint arXiv:2101.03717</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Reshmi</surname>
            , T. S., S. Daniel Madan Raja, and
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Priya</surname>
          </string-name>
          .
          <article-title>"Fake News Detection Using Source Information</article-title>
          and
          <string-name>
            <given-names>Bayes</given-names>
            <surname>Classifier</surname>
          </string-name>
          .
          <source>" IOP Conference Series: Materials Science and Engineering</source>
          . Vol.
          <volume>1084</volume>
          . No.
          <article-title>1</article-title>
          .
          <string-name>
            <given-names>IOP</given-names>
            <surname>Publishing</surname>
          </string-name>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Shushkevich</surname>
            , Elena,
            <given-names>and John Cardiff.</given-names>
          </string-name>
          "TUDublin team at Constraint@
          <fpage>AAAI2021</fpage>
          --
          <source>COVID19 Fake News Detection." arXiv preprint arXiv:2101.05701</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Goldani</surname>
            ,
            <given-names>Mohammad</given-names>
          </string-name>
          <string-name>
            <surname>Hadi</surname>
            , Saeedeh Momtazi, and
            <given-names>Reza</given-names>
          </string-name>
          <string-name>
            <surname>Safabakhsh</surname>
          </string-name>
          .
          <article-title>"Detecting fake news with capsule neural networks</article-title>
          .
          <source>" Applied Soft Computing</source>
          <volume>101</volume>
          (
          <year>2021</year>
          ):
          <fpage>106991</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Shahi</surname>
            ,
            <given-names>Gautam</given-names>
          </string-name>
          <string-name>
            <surname>Kishore</surname>
            , and
            <given-names>Durgesh</given-names>
          </string-name>
          <string-name>
            <surname>Nandini. "FakeCovid--A Multilingual</surname>
          </string-name>
          Cross-domain
          <source>Fact Check News Dataset for COVID-19."</source>
          arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>11343</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Ghosh</surname>
            , Souvick, and
            <given-names>Chirag</given-names>
          </string-name>
          <string-name>
            <surname>Shah</surname>
          </string-name>
          .
          <article-title>"Towards automatic fake news classification</article-title>
          .
          <source>" Proceedings of the Association for Information Science and Technology 55.1</source>
          (
          <year>2018</year>
          ):
          <fpage>805</fpage>
          -
          <lpage>807</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Karimi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hamid</surname>
          </string-name>
          , et al.
          <article-title>"Multi-source multi-class fake news detection</article-title>
          .
          <source>" Proceedings of the 27th international conference on computational linguistics</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Oshikawa</surname>
            , Ray,
            <given-names>Jing</given-names>
          </string-name>
          <string-name>
            <surname>Qian</surname>
          </string-name>
          , and William Yang Wang.
          <article-title>"A survey on natural language processing for fake news detection." arXiv preprint arXiv:</article-title>
          <year>1811</year>
          .
          <volume>00770</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>