<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Short Text Classification Using TF-IDF Features and Fast Text Learner</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zeshan Khan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Umar Naseer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muhammad Atif Tahir</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>zeshan.khan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>umar.naseer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>atif.tahir}@nu.edu.pk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>FAST School of Computing, National University of Computer and Emerging Sciences</institution>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>The spread of the COVID-19 is a challenge for the health sector. This pandemic created health and financial issues for the whole world. The medical experts are working for the diagnostics and reasons behind the COVID-19 disease and its spread. Some conspiracies are being spread related to the COVID-19 disease and its spread. Such conspiracies can be seen on social media including Twitter. In this research, the conspiracies of the COVID-19 have been analyzed from the public tweets. The tweets of the conspiracies have been ifltered from the tweets of the COVID-19 disease, symptoms, and other discussions related to the disease. The analysis of the COVID19 related tweets resulted into three conspiracy classes, the COVID19 tweets without any conspiracy and the conspiracies. A model is presented for the classification of tweets into three conspiracy classes with the Matthews Correlation Coeficient (MCC) of 0.294.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Social media became a source of information sharing from just
closed group chats. The information-sharing generated several
trust issues in the information shared on the social media platforms.
Currently, Twitter is one of the most used public post-sharing
platforms. There are a huge number of tweets being shared daily.
There may have several tweets containing misinformation.</p>
      <p>In the year 2019 a disease, COVID-19 badly damaged human
lives and the economy. There are several solutions proposed for the
treatment and spread control of the disease. The guidelines of the
health organizations are afected by the false information shared
by various individuals. Some of this false information is relating
COVID-19 spread with some technological inventions including
5G. A log of people is making a relationship between COVID-19
disease with the 5G technology towers. The time era of COVID-19
and the 5G technology are the same but that doesn’t show the one
as a cause of other or vise versa.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>The NLP domain is efective from the last decade for various
analyses of textual data. One of the domains of textual analysis is text
classification. The text classification becomes more challenging
when the provided text consists of very short documents. It’s very
dificult to build a context with the short textual document and the
benefit of the short document is the ease of processing.</p>
      <p>
        There is significant work available on the domain of text
classification and especially on short text classification. Some of the
researchers used various textual feature extraction techniques and
then applied classifiers to the textual features. The classifiers of
the SVM [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for the classification using TF-IDF as feature vector
[
        <xref ref-type="bibr" rid="ref12 ref9">9, 12</xref>
        ]. The SVM-based approaches are good in timely detection
or classification of the text with lower detection accuracy. There is
another group of researches done using neural network-based
approaches. The researchers used some pre-trained neural networks
like BERT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] then fine-tuned with the classification dataset [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Another type of research for this task is based on graph neural
networks (GNN). The GNNs are neural networks that can capture
the dependence of the graphs architecture by message passing
between perceptrons of the network. There are various variants of
GNN for the priority of usage in the domain including graph
convolutional network (GCN), graph attention network (GAT), graph
recurrent network (GRN), etc. The GNNs are good in detection
accuracy with a high time and computational cost [
        <xref ref-type="bibr" rid="ref13 ref3 ref7">3, 7, 13</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>The research is based on three diferent methodologies for a diverse
detection of the tweet-class. The Neural Network approaches are
performing well in the current era with a limitation of the high
availability of the data.</p>
      <p>
        The first approach that has been explored in this research is the
usage of cosine similarity between text vectors for the detection of
conspiracies in the tweet texts [
        <xref ref-type="bibr" rid="ref2 ref5">2, 5</xref>
        ]. The idea used in this approach
is to split the tweet texts into sentences and apply the learner for
classification. A similar learner is used to train on the whole tweet
as a single unit. The learner fasttext evaluated both the split and the
combined tweet for the MCC. The architecture is visually presented
in Figure 1.
      </p>
      <p>
        The second approach used in the research is the classification
of the TF-IDF vector [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This methodology converted tweets into
sentences to make the instances smaller. The TF-IDF features are
extracted from the sentences. The TF-IDF feature returned in a
feature vector of 1045 with most of the zero values. A feature
reduction technique of the principal component analysis was performed
to reduce the number of features for computation and accuracy
advantages [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The reduced features vectors were classified using
majority voting of some diverse learners including Decision Tree
Classifier, Linear Discriminant Analysis, and Logistic Regressions.
The architecture is summarised in Figure 2.
      </p>
      <p>
        The third algorithm for the detection of conspiracies in the tweet
is based on a fully connected neural network with the TF-IDF
features of the tweets [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The selection of the importance of words is
done using two phases of the removal of the word from tweets text.
In the first phase the categories of the words that are higher
important for the decision between conspiracy have been selected which
includes the Nouns, verbs, etc. The second phase of the selection of
the important terms is based on the Principal component analysis.
The PCA-based top 500 features have been selected to provide to
the neural network for ternary decision between conspiracy classes.
The detailed architecture of the algorithm can be seen in Figure 3.
The research is conducted using the dataset of the MediaEval 2021
under the task of FakeNews: Corona Virus and Conspiracies
Multimedia Analysis [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. The dataset is a set of tweets by various
Twitter accounts. The tweets consist of text of the tweet for the
detection task. There is some other information available with the
tweet for various objectives. The task of the conspiracy
classification needs only tweets text and conspiracy class for the training
data. The training data provided was of 1511 tweets with various
lengths from a few words to several sentences. The class distribution
of the tweets, class A, B and, C, was 754, 262, and 495 respectively.
The test set was comprised of 266 tweets for the detection of the
classes from three classes. The dataset shows there is a class
imbalance between the three provided classes. Another finding in the
dataset is of class decidability, the class B and C are much closer to
each other than the class A.
      </p>
    </sec>
    <sec id="sec-4">
      <title>RESULTS AND ANALYSIS</title>
      <p>Three approaches were designed to solve the challenge of the
conspiracy detection in the tweets. The first approach based on fasttext
classification was evaluated on the training data with 30% as
validation data. It was evaluated by training with various wordNgrams,
learning rates, dimensions and epochs. The best hyperparameters
for the fasttext resulted in 1-word gram with a learning rate of 0.7
and 800 dimensions. The model is trained for the 50 epochs due
to the limitation of the availability of resources. This approach
resulted in 0.89 accuracies on the validation dataset of 30% extracted
from the training dataset. The same model when applied to the
test dataset resulted in an MCC of 0.294. The second approach for
the computation of conspiracy was based on the TF-IDF vector
classification using a majority voting classifier. This methodology
used the dimensionality reduction technique of PCA with the
selection of the top 500 features out of 7147 features. The methodology
resulted in an accuracy of 0.56 with an MCC of 0.20 when executed
on the training dataset with 30% as validation dataset. The same
approach when applied to the test dataset it resulted in an MCC of
0.03. The third approach that is executed in the research is based on
a fully connected neural network on reduced TF-IDF features. In
this approach, the words of the tweet were selected based on their
categorical/ part of speech (POS) importance in decision making.
We applied various learners to several categories of the words in a
tweet e.g. the verb, nouns, adverbs, adjectives, etc. These learners
were guided about the decidability of the various POS. The selected
set of words is then used for the computation of the TF-IDF and
then PCA is used to reduce the feature vector length from 6898 to
500. The approach is applied to the validation data and the test data
and resulted in 0.17 and 0.07 MCC scores respectively.
6</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSION AND FUTURE WORK</title>
      <p>
        Short text classification is a challenging topic in the domain of
natural language processing. There are several challenges due to the
unavailability of the context of the sentence due to lesser sentences.
Various learners applied on the short text classification and fasttext
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] resulted in the best learner for tweet classification. The results
of fasttext guided the use of neural-network (NN) based approaches
or LSTM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] can give better results for the classification of the short
text.
      </p>
      <p>The results of the approaches above show that the neural
network (NN) based approaches can result in better detection accuracy.
The deep learning (DL) based approaches will be explored further
to improve detection accuracy. The data is limited and the nature
is very close to the various NLP datasets, So, the transfer learning
approaches can also be beneficial e.g. BERT can be used with the
pre-trained weights for a better understanding of the words and
relationships then the tweet data will fine-tune it to decide between
conspiracy classes.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Stanislav</given-names>
            <surname>Glebik</surname>
          </string-name>
          .
          <year>2021</year>
          . FAST TEXT. https://github.com/facebookresearch/ fastText/. [Online; accessed 25-November-2021].
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Abdullah</given-names>
            <surname>Hamid</surname>
          </string-name>
          , Nasrullah Shiekh, Naina Said, Kashif Ahmad, Asma Gul, Laiq Hassan, and
          <string-name>
            <surname>Ala</surname>
          </string-name>
          Al-Fuqaha.
          <year>2020</year>
          .
          <article-title>Fake news detection in social media using graph neural networks and NLP Techniques: A COVID-19 use-case</article-title>
          .
          <source>arXiv preprint arXiv:2012</source>
          .
          <volume>07517</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9</source>
          ,
          <issue>8</issue>
          (
          <year>1997</year>
          ),
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Armand</given-names>
            <surname>Joulin</surname>
          </string-name>
          , Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Fasttext. zip: Compressing text classification models</article-title>
          .
          <source>arXiv preprint arXiv:1612.03651</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Jure</given-names>
            <surname>Leskovec</surname>
          </string-name>
          , Anand Rajaraman, and Jefrey David Ullman.
          <year>2020</year>
          .
          <article-title>Mining of massive data sets</article-title>
          . Cambridge university press.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Hu</given-names>
            <surname>Linmei</surname>
          </string-name>
          , Tianchi Yang, Chuan Shi,
          <string-name>
            <given-names>Houye</given-names>
            <surname>Ji</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xiaoli</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Heterogeneous graph attention networks for semi-supervised short text classification</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          .
          <volume>4821</volume>
          -
          <fpage>4830</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A</given-names>
            <surname>Malakhov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Patruno</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S</given-names>
            <surname>Bocconi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Fake news classification with BERT</article-title>
          .
          <source>In Multimedia Evaluation Benchmark Workshop</source>
          <year>2020</year>
          , MediaEval
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Manfred</given-names>
            <surname>Moosleitner</surname>
          </string-name>
          , Benjamin Murauer, and
          <string-name>
            <given-names>Günther</given-names>
            <surname>Specht</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Detecting Conspiracy Tweets Using Support Vector Machines</article-title>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Konstantin</surname>
            <given-names>Pogorelov</given-names>
          </string-name>
          , Daniel Thilo Schroeder, Luk Burchard, Johannes Moe, Stefan Brenner, Petra Filkukova, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Fakenews: Corona virus and 5g conspiracy task at mediaeval 2020</article-title>
          . In MediaEval 2020 Workshop.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Konstantin</surname>
            <given-names>Pogorelov</given-names>
          </string-name>
          , Daniel Thilo Schroeder, Petra Filkuková, Stefan Brenner, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>WICO Text: A Labeled Dataset of Conspiracy Theory and 5G-Corona Misinformation Tweets</article-title>
          .
          <source>In Proc. of the 2021 Workshop on Open Challenges in Online Social Networks</source>
          .
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Thilo</surname>
          </string-name>
          <string-name>
            <surname>Schroeder23</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Evaluating Standard Classifiers for Detecting COVID-19 Related Misinformation</article-title>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Nguyen</given-names>
            <surname>Manh Duc Tuan and Pham Quang Nhat Minh</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>FakeNews Detection Using Pre-trained Language Models and Graph Convolutional Networks</article-title>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Lipo</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Support vector machines: theory and applications</article-title>
          . Vol.
          <volume>177</volume>
          . Springer Science &amp; Business Media.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Svante</surname>
            <given-names>Wold</given-names>
          </string-name>
          , Kim Esbensen, and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Geladi</surname>
          </string-name>
          .
          <year>1987</year>
          .
          <article-title>Principal component analysis</article-title>
          .
          <source>Chemometrics and intelligent laboratory systems 2</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>3</lpage>
          (
          <year>1987</year>
          ),
          <fpage>37</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>