<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>TweetClass: COVID-19 Vaccine Tweet Classification with scikit-learn</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Baivab Chakraborty</string-name>
          <email>baivab369@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Subhajit Srimani</string-name>
          <email>subhrasrimani2002@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Souvit Biswas</string-name>
          <email>souvitbiswas26@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amity University</institution>
          ,
          <addr-line>Kolkata, West Bengal, Kolkata 700135</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Natural Language Processing, COVID-19 Vaccine Tweets, Multinomial Naive Bayes</institution>
          ,
          <addr-line>Multi-Output</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Netaji Subhash Engineering College</institution>
          ,
          <addr-line>Kolkata, West Bengal, Kolkata 700152</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>as Unnecessary, Mandatory, Pharma</institution>
          ,
          <addr-line>Conspiracy, Political, Country, Rushed, Ingredients, Side-efect</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <fpage>15</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>Our team has proposed TweetClass as a solution in the FIRE 2023 AISoMe Track. We propose to utilize the sci-kit-learn library, which consists of several classifiers that can be used to categorize tweets Inefective and Religious. As the vaccination process continued globally to combat COVID-19, analyzing people's tweets seemed to provide valuable insights into their opinions on the entire vaccination episode. This enormous dataset, on correct utilization, can help the government create efective vaccination strategies in case of future pandemics. Our submitted model achieves a Macro-Fl score of 0.39 and a Metric Jaccard score of 0.46, earning us one of the top spots amongst other submissions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org
Classifier</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>The world faced its toughest challenge in the form of the COVID-19 pandemic. Over time,
vaccines have proven to be a safe and efective way to combat and eradicate infectious diseases.
As a result, there emerged a race to discover efective vaccines that could prevent the havoc of
COVID-19 and eventually, this led to the worldwide availability of these vaccines. However,
the discussions concerning vaccination progress, accessibility, eficacy, and side efects had
been ongoing, and people had both positive and negative opinions about it. Some took to social
media sites, like Twitter, to share their concerns regarding the vaccine and it proves beneficial
for the governments and health organizations, like WHO, to understand people’s thoughts
regarding the new COVID-19 vaccines. They look to use such insights to plan their future
strategies and encourage everyone to get fully vaccinated. It has been crucial to stop spreading
misinformation about these vaccines and appreciate the eforts of governments to have worked
to restrict the pandemic from spreading further. Twitter had also attempted to block tweets
that contained incorrect or misleading information about the virus, its preventative measures,
and treatments. The manual classification of tweets has proved time-consuming and prone to
nEvelop-O
LGOBE
https:// (B. Chakraborty); https://github.com/subhajitsrimani/TweetClass (S. Srimani); http:// (S. Biswas)
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
CEUR
Workshop
Proceedings
error. Therefore, there had been an urgent need to develop machine-learning models that could
assist in classifying tweets about the COVID-19 vaccines.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Task</title>
      <p>The Artificial Intelligence on Social Media (AISoMe) 2023 Track, organized as a part of the 15th
meeting of Forum for Information Retrieval Evaluation (FIRE) 2023, has tasked us to build an
efective multi-label classifier to label a social media post (particularly, a tweet) according to the
specific worries about vaccines that the post’s author stated, and we use this research to show
how we think the problem might be solved. The tweets consist of distinct concerns towards
vaccines owing to reasons such as the politics involved, potential side-efects of vaccines, etc.
To be precise, the tweets are categorized into the following 12 classes, described with instances:
• Side-effect: Suggests that there are side-efects of being vaccinated including death.</p>
      <p>Example: It starts. Please look for secure substitutes for this vaccine. Following patient
illnesses, the UK publishes an allergy warning for the Pfizer COVID-19 vaccination.
• Ineffective: Suggests that the vaccines may be inefective altogether. Example: What
exactly is the point? Are you giving your old and frail patients a second shot because Pfizer
claims the vaccine is worthless if administered outside of the recommended time frame?
• Religious: Suggests that there are religious restrictions against the vaccines. Example:</p>
      <p>A vaccine would go against my religion. The strongest defense available is the 91st Psalm.
• None: The tweet doesn’t coincide with any of the above reasons or with those not stated
above. Example: Total garbage. I could have taken the vaccine but will pass this one</p>
    </sec>
    <sec id="sec-4">
      <title>3. Related Work</title>
      <p>
        Information extraction from social media posts containing textual data has become an integral
part of social computing. Social media upholds the diverse opinions and expressions of people
regarding almost every aspect of this world and, the concerns about the COVID-19 vaccines is
no exception. Hence, the use of textual tweets as data to classify people’s stances on vaccines
employing machine learning methods like Multinomial Naive Bayes Classifier [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and
MultiOutput Classifier [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is very eficient and useful.
      </p>
      <sec id="sec-4-1">
        <title>3.1. Multinomial Naive Bayes</title>
        <p>Multinomial Naive Bayes Classifier is a probabilistic machine learning algorithm based on Bayes’
theorem. and assumes that features are independent and follow a multinomial distribution. It is
vital for roles like text categorization and sentiment analysis, where the frequency of words in
documents is essential for classification.</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Multi-Output Classifier</title>
        <p>A Multi-Output Classifier is a machine learning model designed to handle multiple target
variables simultaneously, making it suitable for multi-label or multi-task learning problems.
Instead of predicting a single output, it produces multiple outputs, each corresponding to
a diferent target variable. Multi-output classifiers are used in various domains, including
natural language processing, where multiple aspects or labels need to be predicted for a single
input. They extend traditional classification and regression techniques to tackle complex,
multidimensional prediction tasks.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Dataset</title>
      <p>A training dataset, containing 9921 tweets exhibiting the varied apprehensions of people about
the vaccines, is provided. Since it is used for training the model, the tweets in this file possess
both their IDs and their respective classes. The test dataset only contains 486 tweets along
with their respective IDs. Our approach makes the best use of the CAVES dataset [? ] through
various pre-processing techniques utilized and application of appropriate classifiers to develop
an eficient model.</p>
      <sec id="sec-5-1">
        <title>4.1. Trends in the Dataset</title>
        <p>On examining the training dataset, it can be stated that side-efects emerges as the most frequent
class of concerns with a whopping 38.4% portion, followed by inefective (16.9%), rushed (14.9%),
pharmaa (12.8%), mandatory (7.9%), unnecessary (7.3%), none (6.3%), political (6.3%), conspiracy
(4.9%), ingredients (4.4%), country (2.0%) and religious (0.7%).</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Pre-processing</title>
      <p>A textual tweet tends to comprise of plain text, special characters and emojis which must be dealt
with otherwise, this suppresses the performance of the model. Therefore, we pre-processed
the tweets to enhance the quality of the data for further processing and reduce any chances of
hampering performance. The following are the steps of pre-processing adopted in our model:
• Feature and Lable Extraction: Feature extraction selects and transforms the raw
data from the tweets into a set of features that can be used as input for the applied
machine learning algorithms. Similarly, the labels are also extracted which the model
will predict.
• Vectorization: The TF-IDF Vectorizer further transforms the text data into a numerical
format that the machine learning algorithms, the model uses, can understand. It also
removes the English stop words from the text that does not add any extra meaning to the
sentence on their own.
• Binarization: The Multi-Label Binarizer assists in the multi-label classification task as
required in case of our dataset where each tweet can belong to more than one class. It
converts these categories into binary labels (0 or 1) for each class.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Methodology</title>
      <p>6.1. Model
We have made a model that executes multi-label text classification. We have basically used
the scikit-learn library and its versatile features such as, the TF-IDF Vectorizer (converts the
collection of raw documents to a matrix of TF-IDF features) and the Multinomial Naive Bayes
Classifier (particularly useful in this case as the data set involves text data with discrete
features such as word frequency counts) to perform multi-label text classification. Moreover,
the implementation of Multi-Output Classifier employs one classifier per target (multi-target
classification) and Multi-Label Binarizer converts the labels into a binary matrix representation,
facilitating the multi-target classification process. Overall, the use of such versatile features
from scikit-learn fine tunes our model enabling eficient prediction of the test dataset.</p>
      <sec id="sec-7-1">
        <title>6.2. Experimental Setup</title>
        <p>The training data is split into training and validation sets in the ratio 9:1 and the data are shufled
prior splitting so that the fraction of instances of each class are preserved in both sets. We have
already explained in section 5 about the pre-processing techniques applied on the data and this
training data is used to fine tune our model while the validation data is used for evaluation.</p>
      </sec>
      <sec id="sec-7-2">
        <title>6.3. Prediction</title>
        <p>We have designed our model to analyze and predict the available test dataset, stored in “test_data”
data frame. In the dataset, the “tweet” column stores the extracted text data, and the
corresponding labels are stored in a new column named “pred_labels”. Then the data frame with
predicted labels is saved in a new CSV file named “prediction_file.csv” and the “tweet” column
is dropped from the data frame. We have also used the feature accuracy_score to calculate
accuracy of each predicted label. Finally, the model opens the CSV file with predictions and
reads the contents into a data frame named “result” and displays the output in the console.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>7. Evaluation</title>
      <p>AISoMe FIRE 23 Track results are evaluated using Macro-F1 score for primary evaluation and
Metric Jaccard for secondary evaluation in case of tie with Macro-F1 score. The result of our
two submitted automated run for the prediction of the test dataset is shown in Table 1. Our
model got the 34th and 38th rank based on our two submitted run files.</p>
    </sec>
    <sec id="sec-9">
      <title>8. Conclusion and Future Work</title>
      <p>This paper illustrates our TweetClass model. This model, upon construction and training,
performs multi- label text classification to successfully execute the prediction of an available
test dataset. The model uses the versatile features of scikit-learn to perform natural language
processing. We pre-processed our data through vectorization, binarization and splitting the
same then, performed feature and label extraction on the resultant data. This was followed
by training of our model using the training data and finally using the model for prediction
of the test dataset. We also look forward to further optimizing the model by exploring the
future scopes of machine learning and artificial intelligence. Our future endeavors include using
more domain specific models to improve the versatility of TweetClass thus, resulting in more
optimized outputs and higher accuracy scores.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. P. D.</given-names>
            <surname>Delizo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Abisado</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. I. P</surname>
          </string-name>
          . De Los Trinos,
          <article-title>Philippine twitter sentiments during covid-19 pandemic using multinomial naïve-bayes</article-title>
          ,
          <source>International Journal 9</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Read</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Martino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. M.</given-names>
            <surname>Olmos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luengo</surname>
          </string-name>
          ,
          <article-title>Scalable multi-output label prediction: From classifier chains to classifier trellises</article-title>
          ,
          <source>Pattern Recognition</source>
          <volume>48</volume>
          (
          <year>2015</year>
          )
          <fpage>2096</fpage>
          -
          <lpage>2109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Samad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ganguly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>Caves: A dataset to facilitate explainable classification and summarization of concerns towards covid vaccines</article-title>
          ,
          <source>in: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>3154</fpage>
          -
          <lpage>3164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Basu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>Overview of the fire 2023 track:artificial intelligence on social media (aisome)</article-title>
          ,
          <source>in: Proceedings of the 15th Annual Meeting of the Forum for Information Retrieval Evaluation</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>