<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identifying Situational Information during Mass Emergency</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sumit Anand</string-name>
          <email>sumit.anand@uem.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehuly Chakraborthy</string-name>
          <email>mehuly25@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diptaraj Sen</string-name>
          <email>diptaraj.work@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Engineering and Management Kolkata</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the advent of Natural Language Processing, what finds itself in much use is analysis. This research paper finds itself in reference to the same that enables it in analysing sentiments of a text. The tasks that were covered in working with NLP includes - firstly, diferentiating tweets on the basis of claims and facts, and secondly to create an efective classifier that finds out if a tweet is anti-covid vaccine, pro-covid vaccine or neutral. The beauty of our paper resides in the fact, that we have hit high end accuracies without using hefty algorithms, namely 93% for the first task using Random Forest and 45.4% for the second task using BERT's Algorithm. Our accuracies are the best among all the teams working on the same tasks, which deepens the efect that this paper resonates. The details of the IRMiDis 2021 data challenge have been discussed elaborately here, and we hope our paper marks its significance by virtue of its own merit.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Random Forest</kwd>
        <kwd>BERT's Algorithm</kwd>
        <kwd>Micro-blogging</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>classifier</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In periods of such dire needs, where mankind is at loss of life, humanity struggles to stay put
in every way possible. Computer Science plays a massive role here by trying to make things
more accessible to people, in eficient and sophisticated ways. Social media posts are one of
the most important bullets that help coders analyse how and what need to be done in case of
massive worldwide emergencies. Our tasks have led us to a discovery of what people think
about covid vaccines, leading us to understand what has to be further coped up with to increase
awareness in society, and also to create a model that separates claims and facts. Taking help
from twitter data has solved our purpose meticulously, and we say this with immense confidence
that micro-blogging will be used further in plethora of fields that Computer Science has blessed
us with. Both our tasks are but an analysis that micro-blogging has provided us with. To explain
more, let’s sketch out a brief overview of tasks 1 and 2. For the first task, we were required to
diferentiate facts and non-facts, the data sets being extracted from twitter. The second task
consisted of 2 data sets- the train and the test- that were also extracted from micro-blogging
sites, where a classifier was to be built to separate opinions-for, against and neutral-regarding
the covid vaccine. With the help of two eficient algorithms, namely Random Forest [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and
BERT’s [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], we have conjured accuracies of 0.93 and 0.45 on tasks 1 and 2 respectively. The
paper, hence describes intricately the procedures that we’ve undertaken to get the highest
accuracies among the other teams working on the same IRMiDis 2021 data challenge.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Tasks</title>
      <sec id="sec-2-1">
        <title>The tasks were pretty simple, yet immensely engaging.</title>
        <p>The first task involved the data set of 11000 tweets from twitter related to the Nepal Earthquake
in April 2015. Along with the dataset, sample of few claims or fact-checkable tweets and
nonfact-checkable tweets were also provided in text format. What needed to be done was to identify
claims and fact checkable tweets.</p>
        <p>Examples of claims:
1. @mashable some pictures from Norvic Hospital *A Class Hospital of nepal* Patients have
been put on parking lot.</p>
        <p>2. @ Refugees: UNHCR rushes plastic sheeting and solar-powered lamps to Nepal earthquake
survivors [url]</p>
        <p>Example of non-fact checkable tweets:
1. Students of Himalayan Komang Hostel are praying for all beings who lost their life after
earthquake!!! Please do...[url]
2. We humans need to come up with a strong solution to create earthquake proof zone’s.</p>
        <p>The second task was in reference to the present scenario of covid wrenched pandemic that
has kept everyone in tatters of their own luck. Two data files were provided namely the Train
dataset and the Test data set. The train dataset contains stances of tweets towards COVID-19
vaccines, crawled between November-December 2020, whereas the test dataset contains tweets
between March-December 2020 with various vaccine related keywords. What needed to be
done was to identify how many people are still skeptical about the covid vaccine. Hence a
classifier was to be built for 3 class classification as stated below:
1. AntiVax - the tweet is against the use of vaccines.
2. ProVax - the tweet supports / promotes the use of vaccines.</p>
        <p>3. Neutral - the tweet does not have any discernible sentiment expressed towards vaccines or
is not related to vaccines</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>
        We have used the datasets provided to us by IRMiDis FIRE 2021 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which include a wide range
of data ranging from Nepal Earthquake in 2015 to tweets deciphering how many people hold
negative views regarding the covid vaccine. For the first task, we have a dataset containing
around 11,000 microblogs [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (tweets) from Twitter that were posted during the Nepal earthquake
in April 2015. Along with the dataset, sample of few claims or fact-checkable tweets and
nonfact-checkable tweets are also provided as text files, in the following format– Tweetid &lt;||&gt;
Tweettext
      </p>
      <p>Example:
592568567247212544&lt;||&gt;RT @NewEarthquake: 4.7 earthquake, 25km S of Kodari, Nepal. Apr
26 13:21 at epicenter (21m ago, depth 10km).</p>
      <p>For the second task, IRMiDis FIRE 2021 has contributed 2 datasets – namely the training and
the testing. To describe more of it –</p>
      <p>The training data set consisted of stances of tweets towards COVID-19 vaccines, crawled
between November-December 2020. From this dataset 2,792 crawled tweets texts along with the
tweet-IDs and the classes were produced for the same. The testing data set, on the other hand,
had in it tweets between March-December 2020 with various vaccine-related keywords. There
were tweets annotated by three crowd workers. For 1600 tweets, there was at least majority
agreement, i.e., at least 2 out of the 3 annotators provided the same label. The test dataset is
formed of these 1600 tweets; each tweet was tagged in with the tweet ID and the tweet text.
Referring to these master datasets, the IRMiDis 2021 data challenge was accomplished with
great results, that helped us analyse skilfully with no hindrance in the least.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>This section, describes in details the process that we followed to get to the desired results. We
have tried to apply Random Forest and BERT’s Algorithm to build up the basics of our task, and
have been successful in leading out predictions at higher accuracies.</p>
      <sec id="sec-4-1">
        <title>4.1. Task 1: Classifying tweets into Facts and Non-facts</title>
        <p>The first task, as mentioned before, asks us to diferentiate tweets into facts and non- facts.
Hence, to perform this, we have figured out a proper model algorithm that gives us an accuracy
of 93%. This task can be divided into 3 non lapping phases, namely – Preprocessing, Feature
Selection and Model Selection.
4.1.1. Preprocessing
The first phase helps us clear unnecessary data that hold no relevance to our task of interest.
The dataset, hence, was pre-processed before moving on to further phases. We removed links
that were present in the data list, along with stop words and tweeter id. Also, user id and
punctuations of any kind were removed, hence what was left was a dataset that simply has
letters and numbers. Then for the ease of our working, we converted the texts to lower case.
This was our pre-processed dataset.
4.1.2. Feature Selection
Proceeding to the next step, we come to Feature Selection, where we try to extract certain
features from our newly reduced dataset. This helps us to understand the dataset in a clearer
fashion. Here, after pre-processing, we have added an additional column which consists of 1
and 0. The assigning of 1 and 0 is done in the following way – if a tweet has a total number of
digits to be 5 or more, we have assigned a 1 in the respective column, while for total digits 4 or
lesser, it has a designated 0 in the column.</p>
        <sec id="sec-4-1-1">
          <title>For example,</title>
          <p>1. ‘Nepal is the only Hinu country in the world, we need to protect and provide relief in this
crisis, hats of to.’</p>
          <p>2. ‘:Earthquake helpline at the Indian embassy in Kathmandu: +977 98511 07021, +977 98511
35141’</p>
          <p>
            Here in the first example, the number of digits is 0, hence the assigned value is 0 and it is a
non-fact. In the next example, we see there are 14 digits in total, so likewise the assigned value
in column is 1, and also it is a fact. The reason we thought about this criterion is because, while
manually inspecting the tweets, we realized that most of the facts are those that involve digits.
So now, we experimented with our model by feeding diferent inputs of the digit count starting
from 2,3,4 and 5. What we came to understand is that if the total count of digits is more than or
equal to 5, it has a greater tendency of being a fact. Hence, we implemented this, and proceeded
to the final phase.
4.1.3. Model Selection
This is the final stage of our task, that involves selecting a proper model to train our data
to. After experimenting with many models and algorithm, we found out that Random Forest
Classifier gives us the best results. Training our data list after Feature extraction in Random
Forest, gives us an accuracy of 0.93. We have incorporated five-fold cross validation for better
results. Hence, we could successfully conclude our first task, by eficiently diferentiating tweets
on the basis of facts and non-facts [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ].
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Task 2: 3-class classification on tweets regarding their stance towards</title>
      </sec>
      <sec id="sec-4-3">
        <title>COVID-19 vaccines.</title>
        <p>
          As explained previously, task 2 asks us to build an efective classifier for 3-class classification on
tweets regarding their stance towards COVID-19 vaccines. The 3 classes are namely – AntiVax
(against covid vaccine), ProVax (for covid vaccine) and Neutral. Hence, we have to train our
training data through an algorithm, that will provide us with higher accuracies while tested
with the testing data. Like Task 1, here too we can divide the process in 3 non-overlapping
phases – Preprocessing, Feature Selection and Model Selection.
4.2.1. Preprocessing
Now to get rid of redundant data, we have performed certain functions that helped us lay more
focus on our required task. We started by removing all the usernames and hashtags, removing
URLs and links and all kinds of special characters, emojis and emoticons. The remaining texts
were converted to lowercase to ease our work. Hence our data was pre-processed successfully,
and we led onto the next step.
4.2.2. Feature Selection
In this step, we have tried to extract certain features from our pre-processed data so that training
the data becomes easy. The classes have been level-coded such that – AntiVax=-1 ProVax=1
Neutral=0. Along with this, we have tokenized the tweets for better working.
4.2.3. Model Selection
After much speculation, we applied a pretrained BERT Model to train our data, that has proved
to be immensely efective. Our BERT model is called ‘distilBERT-base-uncased’ that is a
transformers model, smaller and faster than BERT, which was pretrained on the same corpus in a
self-supervised fashion, using the BERT base model as a teacher. This means it was pretrained
on the raw texts only, with no humans labelling them in any way (which is why it can use
lots of publicly available data) with an automatic process to generate inputs and labels from
those texts using the BERT base model. On applying this, the accuracy that we received while
testing our training data is around 45.4%, which is a good accuracy compared to all the other
algorithms that we have tried. Hence, this led us to the end of the second task [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Evaluation</title>
      <p>This section displays the output of all our algorithms and concepts that we have applied to
complete our required tasks. Hence, this consists of our accuracies, precision, recall, MAP, MAP
Overall and macro-F1 score that we have retained after training our data through the models
which we decided to work on. For both tasks our results were inclined towards the higher ends,
that invariably imply that our tasks were a success.</p>
      <sec id="sec-5-1">
        <title>5.1. Task 1: Classifying tweets into Facts and Non-facts</title>
        <p>Task 1 was successfully completed by applying Random Forest Classifier to our pre-processed
data. The accuracy was 93% which is a very good output based on our dataset. Equivalently our
other results have borne amazing outcomes namely a precision (out of 100) of 0.9100, recall (out
of 1000) of 0.2165, MAP (out of 100) of 0.0669, and a MAP Overall of 0.1543. The facts and the
non-facts were eficiently diferentiated, and the results yielded were more than satisfactory.</p>
        <p>Team ID Precision@100 Recall@1000 MAP@100 MAP Overall</p>
        <p>ByteCrackers 0.9100 0.2165 0.0669 0.1543</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Task 2: 3-class classification on tweets regarding their stance towards</title>
      </sec>
      <sec id="sec-5-3">
        <title>COVID-19 vaccines.</title>
        <p>The second task required us to build up an efective model to classify between 3 classes namely
– ProVax, AntiVax and Neutral. We’ve applied a pretrained BERT’s model to train our data into,
called ‘distilBERT-base-uncased’ that has yielded us an accuracy of 45.4%. Likewise, our
macroF1 score is 0.440, which is a good outcome based on our data. The model works dexterously and
hence, we have successfully completed our second task.</p>
        <p>Team ID Accuracy macro-F1 Score</p>
        <p>ByteCrackers 0.454 0.440</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>
        This paper holds an amalgamation of two beautiful working algorithms – the Random Forest
Classifier and the DistilBERT Algorithm [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. As mentioned in the paper, we have achieved
higher end accuracies in both the tasks. This IRMiDis 2021 data challenge has been much more
than a learning experience, for as coders, we have had a working idea about the surrounding
society, and their thoughts about issues troubling the nation. Programmers can further utilize
this data to procreate something better to treat these societal issues [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Our tasks can be further
modified to make it better and eficient. Theoretically we have thought about implementing
clustering along with BERT’s for task 2, and extracting more eminent features other than digits
to complete task 1. We have now a much wider grip on Machine Learning and we hope to
implement our theories to these tasks to build well-structured models that perform much higher
accuracies.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Biau</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scornet</surname>
            <given-names>E.</given-names>
          </string-name>
          <article-title>A random forest guided tour</article-title>
          .
          <source>TEST 25</source>
          ,
          <fpage>197</fpage>
          -
          <lpage>227</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Miller</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveraging</surname>
            <given-names>BERT</given-names>
          </string-name>
          for Extractive Text Summarization on Lectures. arXiv:
          <year>1906</year>
          .
          <article-title>04165 [cs</article-title>
          .CL]
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Basu</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghosh</surname>
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ghosh</surname>
            <given-names>K.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Overview of the FIRE 2018 track: Information Retrieval from Microblogs during Disasters (IRMiDis)</article-title>
          .
          <source>In Proceedings of the 10th annual meeting of the Forum for Information Retrieval Evaluation (FIRE'18)</source>
          .
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . DOI: https://doi.org/10.1145/3293339.3293340
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Dutt</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basu</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghosh</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghosh</surname>
            <given-names>S.</given-names>
          </string-name>
          <article-title>Utilizing microblogs for assisting post-disaster relief operations via matching resource needs and availabilities</article-title>
          ,
          <source>Information Processing and Management</source>
          , Volume
          <volume>56</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>5</given-names>
          </string-name>
          ,
          <year>2019</year>
          , Pages
          <fpage>1680</fpage>
          -
          <lpage>1697</lpage>
          , ISSN 0306-
          <fpage>4573</fpage>
          . DOI: https://doi.org/10.1016/j.ipm.
          <year>2019</year>
          .
          <volume>05</volume>
          .010.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Chatterjee</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deng</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shan</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiao</surname>
            <given-names>W.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Classifying facts and opinions in Twitter messages: a deep learning-based approach</article-title>
          .
          <source>Journal of Business Analytics</source>
          .
          <volume>1</volume>
          .
          <fpage>29</fpage>
          -
          <lpage>39</lpage>
          .
          <fpage>10</fpage>
          .1080/2573234X.
          <year>2018</year>
          .
          <volume>1506687</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Shekhar</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gangisetty</surname>
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <source>Disaster Analysis Through Tweets</source>
          .
          <volume>10</volume>
          .1109/ICACCI.
          <year>2015</year>
          .
          <volume>7275861</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Victor</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lysandre</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Julien</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            <given-names>W.</given-names>
          </string-name>
          <article-title>DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter</article-title>
          . arXiv:
          <year>1910</year>
          .
          <article-title>01108 [cs</article-title>
          .CL]
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Elaziz</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hosny</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salah</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darwish</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            <given-names>S.</given-names>
          </string-name>
          , et al. (
          <year>2020</year>
          )
          <article-title>New machine learning method for image-based diagnosis of COVID-19</article-title>
          . PLOS ONE
          <volume>15</volume>
          (
          <article-title>6): e0235187</article-title>
          . https://doi.org/10.1371/journal.pone.0235187
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>