<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Vaccine Vision: A deep learning approach towards identifying societal concerns regarding vaccines</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kaustav Das</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shruti Biswas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amity University</institution>
          ,
          <addr-line>Kolkata</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the wake of the COVID-19 pandemic, discerning public sentiments regarding vaccination stances - Pro-Vax, Anti-Vax, or Neutral - emerged as a pivotal undertaking. Leveraging machine learning on COVID-19 related tweets, this study focuses on categorizing individuals based on their vaccination perspectives. However, delving beyond the surface, the analysis uncovers a mosaic of concerns within Anti-Vax sentiments that surpass a mere dichotomy. These concerns encompass a diverse spectrum, spanning from conspiracies and political suspicions to multifaceted uncertainties. In response, this work employs a nuanced multi-label classification approach, aiming to thoroughly comprehend and classify the varied concerns explicitly articulated within Anti-Vax tweets. By scrutinizing a corpus of COVID-19 related tweets, this study endeavors to shed light on the intricate landscape of vaccine hesitancy, providing a comprehensive understanding of the multifaceted reasons underlying Anti-Vax sentiments.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Twitter</kwd>
        <kwd>microblogs</kwd>
        <kwd>COVID-19</kwd>
        <kwd>vaccine concerns</kwd>
        <kwd>tweet</kwd>
        <kwd>multi-label classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Vaccine hesitancy is defined as “delay in acceptance or refusal of vaccination despite the
availability of vaccination services”[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It is seen as the primary cause of decreasing vaccine
rates and the resurgence of vaccine-preventable illnesses in many countries.
      </p>
      <p>
        To combat the COVID-19 pandemic, researchers and pharmaceutical companies came up with
a number of vaccines using the S-protein of SARS-CoV-2. But regardless of the eforts made,
number of concerns rose, which led to a significant decrease in the vaccination drives. In the fight
against COVID-19, vaccine hesitancy has been identified as a significant hindrance.[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The
reason for most of this hesitancy can be traced back to social media platforms, while social media
platforms themselves didn’t directly influence the general public, a lot of misinformation by
unregulated users whose opinions when voiced through these platforms caused mass paranoia
and eventually a distrust towards the public health services. Some factors that contributed
to the hesitancy also include the fact that the vaccines being administered, reportedly came
with mild to quite severe side efects. Even before the emergence of SARS-CoV-2, WHO had
already highlighted vaccine hesitancy as one of the ten leading threats to global health.[3]
Thus, it is quite evident that vaccine hesitancy is a significant and complex phenomenon that
cannot be avoided. To aid such a situation, a thorough analysis of the general public concern
needs to be done. By virtue of social media, these concerns have been voiced by people in the
form of tweets and other social media posts. Therefore, in this paper with the help of deep
learning techniques and natural language processing, we have built an eficient multi-label
classifier to identify some of the concerns associated with these tweets. Most of the tweets are
primarily focused on vaccines associated with COVID-19 like Moderna, Pfizer, Astrazeneca
etc. and most of them also voice more than one concern(label) hence the need for a multi-label
classifier.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Dataset</title>
      <p>With our requirements and concerns(labels) defined let’s have a look at the dataset.The dataset
used for this work can be credited to "CAVES: A Dataset to facilitate Explainable
Classification and Summarization of Concerns towards COVID Vaccines"[4]. This dataset has
been provided as part of the FIRE conference [5]. Here are some examples of tweets per label:</p>
      <p>CAVES Dataset
Original Tweet
@SenSanders We have immune systems. Vaccines are filled with animal
and human DNA, disinfectants, heavy metals, chemicals. For over 30 years
the US has paid out over 4 billion In compensations for deaths and injuries
because the US gave Pharma liability freedom in 1986. #VaccinesKill
@barbieashdown @leslieh707 Oh my! One does not have to be an expert at conspiracy
reading body language to know he is covering something up. Depopulation
is his game. No CV19 vax for me. No Moderna vax for anyone.
@PetenyiSandor @katka_cseh Your choice, go and get a Chinese vaccine. I country
will wait for vaccines which are approved by the EU. We will have enough
by the summer. Let’s see how much the Chinese actually will deliver, how
many people willing to accept it and if Orban can organize mass vaccination.
I doubt it.</p>
      <p>Canadians who received Vaccine vaccine excluded from seeing Bruce Spring- side-efect,
inefecsteen, as Broadway opens up via people are against vaccine but Dr. Bonnie tive
Henry stands by all vaccines all are safe seems they have no idea if they are
or not folks
@TraceyLouise @slimschutte @AefreBetty @jon_severs @StevenCross81
Ah yes, that good old fashioned vaccine that was trialled over the correct
amount of time, and not forced upon every living being on the planet... Even
without this new thing, if you’re healthy, you’ll still be fine.</p>
      <p>Big Pharma spent close to 7 Billion euros on lobbying in France...which as pharma
we know is utterly corrupted in terms of Covid response: blocked
hydroxychloroquine - like all other Western countries - and embraced remdesivir
and vaccines.</p>
      <p>Labels
ingredients,
pharma, side-efect
unnecessary,
rushed, mandatory
@realDonaldTrump Guess who’s gonna get rich before he leaves the white
house of of 71 million people he knows will take a shot in the arm if he says
so? That’s right. Donald Trump now wants science on his side. How much
stock in Pfizer did you buy huh?
AUTHORITARIAN NAZI-wannabes of Democrat party are COMPLETELY
ANTI VOTER ID to PREVENT FRAUD. Those very same LIBERAL TYRANTS
want YOU to be DIGITALLY marked under the pretense of getting a "vaccine."
NO THANKS! VOTER ID or NO VOTE! Keep both your vaccine and SATAN’S
MARK!
Now about this Pfizer "vaccine". Who is the government tryna give this to
again. By my calculations, the people who would want it are the folks who
never got Covid19. Be careful of folks giving you a vaccine that will "protect"
you from a virus for which you already had it. https://t.co/zdJqKENupA.
@NeilClark66 Mad fascist dictator Johnson now wants to force you to take
an experimental vaccine with no long-term safety profile. All this for a virus
with 0.3% IFR. Still think this is about a virus? 1922 committee must step in
NOW and REMOVE this communist lunatic!</p>
      <sec id="sec-2-1">
        <title>2.1. Data Pre-processing</title>
        <p>For the pre-processing step, standard text-cleaning steps were followed where lexicons like
URLS, usernames, and emojis were removed. Here are the steps:
• Tweet id columns were dropped.
• The label columns were binarized/ one hot encoded using the sklearn multi-label binarizer
so that they can be used readily for the future. machine learning steps. (https://scikit-learn.
org/stable/modules/generated/sklearn.preprocessing.MultiLabelBinarizer.html)
• Text standardization:
– Removing HTML elements
– Replacing non-standard punctuation with standard version.
– Replacing '\r', '\n' and '\t' with white spaces
– Removing all control characters
– Removing duplicate white spaces
• Removing contractions from text using contractions library (https://pypi.org/project/
contractions/)
• Replacing usernames and URLs with fillers ’username’ and ’url’
• Removing unicode and accented characters from text
• Emojis were replaced with their meaning using the emoji.demojize() method of emoji
library. (https://pypi.org/project/emoji/)
These steps were also taken in compliance with the standard cleaning steps taken for
CTBERT[6], for getting the best quality bert-embeddings.1</p>
        <p>Cleaned Data
Original Tweet Cleaned Tweet
@PaolaQP1231 Well, I mean congratu- username well, i mean congratulations
lations Covid19 for being the first ever covid19 for being the first ever "thing" to
âoethingâ to eradicate influenza. In other eradicate influenza. in other news, covid
news, Covid vaccines will spur the rise of vaccines will spur the rise of influenza in
influenza In 2021-2022 season. Influenza 2021-2022 season. influenza will be
returnwill be returning for a shot at the title belt. ing for a shot at the title belt. nov 2, 2021
Nov 2, 2021 on pay per view. Order today. on pay per view. order today.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>Before starting with the model training step, we must handle the data imbalance present in
our dataset. Thus, to handle the imbalance a custom loss function called Distribution Balanced
Loss [7] was used. COVID Tweeter-BERT (CT-BERT)[6]: CT-BERT is a domain-specific
transformer-based model, pre-trained using a sizable corpus of tweets about COVID-19 that
were posted between January 12 and April 16, 2020. It is initialized with BERT-Large[8] weights
and then pre-trained/fine-tuned using 160 million tweets regarding the coronavirus. The major
advantage of using CT-BERT is that all of our vaccination tweets in the dataset are concerned
with COVID-19, for which CT-BERT produces domain-specific, quality embeddings. CT-BERT
was our primary model and performed best for the classification task. Other models like
Bi-LSTM[9], and BERT variants like ROBERTA were considered which did not perform as well.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <p>The primary evaluation metric for our model performance was macro-f1 The model was
evaluated on the held-out validation set compromising of 1984 tweets containing tweet instances
from each label class.</p>
      <p>Bi-LSTM(Multi-smote + word2vec)
Baseline CT-BERT(without DB-LOSS)
CT-BERT(with DB-Loss + without cleaning)</p>
      <p>CT-BERT(with DB-LOSS)
1CT-BERT is our final model used for classification, it will be explained later on in the methodology section
From the results seen, it can be concluded that CT-BERT with DB-LOSS and proper cleaning
produces the best result achieving the highest macro-F1 score.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>To finally summarize and prove the efectiveness of our model we will test our model over
a completely unknown distribution of vaccine-related tweets. This dataset 2 contains tweets
associated with COVID-19 vaccines as well as some other non-COVID vaccines.</p>
      <p>Here is the model performance of CT-BERT+DB-LOSS on the same:</p>
      <p>Run File
run_submission.csv</p>
      <p>Methodology
CT-BERT with DBLOSS</p>
      <p>Macro-F1 score Jaccard score
0.71 0.70</p>
      <p>This shows a fairly respectable performance, considering the dataset consisted of many
instances previously unseen by our model. The dataset and code for model training and setup
of our experiment can be found in this repository.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This work further solidifies the efectiveness of pre-trained models like CT-BERT when combined
with other techniques like Distribution Balanced Loss. Also, from a societal angle, the work
further proves the extent to which such a seemingly exhaustive task of manually analyzing
tweets can be automated with the amalgamation of cutting-edge A.I. systems, to further help
the concerned authorities act and mitigate the concerns of the general public. Thus, yet again a
deep understanding of society on a global scale can be done with the power of deep learning
and NLP, proving that AI can be a fast and impactful solution to many of today’s problems.</p>
      <sec id="sec-6-1">
        <title>6.1. Future Work</title>
        <p>The next step would be to make an AI system with the model built here running its background.
Also, the model’s performance can be further improved with the classification being done in
two levels that is by first segregating the ’none’ labeled tweets by classification on the first level
and then classifying the rest of the tweets on the next level.</p>
        <p>2This dataset has been provided by FIRE [5] as part of model evaluation for each team.
[3] E. Robertson, K. S. Reeve, C. L. Niedzwiedz, J. Moore, M. Blake, M. Green, S. V. Katikireddi,
M. J. Benzeval, Predictors of covid-19 vaccine hesitancy in the uk household longitudinal
study, Brain, Behavior, and Immunity 94 (2021) 41–50. URL: https://www.sciencedirect.com/
science/article/pii/S0889159121001100. doi:https://doi.org/10.1016/j.bbi.2021.
03.008.
[4] S. Poddar, A. M. Samad, R. Mukherjee, N. Ganguly, S. Ghosh, Caves: A dataset to facilitate
explainable classification and summarization of concerns towards covid vaccines, in:
Proceedings of the 45th International ACM SIGIR Conference on Research and Development
in Information Retrieval, 2022, pp. 3154–3164.
[5] S. Poddar, M. Basu, K. Ghosh, S. Ghosh, Overview of the fire 2023 track:artificial intelligence
on social media (aisome), in: Proceedings of the 15th Annual Meeting of the Forum for
Information Retrieval Evaluation, 2023.
[6] M. Müller, M. Salathé, P. E. Kummervold, Covid-twitter-bert: A natural language processing
model to analyse covid-19 content on twitter, Frontiers in Artificial Intelligence 6 (2023).
URL: https://www.frontiersin.org/articles/10.3389/frai.2023.1023281. doi:10.3389/frai.
2023.1023281.
[7] T. Wu, Q. Huang, Z. Liu, Y. Wang, D. Lin, Distribution-balanced loss for multi-label
classification in long-tailed datasets, 2021. arXiv:2007.09654.
[8] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional
transformers for language understanding, 2019. arXiv:1810.04805.
[9] R. C. Staudemeyer, E. R. Morris, Understanding lstm – a tutorial into long short-term
memory recurrent neural networks, 2019. arXiv:1909.09586.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N. E.</given-names>
            <surname>MacDonald</surname>
          </string-name>
          , et al.,
          <article-title>Vaccine hesitancy: Definition, scope and determinants</article-title>
          ,
          <source>Vaccine</source>
          <volume>33</volume>
          (
          <year>2015</year>
          )
          <fpage>4161</fpage>
          -
          <lpage>4164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Courtney</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.-M. Bliuc</surname>
          </string-name>
          ,
          <article-title>Antecedents of vaccine hesitancy in weird and east asian contexts</article-title>
          ,
          <source>Frontiers in psychology 12</source>
          (
          <year>2021</year>
          )
          <fpage>747721</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>