<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Forum for Information Retrieval Evaluation, December</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>VaxiBERT: A BERT-Based Classifier for Vaccine Tweets with Multi-Label Annotations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shivangi Bithel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Samidha Verma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Prachi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rajat Singh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indian Institute of Technology</institution>
          ,
          <addr-line>Delhi</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>5</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>Vaccination has long been seen as an essential component of public health, providing a critical line of defense against infectious diseases. Our primary objective is to build a robust multi-label classification system capable of categorizing individual social media posts, specifically tweets, based on the numerous vaccine-related concerns stated by their writers. These reservations cover many issues, including misgivings about necessity, safety, and political intentions. We discuss our approach and evaluation results, shedding light on the intricate interplay between feeling, society, and science in the arena of the vaccine debate, using cutting-edge models such as Covid-Twitter-BERT and OpenLLaMA-7B. Our best-submitted run achieved a 0.67 macro-F1 Score and 0.70 Jaccard score. Github Code: https://github.com/shivangibithel/VaxiBERT_AISoMe2023</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Sentiment Analysis</kwd>
        <kwd>COVID-19 Vaccine Tweets</kwd>
        <kwd>COVID-Twitter-BERT</kwd>
        <kwd>Large Language Model</kwd>
        <kwd>LoRA PEFT</kwd>
        <kwd>Multi-label Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>A key component of public health for decades has been vaccination, a strong defense against
the spreading of infectious diseases. Its significance in preventing outbreaks and safeguarding
local populations cannot be emphasized. The crucial role that vaccination played in containing
the COVID-19 pandemic, a global emergency that has brought vaccinations into the public
eye like never before, has served as a reminder of the need for vaccination in the modern era.
Beyond the pandemic reaction, widespread vaccination acceptance, especially on a societal
level, continues to be essential in preventing disease resurgence, preventing childhood diseases,
and reducing the yearly assault of seasonal illnesses like influenza.</p>
      <p>However, the vaccine environment is defined by the complexity that reaches far beyond the
scientific arena. A distinct undercurrent of suspicion has emerged, spurred by various issues
ranging from politics to alleged side efects. This skepticism is a severe obstacle that must
be addressed as we strive for widespread protection through vaccination. Understanding the
complex issues surrounding vaccinations is critical, and in this age of digital connectedness,
social media platforms have emerged as a great source of information.</p>
      <p>In this context, our work aims to navigate the complex web of public opinion on vaccines,
mainly expressed in social media’s unfiltered and dynamic arena. Our primary goal is to create
a solid and dynamic multi-label classification system capable of categorizing individual social
media posts, specifically tweets, based on the various vaccine-related concerns raised by their
authors. It is critical to recognize that these worries are not uniform; a single tweet may include
numerous separate vaccine-related concerns.</p>
      <p>To simplify our classification task, we have a comprehensive set of concern labels that capture
the wide range of anxieties permeating the vaccine discourse. These terms cover a variety of
concerns, including skepticism about the necessity and safety of vaccinations, suspicions of
larger conspiracies, political motivations behind vaccination mandates, and uncertainty about
the efectiveness of vaccinations. Concerns also include vaccines’ origins, make-up, and alleged
negative efects, with personal religious beliefs influencing opinions.</p>
      <p>In a time when information travels through digital channels at unprecedented speeds, our
study aims to use the vast amounts of data generated on social media platforms to shed light
on the complex world of vaccine apprehension. We aim to provide insightful contributions
that can inform public health strategies, enhance vaccine communication, and foster a more
nuanced understanding of the complex interplay between science, society, and sentiment in the
ifeld of vaccination by analyzing the concerns raised by individuals in their tweets.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task</title>
      <p>
        The task " Building an efective multi-label classifier to label a social media post(particularly,
a tweet) according to the specific concern(s) towards vaccines as expressed by the author of
the post" organized as a part of AISoMe (Artificial Intelligence on Social Media) Track in the
FIRE (Forum for Information Retrieval Evaluation) 2023, we present an efective approach in
this paper [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. A tweet can have more than one label (concern); e.g., a tweet expressing
three diferent concerns towards vaccines will have three labels. The tweets are classified into
multiple classes described below with examples:
• Unnecessary - "The tweet indicates vaccines are unnecessary or that alternate cures are
better."
• Mandatory - "Against mandatory vaccination — The tweet suggests that vaccines should
not be made mandatory."
• Pharma - "Against Big Pharma — The tweet indicates that the Big Pharmaceutical
companies are just trying to earn money, or the tweet is against such companies in
general because of their history."
• Conspiracy - "Deeper Conspiracy — The tweet suggests some deeper conspiracy, and
not just that the Big Pharma wants to make money (e.g., vaccines are being used to track
people, COVID is a hoax)"
• Political - "Political side of vaccines — The tweet expresses concerns that the
governments/politicians are pushing their own agenda through the vaccines."
• Country - "Country of origin — The tweet is against some vaccine because of the country
where it was developed/manufactured"
• Rushed - "Untested / Rushed Process — The tweet expresses concerns that the vaccines
have not been tested properly or that the published data is inaccurate."
• Ingredients - "Vaccine Ingredients/technology — The tweet expresses concerns about
the ingredients present in the vaccines (e.g., fetal cells, chemicals) or the technology used
(e.g., mRNA vaccines can change your DNA)"
• Side-efects - "Side Efects / Deaths — The tweet expresses concerns about the side efects
of the vaccines, including deaths caused."
• Inefective - "Vaccine is Inefective — The tweet expresses concerns that the vaccines
are not efective enough and are useless."
• Religious - "Religious Reasons — The tweet is against vaccines because of religious
reasons"
• None - "No specific reason stated in the tweet, or some reason other than the given ones."
Given below are a few examples of tweets along with their labels:
• "FYI....there are plenty of people walking around without vaccines for all sorts of
contagious diseases/viruses. Why is Covid so diferent? We must ask why a mandatory vaccine
card is even a consideration if the ones who are vaccinated feel that it protects them." :
Mandatory Unnecessary
• "So there have been issues, but FDA are so desperate they deny it’s the vaccine FDA
reports facial paralysis in 4 volunteers for Pfizer’s Covid-19 vaccine, but FDA denies
vaccine is the cause - Business Line https://t.co/nD8gwuxbvu" : side-efects
• "If this is seen as the deadliest disease in our lifetimes, and consequently the vaccine
viewed as a miraculous panacea, why is Pfizer’s stock price virtually unchanged from the
beginning of the year?" : pharma
• "@MelanieMetz6 @XSOmegaMkII Inovio...look it up as well as Moderna. All 3 of these
delivery methods have nano technology that can deliver DNA/RNA gene coding &amp;
mutation. This isnt a joke or up for speculation....its way beyond that now. I have
leukaemia with 17q deletion So No." : side-efect ingredients conspiracy
• "Doctors Around the World Issue Dire WARNING: DO NOT GET THE COVID VACCINE!!
https://t.co/JD5mlPTbVt via @Prepare_Change" : none
• "This is the same CEO that sold 60+% of his stock in Pfizer on the day of the vaccine
announcement. Sell the news, don’t take the vaccine, he seems super bullish on the long
term successful prospects if this vaccine. https://t.co/m5dS8Y9Q9t" : inefective pharma
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Related Work</title>
      <p>
        Users express their opinions regarding healthcare, diseases, treatments, vaccines, and
immunization campaigns on microblogs like Twitter. In social computing, information extraction
from these text-based tweets is increasingly popular. Classical machine learning techniques
such as linear classifiers, Naive-Bayes classifiers, support vector machines, and deep neural
techniques such as Long Short Term Memory(LSTMs) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Bidirectional RNN [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], BERT(Bidirectional
Encoder Representations from Transformers) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and RoBERTa [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. For natural language
processing, more modern language models include large pre-trained models like T5 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], GPT3 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
LLaMA [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], PALM [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and many more.
3.1. BERT
BERT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is a highly efective transformer-based architecture that adapts well to numerous tasks
involving natural language processing. BERT allows for the pre-training of deep bidirectional
representations from unlabeled text, which preserves more of the context and logical flow of
the text. The model is pre-trained using next-sentence prediction (NSP) tasks and Masked
Language Modelling (MLM). By including an additional output layer and achieving cutting-edge
performance, the BERT model may be fine-tuned for a variety of jobs.
3.2. LLaMA
The Large Language Model Meta AI [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], abbreviated as LLaMA, represents a significant
advancement in the realm of natural language processing. This collection of state-of-the-art
foundation language models spans a spectrum of sizes, ranging from 7 billion to 65 billion
parameters. What sets LLaMA apart is its ability to deliver exceptional performance while
maintaining a comparatively smaller model size, thereby reducing the computational demands
typically associated with cutting-edge language models. LLaMA’s foundation models have
been meticulously trained on a diverse and extensive range of unlabeled datasets. This training
corpus includes data from sources such as CommonCrawl, C4, GitHub, Wikipedia, books, ArXiv,
StackExchange, and more. The amalgamation of these varied datasets has empowered LLaMA
to attain state-of-the-art performance, rivaling other top-performing models like Chinchilla-70B
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and PaLM-540B [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Dataset</title>
      <p>
        This work uses a training dataset created as part of the research project "CAVES: A dataset to
facilitate explainable classification and summarization of concerns towards COVID-19 vaccines."
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This carefully managed training dataset includes a sizable corpus of 9,921 tweets criticizing
the COVID-19 vaccination. These tweets were collected between 2020 and 2021 and have
undergone meticulous manual annotation by subject-matter specialists. The issue categories
in our research objectives have been carefully assigned to each tweet in this dataset. To
assess the generalizability and robustness of our classification system, the test set encompasses
approximately 500 tweets obtained from diverse sources. These tweets are not exclusively
centered on COVID-19 vaccines; they span a broader spectrum, incorporating discussions on
other vaccine types, such as the MMR and the flu.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Pre-processing</title>
      <p>
        In line with prior research [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ], we conducted extensive pre-processing of the tweet data to
enhance the quality of word embeddings. Tweets inherently feature unique lexicons, including
elements such as HASHTAGS, @USER mentions, HTTP-URLs, and EMOJIS. These elements
often introduce noise if left unattended and adversely afect model performance. Therefore,
we implemented a comprehensive data-cleaning pipeline as part of our tweet pre-processing
procedure, encompassing the following key steps:
• Stop Word Removal: To streamline the text and emphasize essential information, we
eliminated common stop words such as "the," "a," "an," and "in." These words typically do
not contribute significant meaning to the text.
• Lowercasing: Given the informal nature of tweets, we converted all words to lowercase.
      </p>
      <p>
        This practice standardizes the text and ensures that each word is represented consistently,
facilitating more efective text analysis.
• Emoticon Conversion: Emojis are frequently employed on Twitter to express emotions
and sentiments. Recognizing their importance, we refrained from outright removal and
instead converted emojis to their corresponding textual representations. This
transformation retained the sentiment and emotional context of the text. The ’emoji’ library
(https://pypi.org/project/emoji/) aided in this process.
• Contractions Expansion: We systematically expanded contractions to their original,
uncontracted forms to promote text standardization. For example, "don’t" was expanded
to "do not." We expanded this expansion by leveraging the ’contractions’ library( https:
//pypi.org/project/contractions/).
• Non-Alphanumeric Character Removal: Extraneous non-letter characters, including
brackets, colons, semi-colons, @ symbols, and the like, were removed from the text. This
step contributed to text cleanliness and coherence.
• URL Removal: URLs unrelated to sentiment analysis were purged from the text using
regular expressions. This exclusion aided in focusing the analysis on the textual content
pertinent to sentiment assessment.
6. Methodology
• Run1: COVID-Twitter-BERT (CT-BERT): We used a domain-specific
transformerbased model called CT-BERT[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. We chose this model specifically because BERT-Large
is trained on Wikipedia data, and using a pre-trained model in the same domain, in this
case, COVID-19-related tweets would give more significant results after fine-tuning with
the provided training data. We shufled the training data, then split it into training and
validation sets in the ratio 80:20 such that the percentage of instances of each class was
preserved in both sets. Both training and validation instances were pre-processed, as
explained in section 5. The resulting training data was used for fine-tuning CT-BERT[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
while validation data was used for evaluation. We trained the model for 15 epochs with a
learning rate of 2e-5. The test data was also pre-processed using the same steps as training
and validation data first to generate the embeddings for the tweet and then predict the
probability scores of each tweet against all the classes. We used the sigmoid function
over probability values with a threshold of 0.5 to predict the label. The final
prediction file containing the Tweet ID and the predicted class was submitted as run1 for the task.
• Run2: OpenLLaMA-7B: We use the OpenLLaMA-7B [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] model variant to finetune
for the task at hand. The training data instances were pre-processed, as explained in
section 5. Following the same methodology as given in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], we use the Prefix Tuning
technique, which falls into the larger category of PEFT (Parameter Eficient Fine Tuning)
approaches. In this, we learn a set of adaption prompt tokens, which is appended at the
beginning of some top-N transformer layers. While finetuning, only these prompt tokens
are finetuned for a specific downstream task, while the rest of the LLM parameters
remain frozen. Also, a zero-initialized, zero-gated attention mechanism is used to inject
the finetuned prompt token knowledge into the existing model so that the original
model parameters don’t deviate too much due to noise in the initial learning phase.
We use 10 extra learnable prompt tokens in our setting and append them to the top
30 transformer layers. This adds only an extra 1.2M parameters over the existing 7B
frozen parameters, requiring 10 minutes to train for 9,921 data points using batch
size 4, 512 as max sequence length for 5 epochs using a learning rate of 9e-3. We
generate the classification labels using the pre-processed test data as a text generation task.
• Run3: OpenLLaMA-7B: In run 3, we fine-tuned the OpenLLaMA-7B model variant for
multi-label classification without pre-processing the tweets. All other details of model
training are similar to the Run2. In this run, we generate the classification labels using
the raw test data as a text generation task.
      </p>
    </sec>
    <sec id="sec-6">
      <title>7. Evaluation</title>
      <p>AISoMe Track results are evaluated using the macro-F1 score and Jaccard Score. The result of
our three submitted runs for the task is shown in Table 1.</p>
      <p>Sr No. Team_ID
Run 1 DSIRC
Run 2 DSIRC
Run 3 DSIRC
macro-F1 score
0.67
0.57
0.55</p>
      <p>Jaccard Score
0.7
0.61
0.6</p>
      <p>Rank
4
13
16</p>
    </sec>
    <sec id="sec-7">
      <title>8. Conclusion and Future Work</title>
      <p>This study employs Covid-Twitter-BERT and Open-LLaMA-7B to categorize vaccination-related
tweets. The transformer-based model outperforms the fine-tuned OpenLLaMA-7B-based
classifiers because its word embeddings are more expressive and yield better results on test data.
Furthermore, because transformer-based models require many data, we recommend looking at
data augmentation solutions to improve the performance of our model. Another aspect would
be to train the model to become more robust against adversaries.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Samad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ganguly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>Caves: A dataset to facilitate explainable classification and summarization of concerns towards covid vaccines</article-title>
          ,
          <source>in: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>3154</fpage>
          -
          <lpage>3164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Basu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>Overview of the fire 2023 track:artificial intelligence on social media (aisome)</article-title>
          ,
          <source>in: Proceedings of the 15th Annual Meeting of the Forum for Information Retrieval Evaluation</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Gref</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Koutník</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Steunebrink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <article-title>Lstm: A search space odyssey</article-title>
          ,
          <source>IEEE Transactions on Neural Networks and Learning Systems</source>
          <volume>28</volume>
          (
          <year>2015</year>
          )
          <fpage>2222</fpage>
          -
          <lpage>2232</lpage>
          . URL: https://api.semanticscholar.org/CorpusID:3356463.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Rumelhart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <article-title>Learning internal representations by error propagation</article-title>
          ,
          <year>1986</year>
          . URL: https://api.semanticscholar.org/CorpusID:62245742.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <source>Association for Computational Linguistics</source>
          , Minneapolis, Minnesota,
          <year>2019</year>
          . URL: https://aclanthology.org/N19-1423. doi:
          <volume>10</volume>
          . 18653/v1/
          <fpage>N19</fpage>
          -1423.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>
          ,
          <source>J. Mach. Learn. Res</source>
          .
          <volume>21</volume>
          (
          <year>2019</year>
          )
          <volume>140</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>140</lpage>
          :
          <fpage>67</fpage>
          . URL: https://api.semanticscholar.org/CorpusID: 204838007.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ramesh</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hesse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , E. Sigler,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Litwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chess</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
          </string-name>
          ,
          <article-title>Language models are few-shot learners</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>2005</year>
          .14165.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Izacard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Martinet</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Rozière</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Hambro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Azhar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Rodriguez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Joulin</surname>
          </string-name>
          , E. Grave, G. Lample,
          <article-title>Llama: Open and eficient foundation language models</article-title>
          ,
          <source>ArXiv abs/2302</source>
          .13971 (
          <year>2023</year>
          ). URL: https: //api.semanticscholar.org/CorpusID:257219404.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chowdhery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bosma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Barham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. W.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gehrmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schuh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tsvyashchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Maynez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Barnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Prabhakaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Reif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Hutchinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pope</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bradbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Austin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Isard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gur-Ari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Duke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Levskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghemawat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Michalewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Misra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fedus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ippolito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Spiridonov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sepassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Omernick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Pillai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pellat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lewkowycz</surname>
          </string-name>
          , E. Moreira,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Polozov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Saeta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Firat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Catasta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. S.</given-names>
            <surname>Meier-Hellstern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Eck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petrov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Fiedel</surname>
          </string-name>
          , Palm:
          <article-title>Scaling language modeling with pathways</article-title>
          ,
          <source>ArXiv abs/2204</source>
          .02311 (
          <year>2022</year>
          ). URL: https://api.semanticscholar.org/CorpusID:247951931.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Borgeaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mensch</surname>
          </string-name>
          , E. Buchatskaya,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rutherford</surname>
          </string-name>
          , D. de Las Casas,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Hendricks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Welbl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hennigan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Noland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Millican</surname>
          </string-name>
          , G. van den Driessche,
          <string-name>
            <given-names>B.</given-names>
            <surname>Damoc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Osindero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Elsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Rae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          , L. Sifre,
          <article-title>Training compute-optimal large language models</article-title>
          ,
          <source>ArXiv abs/2203</source>
          .15556 (
          <year>2022</year>
          ). URL: https://api.semanticscholar.org/CorpusID:247778764.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Samad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ganguly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>Caves: A dataset to facilitate explainable classification and summarization of concerns towards covid vaccines</article-title>
          ,
          <source>in: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR '22,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2022</year>
          , p.
          <fpage>3154</fpage>
          -
          <lpage>3164</lpage>
          . URL: https://doi.org/10.1145/3477495.3531745. doi:
          <volume>10</volume>
          .1145/3477495.3531745.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bithel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Malagi</surname>
          </string-name>
          ,
          <article-title>Unsupervised identification of relevant prior cases</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2107</volume>
          .
          <fpage>08973</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bithel</surname>
          </string-name>
          , Ctc: Covid-
          <article-title>19 tweet classification using ct-bert (</article-title>
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Salathé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Kummervold</surname>
          </string-name>
          ,
          <article-title>Covid-twitter-bert: A natural language processing model to analyse COVID-19 content on twitter</article-title>
          , CoRR abs/
          <year>2005</year>
          .07503 (
          <year>2020</year>
          ). URL: https: //arxiv.org/abs/
          <year>2005</year>
          .07503. arXiv:
          <year>2005</year>
          .07503.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Han,
          <string-name>
            <surname>C</surname>
          </string-name>
          . Liu,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qiao</surname>
          </string-name>
          ,
          <article-title>Llama-adapter: Eficient fine-tuning of language models with zero-init attention</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>16199</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>