<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Biomedical Question Answering with Transformer Ensembles</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Raghav R</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jason Rauchwerk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Parth Rajwade</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tanay Gummadi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eric Nyberg</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Teruko Mitamura</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Language Technologies Institute, Carnegie Mellon University</institution>
          ,
          <addr-line>Pittsburgh, Pennsylvania</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recent advancements in natural language processing, specifically transformers, have shown great promise in improving the performance of question-answering systems. However, we observe that a single transformer model may not achieve suficient accuracy and reliability to meet the stringent requirements of biomedical question answering. Based on our participation in the BioASQ Challenge, we present a comprehensive approach for biomedical question answering using transformers, integrating an end-to-end data processing pipeline with the UMLS Metamap and diferent ensembling techniques. Our findings suggest that transformer ensembles achieve significant performance improvements when compared to individual models.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;biomedical question answering</kwd>
        <kwd>transformer models</kwd>
        <kwd>ensemble learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The rapid growth of on-line biomedical text has stimulated the research and development of
robust, specialized language models that provide reliable, high-accuracy responses to queries
posed against the medical literature. For example, the PubMed database contains more than 35
million citations and abstracts of biomedical articles1. To overcome the challenge of inadequate
contextual representation when matching queries in the biomedical domain, researchers have
turned to transformer-based language models such as BERT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which have demonstrated
remarkable eficacy in capturing contextual information from large corpora. Various adaptations
of BERT, namely Med-BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], SciBERT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and Clinical-BERT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] have been specifically
designed to address the need for context-aware biomedical language models.
      </p>
      <p>In this paper, we explore the hypothesis that an ensemble of transformer models can perform
better than a single transformer alone for specific bioinformatic tasks. We tested our hypothesis
by participating in the eleventh edition of the BioASQ Challenge [5], specifically focusing on
Phase B of Task 11b [6]. Our primary objective is to deliver "exact" answers for yes/no, factoid,
and list question types. The BioASQ Challenge consists of four rounds of test sets, providing
participants with the opportunity to submit up to five systems for each test set. The organizers
provide the dataset for Task 11b [7] in the form of a single training set and four test sets for
each evaluation round. We submitted a total of four systems across three of the test sets. This
allowed us to explore various approaches and methodologies, enhancing our understanding of
the problem space.</p>
      <p>After analyzing the performance of diferent systems in previous editions of the BioASQ
Challenge, we decided to ensemble BioBERT [8] and BioM-Electra [9] for factoid and list
questions. For yes/no questions, we employ BioM-Electra [10]. We use the Unified Medical
Language System (UMLS) MetaMap tool 2 for preprocessing data, and for synonym removal
during post-processing.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        There has been significant prior work done for question-answering in the biomedical domain.
Following the advent of BERT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Lee et al. [8] introduced BioBERT, a language modeling
approach that initializes a BERT model (pretrained on Wikipedia and BookCorpus) and continues
pretraining using masked-language modeling (MLM) and next sentence prediction (NSP) on
PubMed abstracts and PubMed Central (PMC) full-text articles. BlueBERT [11] follows a similar
approach but finds performance improvements, in the clinical domain, by pre-training on
PubMed abstracts and MIMIC-III clinical notes. However, Gu et al. [12] find the aforementioned
mixed-domain pretraining objective to be inferior to domain-specific pretraining from scratch
given the diference in vocabulary from the initial BERT model and the later biomedical context.
      </p>
      <p>PubMedBERT [12] is a new BERT model, trained from scratch using PubMed abstracts, that
outperforms BioBERT and BlueBERT; the authors attribute the performance improvements
to having an in-domain vocabulary which the architecture can model completely in order to
fully optimize for in-domain data. Jeong et al. [13] propose a sequential transfer learning
method for fine-tuning biomedical models on intermediate datasets, before fine-tuning on the
specific biomedical task; this helps to improve performance due to data scarcity for the final
task. Specifically, the authors show a significant F1 gain by training on MNLI [ 14] and SQuAD
[15] before BioASQ, and unifying context-length distributions between fine-tuning tasks. Ting
et al. [16] present a method using BioBERT to generate snippets for ideal answers (BioASQ
Task B), and then using these snippets to predict exact answers for factoid/list questions.</p>
      <p>BioM-ELECTRA and BioM-ALBERT [10] are variants of ELECTRA and ALBERT pretrained
on PubMed abstracts. They subsequently fine-tune on a combination of SQuAD and MNLI,
and finally the BioASQ dataset [ 17], achieving SOTA on BioASQ 10b for list questions [18].
BioLinkBERT [19] takes an alternative approach in adding an additional pretraining objective of
document relation prediction, in order to learn contextual linked concepts between documents
in the form of PubMed hyperlinks. Given limited resources, we selected BioBERT, the most
commonly cited baseline model for bioinformatics [8], and BioM-ELECTRA, the best-performing
model on the most recent BioASQ challenge [18] as the two baselines for our work.
2https://www.nlm.nih.gov/research/umls/implementation_resources/metamap.html</p>
    </sec>
    <sec id="sec-3">
      <title>3. Model Overview</title>
      <sec id="sec-3-1">
        <title>3.1. List and Factoid Questions</title>
        <p>Following the methodology of Alrowili and Vijay-Shanker [9], we merge factoid and list
questions to overcome the limited number of training examples. We split the lists into multiple
factoid questions and search the golden snippets for the spans that contain an exact string
match for the answer. Because there are multiple snippets for each question, this creates many
snippet-answer pairs for each original question. Each of these pairs are rewritten into the
SQuAD format to be fed into our models. We use BioBERT and BioM-ELECTRA for these
questions.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Yes/No Questions</title>
        <p>We treated yes/no questions as a binary classification problem. We concatenate all of the golden
snippets to create a paragraph and feed this context and the question to the model. We did not
attempt to answer yes/no questions in our first batch, but submitted a model for the second and
third batches to better compare our performance with submissions from previous years. We
started with DistilBERT [20], BioBERT, and BioM-ELECTRA for these questions. However, we
chose BioM-ELECTRA for our systems because of its superior performance.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <sec id="sec-4-1">
        <title>4.1. Dataset Preprocessing</title>
        <p>Snippets that come from diferent articles may use diferent names or acronyms to refer to the
same concept. For instance, the protein "transforming growth factor alpha" is variously referred
to as "transforming growth factor alpha", "transforming growth factor", "TGF ", and "TGF- ".
We use to the MetaMap tool to ensure that all answers in the snippets are properly identified.
MetaMap queries the UMLS Metathesaurus (curated by the National Library of Medicine3) to
determine the canonical form for each biomedical term. We run snippets through MetaMap to
expand all acronyms and abbreviations before finding the answer spans.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Synonym Postprocessing</title>
        <p>We also utilize UMLS MetaMap to remove synonyms from the model’s predicted candidate
answer list. In the Metathesaurus, each entity has a Concept Unique Identifier (ConceptUI),
which is shared among all names that can refer to the same entity. Our system sorts candidate
answers by their confidence scores and greedily constructs an answer set while making sure
that all final answers have unique ConceptUIs.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Ensembling</title>
        <p>We believe that ensembling models (which was also proposed by Alrowili and Vijay-Shanker [9]
as a future prospect), specifically BioBERT and BioM-ELECTRA, can combine their strengths
and form a better system. We performed a grid search to discover the weighting schemes that
maximize F1 score (for list questions) and MRR (for factoid questions). Our ensembling weights
were (0.004, 0.996) when maximizing F1 and (0.037, 0.963) when maximizing MRR for BioBERT
and BioM-ELECTRA, respectively. We use these weights to compute a linear combination of
weights and confidence scores for each predicted answer. We rerank and filter the candidate
answers to return our ensembled predictions.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. BioASQ Task 11b Systems</title>
      <p>We performed an 80-20 split on the training set for our validation purposes (internal testing).
The results are shown in Table 1.</p>
      <p>We participated in the BioASQ Challenge under the team name ‘AsqAway’, submitting 4
systems to the task. Our systems are described in Table 2. The BioASQ test performance of our
systems is described in Tables 3, 4, and 5. The evaluation datasets are small and there is high
variance in model performance across the batches, making it dificult to compare them directly.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>One of our initial observations was that the models returned a large number of probable answers,
despite having set probability thresholds. While this meant that all possible answers were being
covered in most of the cases, there were a large number of false positives which led to a drop in
precision. Upon further observations, we noticed that both acronyms and their expanded forms
were included in the training data. This formed the motivation behind the UMLS Preprocessing
step as described in Section 4.1.</p>
      <p>Another observation was that a lot of the answers returned by our models were synonyms
of each other. Since the challenge requires the systems to remove synonyms in the candidate
answers, we performed the Synonym Postprocessing step as described in Section 4.2. We
present a quantitative analysis of our results, based on UMLS Preprocessing and Synonym
Postprocessing in Tables 6 and 7 respectively. Figures 2 and 3 show a qualitative example of the
same. The number of answers returned is reduced significantly while maintaining accuracy.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Future Work</title>
      <p>We anticipate significant opportunities for improvement both on the data and modeling side.
As the model is currently only pre-trained on full-text articles and abstracts, there are instances
where our architecture chooses an adjacent but incorrect medical term (low precision) or returns
a nonsensical answer given the sparseness of the correct term in the full-text (low recall). The
former case is more prevalent among our experiments than the latter and we hypothesize adding
titles to the pre-training procedure can improve both precision and recall; precision is improved
as the title acts efectively as a distillation for the context and recall is improved in the form of
providing additional context for the model to train upon. Similarly, in the training procedure,
we can leverage a combination of the abstract from the provided documents rather than solely
relying on snippets as context.</p>
      <p>Moradi and Samwald [21] find vulnerabilities in BioBERT when exposed to word-level and
character-level noise; we corroborate this observation in instances where training data has key
medical terms misspelled or misused. Adversarial training ofers robustness to such errors: Jia
and Liang present the "AddSent" model-independent procedure from [22] and Du et al. [23]
ifnds performance gains in the context of BioASQ. However, the alternative "AddAny" procedure
by Jia and Liang [22] can also be implemented for more rigorous examples.</p>
      <p>Using larger models (e.g. Large, X-Large, XX-Large variants) for BioM-ELECTRA and
BioMALBERT empirically leads to incremental performance gains [10], however don’t address a
strata of error. Alternative models such as LinkBERT [19] show similar performance as the
Bio-M variants from preliminary experimentation. Changing the order of finetuning procedures
mentioned by Jeong et al. [13] in swapping the ordering of MNLI and SQuAD prior to tuning
for the BioASQ data has potential for marginal gains.</p>
      <p>We hypothesize the use of adversarial methods to have the most promise in delivering
performance improvements, followed by the use of alternative architectures.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusion</title>
      <p>In this paper, we address the challenge of generating accurate information retrieval systems for
biomedical information, specifically focusing on Phase B of Task 11b in the eleventh BioASQ
Challenge. Given the complex and sensitive nature of biomedical data, we adopt a novel approach
that involves ensembling state-of-the-art transformer models that have previously performed
well in BioASQ challenges, along with implementing data processing techniques based on UMLS
MetaMap. Our eforts aim to contribute towards the development of highly precise answers for
the list and factoid question types. Our approach yields promise for data-oriented techniques
towards improving performance on the task. Our code files are publicly available in a GitHub
repository4.
[5] A. Nentidis, G. Katsimpras, A. Krithara, S. Lima-López, E. Farré-Maduell, L. Gasco,
M. Krallinger, G. Paliouras, Overview of BioASQ 2023: The eleventh BioASQ challenge on
Large-Scale Biomedical Semantic Indexing and Question Answering, in: Experimental
IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fourteenth
International Conference of the CLEF Association (CLEF 2023), 2023.
[6] A. Nentidis, G. Katsimpras, A. Krithara, G. Paliouras, Overview of BioASQ Tasks 11b and
Synergy11 in CLEF2023, in: Working Notes of CLEF 2023 - Conference and Labs of the
Evaluation Forum, 2023.
[7] A. Krithara, A. Nentidis, K. Bougiatiotis, G. Paliouras, BioASQ-QA: A manually curated
corpus for Biomedical Question Answering, Scientific Data 10 (2023) 170.
[8] J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, J. Kang, Biobert: a pre-trained biomedical
language representation model for biomedical text mining, Bioinformatics 36 (2020)
1234–1240.
[9] S. Alrowili, K. Vijay-Shanker, Exploring biomedical question answering with
biomtransformers at bioasq10b challenge: Findings and techniques, in: Conference and Labs of
the Evaluation Forum, 2022.
[10] S. Alrowili, K. Vijay-Shanker, Biom-transformers: building large biomedical language
models with bert, albert and electra, in: Proceedings of the 20th Workshop on Biomedical
Language Processing, 2021, pp. 221–227.
[11] Y. Peng, S. Yan, Z. Lu, Transfer learning in biomedical natural language processing: an
evaluation of bert and elmo on ten benchmarking datasets, arXiv preprint arXiv:1906.05474
(2019).
[12] Y. Gu, R. Tinn, H. Cheng, M. Lucas, N. Usuyama, X. Liu, T. Naumann, J. Gao, H. Poon,
Domain-specific language model pretraining for biomedical natural language processing,
ACM Transactions on Computing for Healthcare (HEALTH) 3 (2021) 1–23.
[13] M. Jeong, M. Sung, G. Kim, D. Kim, W. Yoon, J. Yoo, J. Kang, Transferability of natural
language inference to biomedical question answering, arXiv preprint arXiv:2007.00217
(2020).
[14] A. Williams, N. Nangia, S. R. Bowman, A broad-coverage challenge corpus for sentence
understanding through inference, arXiv preprint arXiv:1704.05426 (2017).
[15] P. Rajpurkar, J. Zhang, K. Lopyrev, P. Liang, Squad: 100,000+ questions for machine
comprehension of text, arXiv preprint arXiv:1606.05250 (2016).
[16] H.-H. Ting, Y. Zhang, J.-C. Han, R. T.-H. Tsai, Ncu-iisr/as-gis: Using bertscore and snippet
score to improve the performance of pretrained language model in bioasq 10b phase b
(2022).
[17] S. Alrowili, V. Shanker, Large biomedical question answering models with albert and
electra., in: CLEF (Working Notes), 2021, pp. 213–220.
[18] S. Alrowili, K. Vijay-Shanker, Exploring biomedical question answering with
biomtransformers at bioasq10b challenge: Findings and techniques (2022).
[19] M. Yasunaga, J. Leskovec, P. Liang, Linkbert: Pretraining language models with document
links, arXiv preprint arXiv:2203.15827 (2022).
[20] V. Sanh, L. Debut, J. Chaumond, T. Wolf, Distilbert, a distilled version of bert: smaller,
faster, cheaper and lighter, 2020. arXiv:1910.01108.
[21] M. Moradi, M. Samwald, Improving the robustness and accuracy of biomedical language
models through adversarial training, Journal of Biomedical Informatics 132 (2022) 104114.
[22] R. Jia, P. Liang, Adversarial examples for evaluating reading comprehension systems,
arXiv preprint arXiv:1707.07328 (2017).
[23] Y. Du, J. Yan, Y. Lu, Y. Zhao, X. Jin, Improving biomedical question answering by data
augmentation and model weighting, IEEE/ACM Transactions on Computational Biology
and Bioinformatics (2022).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Rasmy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhi</surname>
          </string-name>
          ,
          <article-title>Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction</article-title>
          ,
          <source>NPJ digital medicine 4</source>
          (
          <year>2021</year>
          )
          <fpage>86</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohan</surname>
          </string-name>
          ,
          <article-title>Scibert: A pretrained language model for scientific text</article-title>
          , arXiv preprint arXiv:
          <year>1903</year>
          .
          <volume>10676</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pei</surname>
          </string-name>
          ,
          <article-title>Clinical-bert: Vision-language pre-training for radiograph diagnosis and reports generation</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>36</volume>
          ,
          <year>2022</year>
          , pp.
          <fpage>2982</fpage>
          -
          <lpage>2990</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>