<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing Biomedical Question Answering with Parameter-Eficient Fine-Tuning and Hierarchical Retrieval Augmented Generation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yichen Gao</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Licheng Zong</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yu Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The CUHK Shenzhen Research Institute</institution>
          ,
          <addr-line>Hi-Tech Park, Shenzhen, 518057</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The Chinese University of Hong Kong</institution>
          ,
          <addr-line>Hong Kong SAR, 999077</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper reports the work done by the CUHK-AIH team in the 12th BioASQ Challenge task 12b, which involves Phases A, A+, and B. In Phase A, we build BM25 indexes for all documents from PubMed Central (PMC). When an input question is received, the system uses the question as a keyword to retrieve relevant documents from PMC via the BM25 retriever, obtaining a list of targeted documents. For Phase A+, we construct a hierarchical Retrieval-Augmented Generation (RAG) pipeline based on the Llama2-chat-7B model. The model is fine-tuned on the BioASQ training set using a Parameter-eficient Fine-Tuning (PEFT) method called Low-Rank Adaptation (LoRA). The system further refines the search results from Phase A by employing an ensemble retriever that combines sparse and dense retrievers to identify the most relevant chunks. Finally, the system feeds the question and the most relevant chunks into the base model to generate the answer using appropriate prompts. In Phase B, the answer generation pipeline is similar to Phase A+, with the main diference being that we directly build indexes for the questions and their relevant snippets, treating snippets as the basic retrieval unit. We conducted detailed ablation studies and analyses on the model types and retrieval techniques, which indicate that PEFT and RAG can significantly improve the performance in biomedical Question Answering (QA) tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Biomedical Question-Answering</kwd>
        <kwd>Retrieval-Augmented Generation</kwd>
        <kwd>Large Language Model</kwd>
        <kwd>Parameter-Eficient Fine-Tuning</kwd>
        <kwd>BioASQ</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, Large Language Models (LLM) have undergone significant development and have gained
widespread adoption [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5">1, 2, 3, 4, 5</xref>
        ], particularly in the biomedical domain including medical informatics
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], medical imaging [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and bioinformatics [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8, 9, 10</xref>
        ]. To address the requirements of healthcare
professionals and medical education, researchers have embarked on investigating the utilization of
Large Language Models (LLM) for medical knowledge queries. Their objective is to integrate intricate
medical knowledge into LLMs, thereby providing a knowledge query system with more accurate and
comprehensive medical information. Techniques like In-context Learning [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and Retrieval Augmented
Generation (RAG) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] have been proposed, which introduces a novel paradigm for tackling knowledge
queries and responses within specialized domains.
      </p>
      <p>
        In the field of medical information retrieval, a critical issue lies in constructing a robust information
retrieval system to efectively manage a massive corpus of medical literature [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. BioASQ [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] presents
a platform to address this challenge, encouraging researchers to develop better intelligent retrieval
systems. This paper focuses on task 12b [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], particularly the information retrieval, and the
questionanswering tasks, and tackles Phase A, Phase A+, and Phase B. Phase A involves finding relevant
documents or snippets to a biomedical question. Phase B provides biomedical inquiries alongside
relevant snippets (one or several sentences) and tasks participants with generating either the exact or
ideal answers based on these snippets. Phase A+, which is a new phase this year, will directly measure
the performance of the system while providing only the biomedical question. It’s worth noting that
in this work, we only focus on the ideal answer part and will not be involved in the evaluation of the
exact answer.
      </p>
      <p>
        Here we utilize the BM25 retriever to search documents in Phase A, and the Parameter-Eficient
Fine-Tuning on the BioASQ corpus [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] to generate high-quality answers in Phase B. In Phase A+, we
propose a hierarchical Retrieval-Augmented Generation pipeline combined with the mentioned BM25
retriever and PEFT for searching and generation. So, our system is called Corpus PEFT Searching (CPS).
For convenience, we defined the systems we used in Phase A, B, and A+ as CPS-A, CPS-B, and CPS-A+
respectively. By illustrating the results of our system and conducting detailed ablation studies, we
conclude that the PEFT and hierarchical RAG bring significant improvement on the model performance.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Since this paper can be divided into pipeline-related research and retrieval-related research, this section
will outline two parts of related work in Section 2.1 and Section 2.2. Specifically, we will begin by sorting
out the related work of the main pipeline, which consists of LLMs in medical informatics,
parametereficient model fine-tuning, and retrieval augmented generation. After that, we will introduce the
structure of the retrieved text in RAG, followed by a work that converts the retrieved text composition.</p>
      <sec id="sec-2-1">
        <title>2.1. Construction of Pipeline: LLM, PEFT, and RAG</title>
        <sec id="sec-2-1-1">
          <title>2.1.1. Large Language Models (LLMs) in Medical Informatics</title>
          <p>
            In recent years, the advancement of Large Language Models (LLMs), exemplified by GPT[
            <xref ref-type="bibr" rid="ref1">1</xref>
            ], BERT[
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]
and Llama[
            <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
            ], alongside the abundant public medical text data from platforms like PubMed1, has
encouraged researchers to explore diverse applications of LLMs in medical informatics. Agrawal et al.
[
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] reveals that sizable language models, even without explicitly training for medical fields, exhibit
the capacity to extract clinical information in few-shot settings. The latest iteration, Med-PaLM2 [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ],
ifne-tuned from Flan-PaLM [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ], attained an accuracy of 0.865 on the MedQA [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] benchmark designed
for generating biomedical LLMs. In addition, Venigalla et al. [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ] further conducted full-parameter
training of the model using data from PubMed, enhancing the model’s capabilities in medical
questionanswering. Despite the apparent complexity of training models with a substantial number of parameters,
the application of Parameter-Eficient Fine-Tuning (PEFT) techniques such as Low-Rank Adaptation
(LoRA) [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] has empowered researchers to eficiently fine-tune various models both in English and
Chinese medical context [
            <xref ref-type="bibr" rid="ref20 ref21 ref22 ref23">20, 21, 22, 23</xref>
            ], including our work.
          </p>
        </sec>
        <sec id="sec-2-1-2">
          <title>2.1.2. Parameter-Eficient Fine-Tuning (PEFT)</title>
          <p>
            Parameter-Eficient Fine-Tuning (PEFT) techniques encompass a series of fast, efective, and low GPU
computing and memory demand model fine-tuning methods. They are widely applied in the research of
Large Language Models, especially in specific domain applications [
            <xref ref-type="bibr" rid="ref22 ref23">22, 23</xref>
            ]. Adapter tuning [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ] involves
introducing small neural network modules (adapters) into the Transformer model and using a bottleneck
architecture to compress and restore the original feature vectors, adapting the language model to new
tasks. Prefix Tuning [
            <xref ref-type="bibr" rid="ref25">25</xref>
            ], proposed by X. L. Li and P. Liang, arranges a series of trainable prefixes
for each Transformer layer to improve performance in specific domains. Meanwhile, by freezing the
original matrix and approximating the parameter updates through low-rank decomposition, Low-Rank
Adaptation (LoRA) [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] fine-tunes models with eficient memory and storage usage.
          </p>
        </sec>
        <sec id="sec-2-1-3">
          <title>2.1.3. Retrieval Augmented Generation (RAG)</title>
          <p>
            To supplement the knowledge required by large language models when handling knowledge-intensive
tasks, the system can retrieve information from large corpora like Wikipedia to generate prompt
information [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. The retrieval process can be performed using sparse retrieval, utilizing methods
such as BM25 [
            <xref ref-type="bibr" rid="ref26">26</xref>
            ] or through dense retrieval [
            <xref ref-type="bibr" rid="ref27">27</xref>
            ] using vector similarity based on embedding models
like Bge [
            <xref ref-type="bibr" rid="ref28">28</xref>
            ]. To capture text relationships and knowledge across multiple documents, recent work
[
            <xref ref-type="bibr" rid="ref29 ref30 ref31">29, 30, 31</xref>
            ] treats multiple documents as a graph and conducts subgraph retrieval within the graph.
There are also applications and frameworks, such as ChatPDF 2 and LangChain 3, parsing large segments
of prompt information for question-answering systems by inputting and summarizing entire documents
in segmented fashion, thereby augmenting knowledge.
          </p>
          <p>
            In recent years, RAG also gains much attention in the biomedical field. MedCPT [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ] is proposed to
generate better biomedical sentence representations for zero-shot semantic information retrieval. It
consists of a pair of Transformer-based retriever and re-ranker pre-trained on PubMed user click logs
by contrastive learning. Xiong et al. [
            <xref ref-type="bibr" rid="ref33">33</xref>
            ] proposed MIRAGE and MEDRAG to conduct systematic and
large-scale experiments in the field of medical RAG, which shows that RAG based on biomedical corpus
can improve the performance of GPT-3.5 and Mixtral to the level of GPT-4.
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Retrieval Units: Chunk and Snippet</title>
        <sec id="sec-2-2-1">
          <title>2.2.1. Chunk: Generally Used Basic Unit of Retrieval</title>
          <p>
            Recently, numerous studies have focused on tasks within the medical domain, enhancing the capacity of
Large Language Models (LLMs) to tackle Question Answering (QA) tasks related to medical knowledge
through the application of Retrieval Augmented Generation (RAG) on biomedical databases. In these
studies, due to the substantial size of their databases and the absence of structured organization
or sentence extraction methods for Question Answering inquiries, the entirety of text or chunks is
employed as the basic unit of retrieval. GeneGPT[
            <xref ref-type="bibr" rid="ref34">34</xref>
            ] obtains knowledge for LLM in the genetic field
by calling the web API provided by the National Center for Biotechnology Information (NCBI). When
GeneGPT organizes this knowledge, it uses the Codex [
            <xref ref-type="bibr" rid="ref35">35</xref>
            ] model with long context length (8k tokens)
to handle structured API return results. Based on the Chain of Thought [36], the problem that requires
multiple returns from the API is decomposed into multiple sub-problems. The RAG pipeline that uses a
vectorized database, such as Chat-Orthopedist [37], splits the complete document in the database into
chunks with the same length, which has a suitable length for retrieval and input to LLM. In addition,
for RAG tasks in non-medical fields, approaches dividing documents into chunks are also very typical.
Usually, vector databases such as chromaDB 4 and Pinecone 5 are employed as the basis for construction.
          </p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2.2.2. Snippet: Basic Unit of Retrieval in BioASQ</title>
          <p>Besides the exemplary medical knowledge Question Answering Large Language Model (LLM) pipelines,
which stand out in the research domain, we also examined other systems featured in the BioASQ task
11b’s publication (CLEF-WN 2023 [38]). These systems primarily utilize the provided snippets directly,
but there are some slight variations in specific applications. UR-gpt [ 39], which performs very well
in the rankings, inserts snippets directly into prompts for in-context learning. The Sentence-based
Ranking [40] system’s method is close to the previous system, but they used the top 5 most relevant
sentences due to the input length limit of its model. The IISR [41] firstly selects the top 5 most relevant
snippets and uses ChatGPT to truncate or summarize them according to a fixed length. The context is
input as the ChatGPT conversation history instead of fusing it together with the question in IISR.
2https://www.chatpdf.com
3https://www.langchain.com/
4https://www.trychroma.com/
5https://www.pinecone.io/</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>This section details the pipelines we employed in each phase. Our system is called Corpus PEFT
Searching (CPS), for our main contribution is to integrate Parameter-Eficient Fine-Tuning (PEFT) on
the PMC corpus with a hierarchical retrieval-based searching method in the BioASQ 12b task. Since the
inputs and targeted outputs for submission are diferent in the three phases, we built three pipelines
named CPS-A, CPS-B, and CPS-A+ for Phase A, B, and A+ submissions respectively. Table 1 shows the
systems used in diferent phases and the sections and figures that describe them.</p>
      <p>In Phase A (Section 3.1) we leveraged the BM25 retriever to search relevant documents and return
the document list. Then in Phase B (Section 3.2), we used the BioASQ training set as the corpus to
ifne-tune a Large Language Model (LLM) based on a Parameter-Eficient Fine-Tuning (PEFT) method
to generate more reasonable answers. Finally, in Phase A+ (Section 3.3), we proposed a hierarchical
Retrieval-Augmented Generation (RAG) method combined with the methods mentioned above to
construct a general biomedical Q&amp;A system, which could serve as a chatbot for any other biomedical
questions.
3.1. Phase A</p>
      <p>
        As illustrated in Figure 1, in Phase A, we first downloaded all the abstracts available from PubMed
Central (PMC)6 and employed the package Pyserini[42] to build BM25 indexes of these abstracts, treating
each abstract as an individual document. Subsequently, we applied the BM25 retriever to search related
documents using the query questions as the keywords. After being filtered by an appropriate similarity
score threshold, the final list of relevant articles is obtained.
3.2. Phase B
Our work in Phase B can be divided into two parts: Pre-processing and Retrieval &amp; Generation, as
shown in Figure 2. In the Pre-processing stage, we built BM25 indexes of the golden enriched relevant
snippets using Pyserini[42] for later further retrieval. To help generate more accurate answers, we
used a Parameter-Eficient Fine-Tuning (PEFT) method, LoRA[ 43] to fine-tune the Llama2-chat-7B
model[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] on the BioASQ corpus, which is the Question-Answering pairs from BioASQ task 12b training
Golden Enriched
      </p>
      <p>Relevant Snippets
(Provided in Phase B)
set (consisting of 5049 Q&amp;A pairs). Due to the computing resource limit, we were unable to directly
ifne-tune the whole Large Language Model on the corpus. The Low-Rank Adaptation (LoRA)[ 43] helps
reduce the fine-tuning time and resources significantly while ensuring the model keeps a satisfying
performance. The work in this stage was done before the test sets were released and wouldn’t be
repeated later.</p>
      <p>The Retrieval &amp; Generation stage is the process to generate answers for the query questions in Phase
B. Our system leverages an ensemble retriever combining sparse (BM25) and dense (vector similarity)
retrievers for searching the most relevant snippets and provides them as references to the LLM. Bge7
was used to embed the chunks into vectors for similarity searching. The reason we constructed the
retriever is that we found the performance of directly inputting all snippets into the model are not good
enough, which we will discuss in Section 4.3.3.</p>
      <p>Given appropriate prompts with the retrieved snippets and the query question, the fine-tuned model
can generate a response for the query question to be submitted in Phase B. The prompt template is
shown below, which we modified from Alpaca 8 by adding a role definition at the beginning.</p>
      <p>You are an expert in the field of biomedical science.</p>
      <p>Below is an instruction that describes a task, paired with an input that provides further
context. Write a response that appropriately completes the request.
### Instruction: {question}
### Input: {context}
### Response:
In this template, the position of ‘context’ will be replaced by the relevant snippets found in the retrieval
procedure, while ‘question’ will be replaced by the query question body.
3.3. Phase A+
Phase A+ is more challenging than Phase B since there are no golden-enriched snippets. Therefore,
our system needs to first search the relevant documents and snippets from the whole PMC database,
and then generate answers. Here, we implement a hierarchical Retrieval Augmented Generation (RAG)
pipeline combining the insights from Phase A and B (shown in Figure 3).
7https://huggingface.co/BAAI/bge-large-en
8https://github.com/tatsu-lab/stanford_alpaca</p>
      <p>Pre-processing</p>
      <p>Similarly to Phase B, the framework consists of two stages: Pre-processing and Retrieval &amp;
Generation. The fine-tuned model in the Pre-processing stage is identical to that in Phase B, while the
indexes are built from the whole PMC abstracts rather than golden-enriched snippets.</p>
      <p>In the Retrieval &amp; Generation stage, we propose a hierarchical strategy to search question-relevant
texts. Specifically, considering the quicker nature of sparse retrieval compared to dense retrieval, we
still employed the BM25 sparse retrieval indexed on PMC abstracts as the first-level retrieval. This
process selects the most relevant abstracts as the preliminary information screening. These documents
are then segmented into multiple chunks by the TextSplitter from LangChain Package9 for the next
search. Subsequently, an ensemble retriever which is the same as the retriever in phase B is employed
as a more precise second-level retrieval. This ensemble retriever will identify multiple chunks most
relevant to the question as the second level of retrieval. Finally, these most relevant chunks and the
questions are combined by the prompt template mentioned in Section 3.2 to form the model input.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Analyses</title>
      <sec id="sec-4-1">
        <title>4.1. Experiment Settings</title>
        <sec id="sec-4-1-1">
          <title>4.1.1. Datasets</title>
          <p>In the oficial submissions this year, we utilize the training set of BioASQ Task 12b (consisting of 5049
question-answer pairs) as the LoRA fine-tuning (mentioned in Section 3.2 and Section 3.3) dataset
for CPS-B and CPS-A+ submissions. In the ablation study experiments in Section 4.3, we employ the
training set of BioASQ Task 11b (consisting of 4719 question-answer pairs) as the LoRA fine-tuning
dataset and the test set of BioASQ Task 11b batch 1 (consisting of 90 question-answer pairs) as the test
set to do the evaluations.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Model Deployment</title>
          <p>Considering factors such as deployment complexity, open-source availability, and parameter scale,
we chose the Llama2-7b-chat mode as the basic model in this work. For not focusing on the model
9LangChain:langchain_text_splitters.RecursiveCharacterTextSplitter</p>
          <p>Batch 1
Batch 2
Batch 3
Batch 4</p>
          <p>System</p>
          <p>CPS-A
Top Competitor</p>
          <p>CPS-A
Top Competitor</p>
          <p>CPS-A
Top Competitor</p>
          <p>CPS-A
Top Competitor
itself, this study adheres to default settings for model parameters and fine-tuning parameters. In all
our experiments and submissions, we use fp16 precision for both model fine-tuning and inference. We
performed our experiments on a Linux server with two NVIDIA GeForce RTX 3090 GPUs.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. BioASQ Task 12b Oficial Evaluations</title>
        <p>In this section, we present the oficial evaluation results of our system in the four batches of BioASQ
Task 12b. For comparison, we have included the best-performing systems in each batch. As there are
multiple metrics, we selected the system that ranked in the top 5 among all metrics in each batch. If
no system met this criterion, we expanded the selection to include systems ranked in the top 10. The
selected system is referred to as the "Top Competitor" in the tables in Section 4.2.
4.2.1. Phase A
We utilized the BM25 retrieval-based CPS-A system to retrieve relevant documents in phase A
(introduced in Section 3.1). Here we show the submission results of the retrieved document in Table 2. As for
selecting the relevant document list, we filter the retrieval score by an appropriate threshold optimized
on the training set and choose the top 10 documents if the filtered documents are more than 10.
In this section, we demonstrate the ‘ideal answer’ submission results in BioASQ task 12b phase B using
CPS-B in Table 3. Since there is randomness in answer generation, we submitted two or three results in
the challenge and we recorded the best one in the table.
We used CPS-A+ to complete the submission of Phase A+, and the ‘ideal answer’ results are illustrated
in Table 4. We did not submit the result of batch 1 due to the time limit.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Ablation Study</title>
        <sec id="sec-4-3-1">
          <title>4.3.1. On Generation Models</title>
          <p>The capability of the answer generation model has a significant influence on the performance in Phase
B. Here we compare the performance of diferent generation models when taking questions and relevant
contexts as input (shown in Table 5).</p>
          <p>The experiments in the first three rows of Table 5 were carried out by ourselves. In particular, the
Llama2-7B-chat + LoRA system is identical to the mentioned system CPS-B, but its dataset has been
changed as 4.1.1 stated. The diference between Llama2-7B-chat and CPS-B is that the fine-tuned
model is replaced by the original Llama2-7B-chat model. The GPT-3.5-turbo system is based on the API
of GPT-3.5-turbo, where the input consists of the query question and the context obtained from the
retrieval process of the CPS-B system. In this way, the input in the case of GPT-3.5-turbo is consistent
with that of CPS-B.</p>
          <p>Apart from these systems, we also select baselines and outstanding results in BioASQ task 11b phase
B10. BioASQ Baseline FS is an oficial baseline obtained from BioGPT [ 44], and is detailed described in
[38]. This oficial baseline system uses the concatenation of the question body and the relevant snippets
until the input length is exceeded. UR-gpt4 [39] and UR-gpt3.5 [40] are systems using ChatGPT4 and
ChatGPT3.5 with corresponding snippets. IISR [41] also used ChatGPT as its main model, but used a
diferent snippet input method as desctibed in 2.2.2.</p>
          <p>We can observe that the GPT-4 model is quite powerful and can obtain impressive results without
ifne-tuning (the highest R-2 Recall and the second highest R-SU4 Recall). Besides, our system won
the highest place among the three metrics including R-2 F1-score, R-SU4 Recall, and R-SU4 F1-score.
Therefore, fine-tuning a relatively small LLM can achieve amazing results and beat a very powerful
LLM in specific domains.</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>4.3.2. On Retrieved Context</title>
          <p>Beyond the generation model itself, the retrieved context can improve the accuracy of the generated
answers. Here we evaluated the performance while there is or isn’t context given to the models, which
is shown in Table 6. We can observe that all models didn’t perform very well without relevant context.
Llama2-7B-chat + LoRA and BioASQ Baseline ZS are better since they were fine-tuned on biomedical
corpora. However, when the relevant context is given, all models have a significant improvement even
more than 100% on some metrics. Therefore, Retrieval-Augmented Generation is useful and promising
when dealing with domain-specific Qustion-Answering tasks.
10http://participants-area.bioasq.org/results/11b/phaseB/</p>
          <p>Besides, we explored the influence of the retrieved context quality on the performance from two
aspects: Retrieval Unit and Retrieval Source.</p>
          <p>• Retrieval Unit The golden-enriched snippets provided in Phase B test sets are of course the best
retrieval unit for the answer generation since they are validated by biomedical experts. Therefore,
we build a baseline treating chunks split from the golden-enriched documents as retrieval units.
The ensemble retriever needs to search the most relevant snippets or chunks to provide references
for the answer generation model (Llama2-7B-chat + LoRA). The results are illustrated in Table 7,
which indicates retrieving from the golden-enriched snippets brings better performance.
• Retrieval Source Then we experiment with adding noise to the retrieval process by changing
the retrieval sources. The retrieval source means where the retrieval context comes from. We
conducted experiments in three diferent settings: ‘Test Set’, ‘Training Set’, and ‘Training + Test
Set’. The ‘Test Set’ setting is identical to the setting in Phase B. The ‘Training + Test Set’ setting
means the relevant snippets should be first searched from snippets in the training and test set,
which will bring some noise to the given snippets. The ‘Training Set’ setting means the relevant
snippets should be searched from snippets only in the training set. Notice that snippets in the
training set could be irrelevant in most cases, so it will bring a huge amount of noise. The results
in Table 8 also support the conclusion that higher quality of the retrieved contexts will lead to
better performance.</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>4.3.3. On Context Input Method</title>
          <p>When the generation model and the context were the same, we further explored diferent context input
methods. We constructed a Baseline called ‘Stuf Documents’ here by replacing the ensemble retriever
of the system CPS-B as the ‘create stuf chain’ 11 from the LangChain Package. This chain means stufing
11Langchain API: langchain.chains.combine_documents.stuf.create_stuf_documents_chain
all the relevant snippets of query questions as the context for the model input. The results in Table
9 show that introducing the retrieval process in CPS-B is more efective than just inputting all the
relevant snippets as the context.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>We constructed the Corpus PEFT Searching (CPS) system step by step following diferent phases in the
BioASQ challenge. Firstly, in phase A the system acquires access to the PubMed Central database with
a BM25 retriever. Secondly, in phase B the system integrates the Llama2-7b-chat model fine-tuned by
LoRA with an ensemble retriever. Finally, in phase A+, we used an optimized two-stage hierarchical
retrieval structure to connect the corpus and answer generation LLM. The first stage employs BM25
for a rough retrieval of documents, while the second stage combines BM25 with vector similarity for a
ifne-grained retrieval of document chunks. The CPS system can search documents related to query
questions and generate appropriate answers. It achieved a mid to slightly higher position on the BioASQ
12b leaderboard rankings. The reason may be that the scale and the context window length of our
model are relatively small.</p>
      <p>However, we made great eforts to improve its performance by utilizing diferent techniques including
sparse and dense ensemble retrievers, Parameter-Eficient Fine-Tuning (PEFT), and hierarchical
RetrievalAugmented Generation (RAG). The experiments in 4.3 demonstrate our exploration of the model
ifne-tuning and retrieval techniques. We could conclude that the techniques we employed could
significantly enhance the model performance without altering its architecture, leveraging the full
synthesis capabilities of a 7B-parameter scale LLM. The experiment and analysis elaborate the idea
that if the retrieved context contains all the necessary information to complete the task, a small-scale
LLM fine-tuned by a parameter-eficient method is fully capable of generating ideal answers. Besides,
improving the context input strategy like using ensemble retrievers also has the potential to enhance
system eficiency.</p>
      <p>
        Experiment results illustrated that our method exceeded the best level of last year’s ranking. However,
there is a gap between the system performance in last year and this year’s challenges. In this work, we
only applied the BM25 method to retrieve documents or snippets, since BM25 is still a strong retriever
in the biomedical RAG domain, according to the comparisons in Xiong et al. [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]. Dense retrievers built
for the biomedical area such as MedCPT[
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] have the potential to obtain better retrieval results. Future
endeavors will focus on in-depth research in exploring more powerful retrieval techniques to improve
the generation performance. In addition, the research on LLM has been varying. We will keep up with
the pace and conduct more in-depth research and improvement on the generation model itself.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was supported by the IdeaBooster Fund IDBF23ENG05 of The Chinese University of Hong
Kong.
arXiv:2107.03374 (2021).
[36] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al., Chain-of-thought
prompting elicits reasoning in large language models, Advances in neural information processing
systems 35 (2022) 24824–24837.
[37] W. Shi, Y. Zhuang, Y. Zhu, H. Iwinski, M. Wattenbarger, M. D. Wang, Retrieval-augmented
large language models for adolescent idiopathic scoliosis patients in shared decision-making, in:
Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology,
and Health Informatics, 2023, pp. 1–10.
[38] A. Nentidis, G. Katsimpras, A. Krithara, S. Lima López, E. Farré-Maduell, L. Gasco, M. Krallinger,
G. Paliouras, Overview of bioasq 2023: The eleventh bioasq challenge on large-scale biomedical
semantic indexing and question answering, in: International Conference of the Cross-Language
Evaluation Forum for European Languages, Springer, 2023, pp. 227–250.
[39] S. Ateia, U. Kruschwitz, Is chatgpt a biomedical expert, Exploring the Zero-Shot Performance of</p>
      <p>Current GPT Models in Biomedical Tasks (2023).
[40] A. Aksenova, T. Asamov, P. Ivanov, S. Boytcheva, Improving biomedical question answering with
sentence-based ranking at bioasq-11b, in: Conference and Labs of the Evaluation Forum, 2023.</p>
      <p>URL: https://api.semanticscholar.org/CorpusID:264441330.
[41] C.-Y. Hsueh, Y. Zhang, Y.-W. Lu, J.-C. Han, W. Meesawad, R. T.-H. Tsai, Ncu-iisr: Prompt
engineering on gpt-4 to stove biological problems in bioasq 11b phase b, in: 11th BioASQ Workshop at the
14th Conference and Labs of the Evaluation Forum (CLEF), 2023.
[42] J. Lin, X. Ma, S.-C. Lin, J.-H. Yang, R. Pradeep, R. Nogueira, Pyserini: A Python toolkit for
reproducible information retrieval research with sparse and dense representations, in: Proceedings of the
44th Annual International ACM SIGIR Conference on Research and Development in Information
Retrieval (SIGIR 2021), 2021, pp. 2356–2362.
[43] E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al., Lora: Low-rank adaptation
of large language models, in: International Conference on Learning Representations, 2021.
[44] R. Luo, L. Sun, Y. Xia, T. Qin, S. Zhang, H. Poon, T.-Y. Liu, Biogpt: generative pre-trained transformer
for biomedical text generation and mining, Briefings in bioinformatics 23 (2022) bbac409.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Achiam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Adler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ahmad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Akkaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. L.</given-names>
            <surname>Aleman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Altenschmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Altman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Anadkat</surname>
          </string-name>
          , et al.,
          <source>Gpt-4 technical report, arXiv preprint arXiv:2303.08774</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Izacard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Martinet</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Rozière</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Hambro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Azhar</surname>
          </string-name>
          , et al.,
          <article-title>Llama: Open and eficient foundation language models</article-title>
          ,
          <source>arXiv preprint arXiv:2302.13971</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Stone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Albert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almahairi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Babaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bashlykov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhargava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhosale</surname>
          </string-name>
          , et al.,
          <source>Llama</source>
          <volume>2</volume>
          :
          <article-title>Open foundation and fine-tuned chat models</article-title>
          ,
          <source>arXiv preprint arXiv:2307.09288</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] OpenAI, Introducing chatgpt,
          <year>2023</year>
          . URL: https://openai.com/blog/chatgpt, openAI Blog, OpenAI.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gottweis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sayres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Wulczyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pfohl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cole-Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Neal</surname>
          </string-name>
          , et al.,
          <article-title>Towards expert-level medical question answering with large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2305.09617</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Shen</surname>
          </string-name>
          , Chatcad:
          <article-title>Interactive computer-aided diagnosis on medical image using large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2302.07257</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cramer</surname>
          </string-name>
          ,
          <article-title>Alphafold2 and the future of structural biology</article-title>
          ,
          <source>Nature structural &amp; molecular biology 28</source>
          (
          <year>2021</year>
          )
          <fpage>704</fpage>
          -
          <lpage>705</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shen</surname>
          </string-name>
          , et al.,
          <article-title>Interpretable rna foundation model from unannotated data for highly accurate rna structure and function predictions</article-title>
          ,
          <source>arXiv preprint arXiv:2204.00300</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Hmd-amp: Protein language-powered hierarchical multi-label deep forest for annotating antimicrobial peptides</article-title>
          ,
          <source>arXiv preprint arXiv:2111.06023</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hegselmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sontag</surname>
          </string-name>
          ,
          <article-title>Large language models are few-shot clinical information extractors</article-title>
          ,
          <source>arXiv preprint arXiv:2205.12689</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piktus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Petroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Karpukhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Küttler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          , W.-t. Yih,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rocktäschel</surname>
          </string-name>
          , et al.,
          <article-title>Retrieval-augmented generation for knowledge-intensive nlp tasks</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>9459</fpage>
          -
          <lpage>9474</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>I.</given-names>
            <surname>Klerings</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Weinhandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. J.</given-names>
            <surname>Thaler</surname>
          </string-name>
          ,
          <article-title>Information overload in healthcare: too much of a good thing?</article-title>
          ,
          <source>Zeitschrift für Evidenz</source>
          ,
          <source>Fortbildung und Qualität im Gesundheitswesen</source>
          <volume>109</volume>
          (
          <year>2015</year>
          )
          <fpage>285</fpage>
          -
          <lpage>290</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Katsimpras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lima-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré-Maduell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Loukachevitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Davydova</surname>
          </string-name>
          , E. Tutubalina, G. Paliouras,
          <source>Overview of BioASQ</source>
          <year>2024</year>
          :
          <article-title>The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering</article-title>
          , in: L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Quénot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Schwab</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Soulier</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Maria Di Nunzio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Galuščáková</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fifteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ),
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Katsimpras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Paliouras, Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . García Seco de Herrera (Eds.),
          <source>Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bougiatiotis</surname>
          </string-name>
          , G. Paliouras,
          <string-name>
            <surname>BioASQ-QA</surname>
          </string-name>
          :
          <article-title>A manually curated corpus for Biomedical Question Answering</article-title>
          ,
          <source>Scientific Data</source>
          <volume>10</volume>
          (
          <year>2023</year>
          )
          <fpage>170</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H. W.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Longpre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Fedus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Brahma</surname>
          </string-name>
          , et al.,
          <article-title>Scaling instruction-finetuned language models</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>25</volume>
          (
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Oufattole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-H.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Szolovits</surname>
          </string-name>
          ,
          <article-title>What disease does this patient have? a large-scale open domain question answering dataset from medical exams</article-title>
          ,
          <source>Applied Sciences</source>
          <volume>11</volume>
          (
          <year>2021</year>
          )
          <fpage>6421</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Venigalla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Frankle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Carbin</surname>
          </string-name>
          ,
          <article-title>Biomedlm: a domain-specific large language model for biomedical text</article-title>
          ,
          <source>MosaicML. Accessed: Dec</source>
          <volume>23</volume>
          (
          <year>2022</year>
          )
          <article-title>2</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>H.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <article-title>Doctorglm: Fine-tuning your chinese doctor is not a herculean task</article-title>
          ,
          <source>arXiv preprint arXiv:2304.01097</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          , G. Fan,
          <article-title>Medchatzh: a better medical adviser learns from better instructions</article-title>
          ,
          <source>arXiv preprint arXiv:2309.01114</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>T.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Adams</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-M. Papaioannou</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Grundmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Oberhauser</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Löser</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Truhn</surname>
            ,
            <given-names>K. K.</given-names>
          </string-name>
          <string-name>
            <surname>Bressem</surname>
          </string-name>
          ,
          <article-title>Medalpaca-an open-source collection of medical conversational ai models and training data</article-title>
          ,
          <source>arXiv preprint arXiv:2304.08247</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Xi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Qiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Qin</surname>
          </string-name>
          , T. Liu, Huatuo:
          <article-title>Tuning llama model with chinese medical knowledge</article-title>
          ,
          <source>arXiv preprint arXiv:2304.06975</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>N.</given-names>
            <surname>Houlsby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Giurgiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jastrzebski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Morrone</surname>
          </string-name>
          ,
          <string-name>
            <surname>Q. De Laroussilhe</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Gesmundo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Attariyan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gelly</surname>
          </string-name>
          ,
          <article-title>Parameter-eficient transfer learning for nlp</article-title>
          ,
          <source>in: International conference on machine learning, PMLR</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2790</fpage>
          -
          <lpage>2799</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Prefix-tuning: Optimizing continuous prompts for generation</article-title>
          ,
          <source>arXiv preprint arXiv:2101.00190</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S.</given-names>
            <surname>Robertson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zaragoza</surname>
          </string-name>
          , et al.,
          <article-title>The probabilistic relevance framework: Bm25 and beyond</article-title>
          ,
          <source>Foundations and Trends® in Information Retrieval</source>
          <volume>3</volume>
          (
          <year>2009</year>
          )
          <fpage>333</fpage>
          -
          <lpage>389</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>V.</given-names>
            <surname>Karpukhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Oğuz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Edunov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          , W.-t. Yih,
          <article-title>Dense passage retrieval for open-domain question answering</article-title>
          , arXiv preprint arXiv:
          <year>2004</year>
          .
          <volume>04906</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dou</surname>
          </string-name>
          , J.-Y. Nie,
          <article-title>Retrieve anything to augment large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2310.07554</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yasunaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Linkbert: Pretraining language models with document links</article-title>
          ,
          <source>arXiv preprint arXiv:2203.15827</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yasunaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosselut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <article-title>Deep bidirectional language-knowledge graph pretraining</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>35</volume>
          (
          <year>2022</year>
          )
          <fpage>37309</fpage>
          -
          <lpage>37323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosselut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yasunaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          , Greaselm:
          <article-title>Graph reasoning enhanced language models for question answering</article-title>
          ,
          <source>arXiv preprint arXiv:2201.08860</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Comeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yeganova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Wilbur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lu</surname>
          </string-name>
          , Medcpt:
          <article-title>Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>39</volume>
          (
          <year>2023</year>
          )
          <article-title>btad651</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>G.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Zhang,</surname>
          </string-name>
          <article-title>Benchmarking retrieval-augmented generation for medicine</article-title>
          ,
          <source>arXiv preprint arXiv:2402.13178</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lu</surname>
          </string-name>
          , Genegpt:
          <article-title>Augmenting large language models with domain tools for improved access to biomedical information</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>40</volume>
          (
          <year>2024</year>
          )
          <article-title>btae075</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tworek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. P. d. O.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Edwards</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Burda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Joseph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Brockman</surname>
          </string-name>
          , et al.,
          <article-title>Evaluating large language models trained on code, arXiv preprint</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>