<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <article-id pub-id-type="doi">10.1145/3560815</article-id>
      <title-group>
        <article-title>Can Open-Source LLMs Compete with Commercial Models? Exploring the Few-Shot Performance of Current GPT Models in Biomedical Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Samy Ateia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Udo Kruschwitz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Science, University of Regensburg</institution>
          ,
          <addr-line>Universitätsstraße 31, 93053, Regensburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>3497</volume>
      <fpage>09</fpage>
      <lpage>12</lpage>
      <abstract>
        <p>Commercial large language models (LLMs), like OpenAI's GPT-4 powering ChatGPT and Anthropic's Claude 3 Opus, have dominated natural language processing (NLP) benchmarks across diferent domains. New competing Open-Source alternatives like Mixtral 8x7B or Llama 3 have emerged and seem to be closing the gap while often ofering higher throughput and being less costly to use. Open-Source LLMs can also be self-hosted, which makes them interesting for enterprise and clinical use cases where sensitive data should not be processed by third parties. We participated in the 12th BioASQ challenge, which is a retrieval augmented generation (RAG) setting, and explored the performance of current GPT models Claude 3 Opus, GPT-3.5-turbo and Mixtral 8x7b with in-context learning (zero-shot, few-shot) and QLoRa fine-tuning. We also explored how additional relevant knowledge from Wikipedia added to the context-window of the LLM might improve their performance. Mixtral 8x7b was competitive in the 10-shot setting, both with and without fine-tuning, but failed to produce usable results in the zero-shot setting. QLoRa fine-tuning and Wikipedia context did not lead to measurable performance gains. Our results indicate that the performance gap between commercial and open-source models in RAG setups exists mainly in the zero-shot setting and can be closed by simply collecting few-shot examples for domain-specific use cases. The code needed to rerun these experiments is available through GitHub*.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Zero-Shot Learning</kwd>
        <kwd>Few-Shot Learning</kwd>
        <kwd>QLoRa fine-tuning</kwd>
        <kwd>LLMs</kwd>
        <kwd>BioASQ</kwd>
        <kwd>GPT-4</kwd>
        <kwd>RAG</kwd>
        <kwd>Question Answering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        the LLMs during pre-training, retrieval augmented generation (RAG) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is often used to enable these
models to understand new concepts and be more helpful and grounded in their responses [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Several
software vendors are publishing marketing articles to advertise the usefulness of their solutions to
enable RAG for enterprises 567.
      </p>
      <p>The BioASQ challenge is a great example of a RAG setup in a specialized domain, as the participating
systems first have to find relevant biomedical papers from PubMed and extract snippets that are later
used to generate answers for biomedical questions.</p>
      <p>We set out to explore the usefulness and competitiveness of open-source models to the current SOTA
commercial oferings in a typical domain-specific RAG setup represented by the BioASQ challenge.
Compared to our last year’s approach where we only looked at the zero-shot performance of commercial
models, we now explored few-shot learning because we saw that this enables open-source models to
better follow instructions while also improving overall performance.</p>
      <p>Another aspect that we explored was how additional relevant context retrieved from Wikipedia
might aid the models in generating useful answers or relevant queries, as they might be limited in their
biomedical knowledge about entities and their synonyms.</p>
      <sec id="sec-1-1">
        <title>1.1. BioASQ Challenge</title>
        <p>
          BioASQ is "a competition on large-scale biomedical semantic indexing and question answering"[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. It is
held as a lab at the Conference and Labs of the Evaluation Forum (CLEF) conference8. The current 2024
workshop is the 12th installment of the BioASQ competition9.
        </p>
        <p>
          The 12th BioASQ challenge comprises several tasks:
• BioASQ Task Synergy On Biomedical Semantic QA For Developing Issues [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
• BioASQ Task B On Biomedical Semantic QA [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
• ioASQ Task MultiCardioNER On Mutiple Clinical Entity Detection In Multilingual Medical
        </p>
        <p>
          Content [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
• BioASQ Task BioNNE On Nested NER In Russian And English [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
        </p>
        <p>
          We participated in Task B and Synergy[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. For Task B the participants’ systems receive a list of
biomedical questions that should be answered with a short paragraph style answer and some require
an additional exact answer which can be one of 3 formats, yes/no, factoid (a list of up to 5 entities) or
list (a list of up to 200 entities). Additionally, the systems first have to retrieve relevant papers from the
PubMed annual baseline and extract relevant snippets from these papers that could aid in answering
the questions. This retrieval subtask of Task B is called Phase A while the actual question answering
subtask is called Phase B. For Phase B the systems also receive a set of gold snippets and documents that
should help them answer the question. Task B was scheduled in 4 batches with two weeks in between
and ran from March 28 to May 11.
        </p>
        <p>In the 12th installment, another Phase A+ was introduced, where the systems were supposed to
provide answers to the questions before the gold snippets and documents were provided, relying solely
on their own retrieved documents and snippets.</p>
        <p>For the Synergy task, the systems receive a similar list of questions, for which they also have to retrieve
useful papers and extract snippets and as soon as a question is marked as ready to answer, they also
need to submit answers in the same format as for Task B. The diference between Synergy and Task B is
that initially in the first round no gold set of documents and snippets is provided, instead the submitted
documents and snippets by the systems are evaluated by biomedical experts and selected as gold
reference items for subsequent rounds. This also means that the same questions might be reintroduced
5https://cohere.com/blog/five-reasons-enterprises-are-choosing-rag
6https://www.pinecone.io/learn/retrieval-augmented-generation/
7https://gretel.ai/blog/what-is-retrieval-augmented-generation
8https://clef2024.clef-initiative.eu/
9http://www.bioasq.org/
in subsequent rounds, possibly with additional questions and positive and negative feedback on the
previously submitted documents.</p>
        <p>Following this introduction, we will highlight some related work in Section 2, describe our
methodology in Section 3, report our results in Section 4 and discuss them in Section 5. Section 6 will present
some ethical considerations, and Section 7 ofers our conclusions.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>2.1. GPT Models</title>
        <p>We will briefly introduce the related work that led to the creation of the evaluated models, as well as
the approaches that inspired our methodology.</p>
        <p>
          Nearly all the popular SOTA LLMs that are used across various NLP tasks and use cases today are based
on the transformer architecture [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. With the generative pretrained transformer (GPT) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] being a
popular variant. These models undergo pre-training on vast amounts of text by solving the next-token
prediction task [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Afterwards, the models are fine-tuned to align with human preference data [ 10]
which enables them to follow instructions and be useful in direct interactions with users.
        </p>
        <p>OpenAI was the first company that released such a fine-tuned model to the public in November 2022 10
which sparked massive interest in generative artificial intelligence research and products. Their latest
model at the time of writing that is powering their ChatGPT product was GPT-4 [11]. One interesting
competitor model that we also used during this competition is Claude 3 Opus11 by Anthropic, which
reached GPT-4 level performance (GPT-4-0125-preview) at the time of the BioASQ competition12. The
exact architecture of GPT-4 and Claude Opus 3 and other commercial models is unknown.</p>
        <p>GPT-4 is the most expensive and slowest model that OpenAI is ofering via their API service. A
more afordable alternative that they are ofering is GPT-3.5-turbo. We compared both these models’
performance in last year’s BioASQ competition and were able to show that GPT-3.5-turbo was sometimes
performing better than GPT-4 in some question formats and subtasks of the competition [12].</p>
        <p>For this year’s BioASQ competition we also used Mixtral 8x7B [13] a downloadable open-source model
(Apache 2.0 license) which uses a type of Mixture-of-Experts Architecture [14][15]. This architecture
ofers higher computational eficiency by routing requests to expert subnetworks to generate a response.
Since only some specialized experts are active during the generation, less computation and memory is
needed to serve the request.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Few and Zero-Shot Learning</title>
        <p>Few-Shot learning is the ability of LLMs to learn how to solve a new problem that they were not
specifically fine-tuned for by only showing them a few examples. When GPT-3 was first introduced, its
impressive few-shot learning abilities made the concept popular [16] because it greatly reduces the
need for expensive training data.</p>
        <p>Zero-shot learning [17] takes this concept a step further by only requiring an abstract task description
or direct question, which leads the model to ideally generate a useful completion that solves the task
at hand [18]. In last year’s BioASQ competition, we were able to win some batches while only using
zero-shot learning with SOTA commercial models [12].</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Adapter Fine-Tuning</title>
        <p>Current LLMs have billions of parameters and require specialized hardware with enough GPUs and
VRAM to hold all the model weights in memory. For example, "16-bit finetuning of a LLaMA 65B
10https://web.archive.org/web/20240502090536/https://openai.com/index/chatgpt/
11https://web.archive.org/web/20240516173322/https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/</p>
        <p>Model_Card_Claude_3.pdf
12https://chat.lmsys.org/?leaderboard
parameter model requires more than 780 GB of GPU memory"[19]. Since this makes fine-tuning these
models prohibitively expensive for many researchers and users, several clever techniques have been
invented to reduce these hardware requirements. We wanted to fine-tune Mixtral 8x7B, which roughly
takes up the memory of a 47B model13.</p>
        <p>One popular approach is QLora by Dettmers et al. [19] where the model weights are quantized
to 4 bits and frozen and only some low rank adapters (LoRa)[20] are fine-tuned. This would enable
ifne-tuning Mixtral 8x7B, for example, on only two RTX A6000 GPUs with 2x 48 GB of VRAM.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Retrieval Augmented Generation (RAG)</title>
        <p>
          Retrieval augmented generation (RAG) is a technique [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] that combines information retrieval with
language models to enhance their ability to generate relevant and factual text. In RAG, the language
model is augmented with an external knowledge base or other information source, such as a collection
of documents or web pages. When generating text, the model first retrieves relevant information,
based on the input query, and then uses that information to guide the generation process. This process
is applied in the BioASQ challenge, where the relevant information source is the annual baseline of
PubMed.
        </p>
        <p>
          RAG has been shown to improve the factual accuracy of generated text compared to standalone
language models [21]. It allows the model to access a vast amount of external knowledge and incorporate
it into the generated output. RAG is particularly useful for tasks that require domain-specific knowledge
or up-to-date information [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Professional Search</title>
        <p>Professional search is conducted in a professional context, often to aid in work-related research tasks
[22]. In some professional search settings, highly trained specialists are needed to create documented
and reproducible search strategies, this sets professional search apart from everyday web search [23].
The BioASQ challenge exemplifies one possible professional search setting where biomedical experts
aim to find answers to domain-specific questions with suficient evidence.</p>
        <p>Other examples of professional search might be systematic reviews [24], patent-search or search
conducted by recruitment professionals [25]. All of these settings might require complex search
strategies, where the search expert makes use of a query syntax involving boolean operators on specific
search fields. Systematic reviews, for example, also require the search to be explainable and reproducible,
which makes it dificult to use advanced vector-based retrieval techniques. Formulating traditional
queries but with large language models that might be able to expand synonyms and related terms based
on their semantic representations is therefore an interesting approach that might aid in professional
search settings. We set out to explore this approach in the BioASQ challenge.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>3.1. Model
In this year’s BioASQ competition, we looked at the commercial oferings GTP-3.5-turbo and GPT-4
from OpenAI and also used Antrophics Claude 3 Opus, which was at the time of the run submissions the
only other model that was on a level with GPT-4 according to the LMSYS Chatbot Arena Leaderboard
[26].</p>
      <p>Since last year’s BioASQ competition, some competitive Open-Source models were published. The
most notable ones being the Llama series models by Meta [27], with the latest being Llama 3 [28].
Llama comes with its own custom License which is quite permissive, but even though commercial
use is allowed under this license, as long as the monthly user base does not exceed 700 million users,
13https://mistral.ai/news/mixtral-of-experts/</p>
      <p>Listing 1: Query Expansion Prompt
{"role": "user", "content": f"""Turn the following biomedical question into an efective elasticsearch query
using the query_string query type by incorporating synonyms and additional terms that closely relate
to the main topic and help reduce ambiguity. Focus on maintaining the query’s precision and relevance
to the original question, the index contains the fields ’title’ and ’abstract’, return valid json: ’{question}’
"""}
the license might not be straightforward to adopt for enterprise use cases when licenses have to be
pre-approved by a legal team.</p>
      <p>When we prepared our runs for the competition, the best-performing model with a permissive
open-source license (Apache 2.0) on the LMSYS leaderboard was Mixtral 8x7B [13]. The model also
has a large context length of 32k tokens, which makes it especially interesting for RAG use cases or
few-shot learning. We therefore choose Mixtral 8x7B as our open-source competitor model for this
competition. During the competition, the newer Mixtral 8x22B model was also published, and we used
it in some batches of Task B.</p>
      <p>We used the commercial hosting service fireworks.ai 14 to access and fine-tune Mixtral 8x7B, as the
provided speed was very high and costs for their API usage were low.</p>
      <sec id="sec-3-1">
        <title>3.2. Synergy</title>
        <p>We downloaded and indexed the annual PubMed baseline from the oficial website 15. We indexed both
the title and the abstract of all papers in separate fields of our index using the built-in English analyzer of
Elasticsearch16. For every round of synergy, the most recent snapshots of 2024 up to the date considered
in that round were downloaded and indexed in another similar index, which was then also searched
during the runs.</p>
        <p>For synergy, we used both gpt-4-0125-preview and gpt-3.5-turbo-0125, the newest available versions
of OpenAI’s GPT-4 and GPT-3.5-turbo at the time of the competition. We used 2-shot learning
to generate queries for our Elasticsearch PubMed index and zero-shot learning for extracting and
reranking snippets as well as answering questions. We also wanted to use Mixtral 8x7B in this task,
but the model was unable to follow instructions well enough to produce usable runs, especially in the
zero-shot setting.</p>
        <p>Given a question, we prepended the prompt in Listing 1 with two examples where the same prompt
contained other questions, for example "Is CircRNA produced by back splicing of exon, intron or both,
forming exon or intron circRNA?" and an ideal completion in form of an Elasticsearch query for the
query_string endpoint of Elasticsearch for which an example can be seen in Listing 2. We sent both the
examples and the prompt with our actual question to the model and received back a JSON object that
could be used to directly query our index.</p>
        <p>We ran the generated query to retrieve the top 50 relevant documents from Elasticsearch. We filtered
out documents that were marked as irrelevant in the feedback file for the synergy round. We then sent
each remaining article alongside the question to the model and used a zero-shot prompt to extract a
list of relevant snippets from the article. We then used string matching to insure the returned snippets
were actually present in the article title or abstract and to calculate the ofsets.</p>
        <p>We collected all relevant snippets from all potentially 50 articles and filtered out articles as irrelevant
where the model did not extract any snippets from. We then prompted the model to select the top
10 snippets ranked by helpfulness from the set of snippets. We finally reranked the retrieved articles
according to the snippet order returned by this step.
14https://fireworks.ai/
15https://pubmed.ncbi.nlm.nih.gov/download/
16https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-lang-analyzer.html#english-analyzer</p>
        <p>Listing 2: Query Expansion Competion Example
{"role": "assistant", "content": """
{
"query": {
"query_string": {
"query": "(CircRNA OR \"circular RNA\") \"back splicing\" exon OR intron",
"fields": [
"title^10",
"abstract"
],
"default_operator": "and"
}
"""},</p>
        <p>In the question answering step, we used the identical zero-shot prompts from our last year’s
participation in Task B for this year’s Synergy task [12]. And we also merged the already deemed relevant
snippets from the feedback files into our list of snippets that we passed on to the modal alongside the
question prompt.</p>
        <p>We also sent the same initial system prompt that we used last year [12] to the models. For the
parameters, we set the temperature parameter to 0 to reduce randomness in the completion and supplied
a seed parameter which is a new feature ofered by OpenAI that can help maximize reproducibility for
the model output, but determinism is still not guaranteed17. We also used the new response_format
parameter to insure the model produced valid JSON responses for the prompts where we needed it18.
The exact python notebooks with all the implementation details and prompts used are available in our
GitHub repository.
3.3. Task 12 B
For Task B we reused the indexed PubMed annual snapshot that we created from the synergy task. We
also switched the models, we added Mixtral 8x7B Instruct v0.1 as an open-source model, and we used
Claude 3 Opus instead of GPT-4 because it became available shortly before the Phase started, and we
had access to a free beta evaluation account.</p>
        <p>In batches 1 &amp; 2 we also explored the fine-tuning service of OpenAI and created 6 fine-tuned versions
of GPT-3.5-turbo for each sub problem that our system had to solve. We also created 6 QLoRa fine-tuned
versions of Mixtral 8x7B Instruct v0.1 using the fine-tuning service of fireworks.ai. The training sets
were created from the supplied training data. We also adjusted our code to be able to use these training
ifles as sources for few-shot examples.</p>
        <p>The sub problems that we sampled training sets for were:
• Snippet Extraction
• Snippet Reranking &amp; Selection
• Summary Question Answering
• Exact yes/no Question Answering
• Exact factoid Question Answering
• Exact list Question Answering
17https://platform.openai.com/docs/api-reference/chat/create#chat-create-seed
18https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format</p>
        <p>Listing 3: Wikipedia Retrieval Prompt
prompt = f"""</p>
        <p>Given the question "{question}", identify existing Wikipedia articles that ofer helpful background
information to answer this question.</p>
        <p>Ensure that the titles listed are of real articles on Wikipedia as of your last training cut−of. Wrap
the confirmed article titles in hashtags (e.g., #Article Title#).</p>
        <p>Provide a step−by−step reasoning for your selections, ensuring relevance to the main components of
the question.</p>
        <sec id="sec-3-1-1">
          <title>Step 1: Confirm the Existence of Articles Before listing any articles, briefly verify their existence by ensuring they are well −known topics generally covered by Wikipedia.</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>Step 2: List Relevant Wikipedia Articles After confirming, list the articles, wrapping the titles in hashtags and explaining how each article is relevant to the question. """</title>
          <p>In batches 3&amp;4 we explored how we could use the models to retrieve relevant additional context from
Wikipedia. We hypothesized that supplying these models with knowledge about relevant entities in the
questions might improve their ability to generate correct answers. We wanted to explore the approach
of retrieving such additional information about entities from a Wiki, because such a wiki like could
also be generated in an enterprise setting, potentially closing the knowledge gap for entities that the
models didn’t encounter during pre-training. We chose Wikipedia as a knowledge base even though the
concepts described there might not be novel to the models because it was easy to use, and we hoped
that we could observe an efect even with known concepts. We suspected that these models might just
"know" which Wikipedia articles might be relevant to a question because they might have been highly
trained on Wikipedia and links to Wikipedia articles. The exact zero-shot prompt that was used for
ifnding relevant Wikipedia articles can be seen in Listing 3.</p>
          <p>The titles of the Wikipedia articles returned by this prompt were extracted and, if the Wikipedia
articles actually existed, their content was retrieved and concatenated. Finally, the concatenated articles
were again sent to the model, and it was prompted to make a concise summary to help answer the
question. This summary was then added as additional context to all subsequent prompts such as query
generation or snippet extraction, reranking and question answering.
3.3.1. Phase A
In Phase A, we changed the way we prompted the models for queries compared to the Synergy task.
Instead of expecting the whole valid JSON query object back, we only prompted the models to create
the query string in the valid query_string query syntax to pass on to the Elasticsearch endpoint and
manually controlled weighting of the fields and the document return size.</p>
          <p>We then used the questions from batch 1 from last year’s task, that were provided in the development
set, to create a set of Elasticsearch queries with Claude 3 Opus. Finally, we ran these queries and
evaluated the returned documents against the gold set to select 10 queries with the highest f1 score as
few-shot examples for all models. We used 10 examples for most of our few-shot tasks because they fit
in most models context lengths, with GPT-3.5-turbo having the smallest context length of 16k tokens.</p>
          <p>The rest of our approach to retrieving relevant documents and snippets was mostly in line with our
approach from the Synergy task, except that we re-added the step that we also used last year when we
prompted the models for an improved query when a generated query did not return any results.</p>
          <p>Compared to our system from last year, we changed the query expansion/generation prompt, added
the possibility to prepend few-shot examples to the prompts and used diferent models where some
were also fine-tuned. We also reranked snippets instead of titles, and filtered and ranked the articles
according to their extracted snippets. We also used our own Elasticsearch index of the PubMed annual
baseline instead of the online PubMed search endpoint.
3.3.2. Phase A+ and Phase B
Our approach for Phase A+ and Phase B were mostly identical. We used the same prompts and few-shot
examples taken from the sampled training sets for model fine-tuning. The only diference was the
relevant snippets provided as context to the models. For Phase A+ we used the same snippets as input
for all models, these were taken from the run in Phase A where the most snippets were found. We
opted to use the same snippets, as we wanted to be able to compare the performance of the diferent
models in Phase A+ isolated from their performance in Phase A. For Phase B we took the gold snippets
provided in the run file as input.</p>
          <p>The prompts used for Phase A+ and Phase B were identical to the question answering prompts from
the Synergy task and our last year’s approach. We only added the option to add additional context
before the snippets and supply few-shot examples. In batches 1 &amp; 2 we compared the performance of
ifne-tuned models with their non-fine-tuned counterparts, and in batches 3 &amp; 4 we compared systems
with additional relevant context taken from Wikipedia with systems that did not have this context.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. Synergy</title>
        <p>We participated with our systems in two tasks of the BioASQ challenge: Synergy and Task B On
Biomedical Semantic QA. For the synergy task, we report the results only for batch 3 and 4 as we were
unable to participate in earlier batches. For Task B we competed in all batches and report the results of
sub-tasks A (Retrieval), A+ (Q&amp;A with own retrieved documents) and B (Q&amp;A with gold documents).
The results presented in this section are only preliminary, as the manual assessment of the system
responses by the BioASQ team of biomedical experts is still ongoing. The final results will be available
on the BioASQ homepage once the manual assessment is finished 19.</p>
        <p>We participated with 2 systems in batch 3 and 4 of the Synergy task. The full result table is accessible
on the BioASQ website20. The two systems both used the same 2-shot query expansion and zero-shot
Q&amp;A approach but with diferent commercial models. The system names and corresponding models
are listed below:
• UR-IW-2: gpt-4-0125-preview.</p>
        <p>• UR-IW-3: gpt-3.5-turbo-0125</p>
        <p>We tried to also use Mixtral 8x7B Instruct v0.1 as an open-source alternative in this task, but the
model was unable to follow the zero-shot prompts for query expansion and extracting the snippets
consistently enough to produce a submittable run file. Our retrieval approach for query expansion and
ifltering and reranking via extracted snippets was not competitive in the document retrieval stage of
the task, with both models performing similarly poorly compared to the top competitors, as indicated
in Table 1. The oficial metrics to rank systems in each subtask are highlighted in bold in the following
tables21.
19http://participants-area.bioasq.org/results/
20http://participants-area.bioasq.org/results/synergy_v2024/
21"Top Competitor" are the systems that took the first position in a round or batch that are not ours. They are added as a
reference point for the reported metrics. When "Top Competitor" is missing in a reported batch, one of our systems was the
best-performing one.</p>
        <p>For snippet extraction, the performance of our approach was also poor, except for batch 4 where
gpt-3.5-turbo-0125 was able to achieve second place as can be seen in Table 2. Gpt-4-0125-preview was
unable to extract any snippets in the same batch, which was unusual and could have been due to issues
OpenAI had with serving the preview model via their API.</p>
        <p>For the question-answering stage, both models achieved perfect scores in the exact yes/no answer
format, see Table 3. For factoid answers, gpt-4-0125-preview was able to take first place in batch 4 and
achieved higher placements than gpt-3.5-turbo-0125 over both batches, as can be seen in Table 4. A
similar diference is also observable in Table 5 for the list answer results.</p>
        <p>Completing one run with gpt-4-0125-preview cost around $12 in API fees, while the same run with
gpt-3.5-turbo-0125 was around 10 times cheaper at $1.2. Gpt-4-0125-preview was also quite slow, taking
around 180 minutes to complete one run, while gpt-3.5-turbo-0125 took only a few minutes. Cost
significantly decreased compared to our last year’s participation, while speed increased. This enabled
us to actually use GPT-4 on the snippet extraction task, while last year we were not able to complete
runs with snippet extraction for GPT-4 due to time and cost constraints. We also encountered less to
none API errors during our runs, except for the empty responses in batch 4 for GPT-4, while last year
we often had to rerun questions because the API timed-out or returned other errors.
4.2. Task 12 B Phase A
We participated with 5 systems in all 4 batches of Task 12 B Phase A. The systems either used 1- or
10-shot learning with the plain or a fine-tuned model in batches 1+2 or additional context retrieved
from Wikipedia in batches 3-4. The system names and configurations are listed below.</p>
        <p>Batches 1-2:
• UR-IW-1: Claude 3 Opus + 1-shot
• UR-IW-2: Mixtral 8x7B Instruct v0.1, QloRa Fine-Tuned + 10-Shot
• UR-IW-3: gpt-3.5-turbo-0125 fine-tuned + 1-shot
• UR-IW-4: Mixtral 8x7B Instruct v0.1 + 10-shot
• UR-IW-5: gpt-3.5-turbo-0125 + 10-shot
• UR-IW-1: Claude 3 Opus 1-shot + wiki
• UR-IW-2: Claude 3 Opus 1-shot
• UR-IW-3: Mixtral 8x7B Instruct v0.1 10-Shot + wiki
• UR-IW-4: Mixtral 8x7B Instruct v0.1 10-Shot
• UR-IW-5: Mixtral 8x22B Instruct v0.1 10-Shot + wiki</p>
        <p>The following Tables 6 and 7 show the results of our systems participating in the 4 batches. MAP
was the oficial metric to compare the systems.</p>
        <p>In batches 1 &amp; 2 where we compared fine-tuned versions of GTP-3.5 and Mixtral 8x7b with 10-shot
learning and Claude 3 Opus, no clear trend was observable over batches in the document retrieval
stage of Phase A as can be seen in Table 6. Only Mixtral 8x7B with 10-shot learning was consistently
performing worse than all our other models, While Claude 3 Opus with 1-shot learning was our best
model in batch 2 and second best in batch 1.</p>
        <p>We used 1-shot learning instead of 10-shot learning for Claude 3 Opus due to time constraints because
the model was slow, and for the fine-tuned gpt-3.5-turbo due to cost constraints. Sending 10 examples
of abstracts for snippet extractions per 50 highest-ranked search results would have amounted to quite
some input tokens per run for these models.</p>
        <p>For batch 3 &amp; 4 where we explored if giving the systems additional Wikipedia context while creating
queries for Elasticsearch could improve their performance, we also could not observe a consistent efect
over batches. While in batch 3 the systems with additional Wikipedia context performed better, this
efect was reversed in batch 4.</p>
        <p>In the snippet extraction stage of Phase A, our QloRa fine-tuned version of Mixtral 8x7B was our
worst-performing system in both batch 1 &amp; 2, followed by the 10-shot Mixtral 8x7B version as can
be seen in Table 7. The fine-tuned version of GPT-3.5-turbo was consistently ahead of its 10-shot
counterpart in both batches, while Claude 3 Opus was worse than the GPT-3.4-turbo systems in batch 1
and better in batch 2.</p>
        <p>Additional Wikipedia context did not lead to consistent results across batches. While the systems with
Wikipedia context (UR-IW-1, UR-IW-3) performed better than their counterparts (UR-IW-2, UR-IW-4)
in batch 3, this efect was again reversed in batch 4.</p>
        <p>One run with Claude 3 Opus with 1-shot learning and additional Wikipedia context took around
140 minutes to complete, we did not have to pay for the tokens used because we had an early beta
evaluation account. For Mixtral 8x7B with 10-shot learning and additional Wikipedia context, the runs
took around 14 minutes to complete via the fireworks.ai API and cost around $ 11. Fireworks.ai charged
$0.50 /1M tokens for both input and output tokens as of writing, while Anthropic would have charged $
15 /1M tokens for input and $ 75 /1M tokens for output. So the cost for doing 10-shot learning with
Claude 3 Opus would have been at least 30 times as high while being 10 times slower.
4.3. Task 12B Phase A+
We participated with 5 systems in nearly all 4 batches of Task 12 B Phase A+22. The systems either used
10-shot learning with the plain or a fine-tuned model in batches 1+2 or additional context retrieved
from Wikipedia in batches 3-4. Per batch, we used the same input snippet file for all systems to base
their answers on to ensure that their performance is comparable.</p>
        <p>Batches 1-2:
• UR-IW-1: Claude 3 Opus + 10-shot
• UR-IW-2: Mixtral 8x7B Instruct v0.1, QLoRa Fine-Tuned + 10-Shot
• UR-IW-3: gpt-3.5-turbo-0125 fine-tuned + 10-shot
• UR-IW-4: Mixtral 8x7B Instruct v0.1 + 10-shot
• UR-IW-5: gpt-3.5-turbo-0125 + 10-shot
• UR-IW-1: Claude 3 Opus 10-shot + wiki
• UR-IW-2: Mixtral 8x22B Instruct v0.1 10-Shot
• UR-IW-3: Mixtral 8x7B Instruct v0.1 10-Shot + wiki
• UR-IW-4: Mixtral 8x7B Instruct v0.1 10-Shot
• UR-IW-5: Mixtral 8x22B Instruct v0.1 10-Shot + wiki
22We failed to submit one system run for system number 5 in batch 3.</p>
        <p>For the yes/no exact answer format, Claude 3 Opus with 10-shot learning was our worst-performing
system, while the fine-tuned version of GPT-3.5-turbo was our top-performing system in batch 1 and
only beaten by its 10-shot counter-part in batch 2, as can be seen in Table 8. It was interesting to
see that the open-source models could perform better than the presumably most advanced
commercial model, Claude 3 Opus, in this task.</p>
        <p>For batches 3 &amp; 4, we could show that additional Wikipedia context led to inconsistent results. While
this context improved performance in batch 3 for the Wikipedia enhanced systems (UR-IW-5, UR-IW-3)
over their normal 10-shot counterparts (UR-IW-2, UR-IW-4) it again led to worse performance in batch
4. This result is in line with the results from Phase A where these systems performed similarly for
document retrieval and snippet extraction. We speculate that the models are sensitive to the Wikipedia
context, and the usefulness of the context is highly influenced by both the entities present in the
questions, and its relationship to the relevant snippets.</p>
        <p>In the exact answer factoid format, our worst-performing system in batches 1 &amp; 2 was consistently
Mixtral 8x7B with 10-shot learning while its fine-tuned counterpart performed better as can be seen
in Table 9. This order was reversed for GPT-3.5-turbo, where the fine-tuned version performed worse
than its counterpart.</p>
        <p>The additional Wikipedia context again led to inconsistent results across batches, but this time the
systems with Wikipedia context performed better in batch 4 compared to batch 3, which is contrary
to the observed behavior in the document retrieval and snippet extraction in Phase A as well as the
yes/no answer format in Phase A+.</p>
        <p>For the list exact answer format, Mixtral 8x7B was again our worst-performing system in batch 1 &amp;
2, while its fine-tuned counterpart was competing with the fine-tuned version of gpt-3.5-turbo for the
top positions as can be seen in Table 10.</p>
        <p>In batches 3 &amp; 4 Claude 3 Opus with 10-shot learning and additional Wikipedia context was the
bestperforming system in both batches while Mixtral 8x7B with 10-shot learning and additional Wikipedia
context was the worst-performing one. Overall no consistent efect of the Wikipedia context was
observable across models and batches.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.4. Task 12B Phase B</title>
        <p>We participated with 5 systems in all 4 batches of Task 12B Phase B. The systems used 10-shot learning
with the plain or a fine-tuned model in batches 1 &amp; 2 or additional context retrieved from Wikipedia in
batches 3 &amp; 4.</p>
        <p>Batches 1-2:
• UR-IW-1: Claude 3 Opus + 10-shot
• UR-IW-2: Mixtral 8x7B Instruct v0.1, QLoRa Fine-Tuned + 10-Shot
• UR-IW-3: gpt-3.5-turbo-0125 fine-tuned + 10-shot
• UR-IW-4: Mixtral 8x7B Instruct v0.1 + 10-shot
• UR-IW-5: gpt-3.5-turbo-0125 + 10-shot
Batch 3
Batch 4
• UR-IW-1: Claude 3 Opus 10-shot + wiki
• UR-IW-2: Mixtral 8x22B Instruct v0.1 10-Shot
• UR-IW-3: Mixtral 8x7B Instruct v0.1 10-Shot + wiki
• UR-IW-4: Mixtral 8x7B Instruct v0.1 10-Shot
• UR-IW-5: Mixtral 8x22B Instruct v0.1 10-Shot + wiki
• UR-IW-1: Claude 3 Opus 10-shot + wiki
• UR-IW-2: Claude 3 Opus 10-shot
• UR-IW-3: Mixtral 8x7B Instruct v0.1 10-Shot + wiki
• UR-IW-4: Mixtral 8x7B Instruct v0.1 10-Shot</p>
        <p>Batch
Test Batch 1
Test Batch 2</p>
        <p>In the exact yes/no answer settings of Phase B, Claude 3 Opus with 10-shot learning and gpt-3.5
turbo with 10-shot learning were sharing first place in batch 1 while the fine-tuned version of
gpt-3.5-turbo was the best-performing system in batch 2 as can be seen in Table 11. The fine-tuned
version of Mixtral 8x7B was also competitive, taking the 5th position in batch 1 and 8th position in
batch 2.</p>
        <p>In batch 3 &amp; 4 with additional wikipedia context the systems performed better than their counterparts
in batch 3 while this result was mixed in batch 4, leading again to inconsistent results.</p>
        <p>For the exact factoid answer format in Phase B, Claude 3 Opus and our fine-tuned Mixtral 8x7B
model where sharing first place in batch 2 while in batch 3, Claude 3 Opus and gpt-3.5-turbo where
the on the same level as can be seen in Table 12.</p>
        <p>For Mixtral 8x7B, additional Wikipedia context improved the outcome in batch 3 but led to worse
results in batch 4, while a similar efect was observable for Claude 3 opus in batch 4 23.</p>
        <p>For the exact list answer format in Phase B, Mixtral 8x7B with 10-shot learning took first place in
batch 2 while being on place 22 of 39 in batch 1 where gpt-3-5-turbo with 10-shot learning was our
best-performing one as can be seen in Table 13.</p>
        <p>For batches 3 &amp; 4 both Claude 3 Opus and Mixtral 8x7B performed worse with additional Wikipedia
context across batches. In batch 3 our best-performing system was Mixtral 8x7B with 10-shot learning
and in batch 4 it was Claude 3 Opus with 10-shot learning.</p>
        <p>The costs for completing runs in Phase B were lower than in Phase A and the runs were faster
because we did not have to do snippet extraction for 50 documents per question, times the number of
few-shot examples, we therefore were able to also do 10-shot learning with Claude 3 Opus and the
ifne-tuned version of GPT-3.5-turbo. The processing with Mixtral 8x7B via the Fireworks.ai API only
took around 30 seconds for plain 10-shot examples and around 2 minutes for 10-shot examples and
additional Wikipedia context.</p>
        <p>We also submitted ideal answers for Task B and A+, but do not report on the preliminary results
here, as the oficial judging metric for this answer type is based on the manual judgements that are not
available yet.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Future Work</title>
      <p>While testing both commercial and open-source models, we could observe that there was no clear
dominating model across batches or sub-tasks. Even our presumably weakest model, Mixtral 8x7B
Instruct v0.1 with 10-shot learning was able to secure some leading spots in some batches of the
competition, beating all other competing systems, (see batch 3 in Table 8 and batch 2 in Table 13). We
23The run of Mixtral 8x22B with Wikipedia context was not successfully submitted to batch 3, we might have overlooked to
upload them.
speculate that both the RAG setting and 10-shot learning might level the playing field a bit between
commercial and open-source models, and it indicates that there is clear potential for creating
state-ofthe-art systems even with cheaper, faster and presumably smaller open-source models, if they are used
in the right way.</p>
      <p>While the Mixtral model weights are publicly available, their training data is not published, which
makes these models not ideal candidates for scientific research. A truly open-source LLM alternative is
OLMo [29] published by the Allen Institute for AI. We choose Mixtral nevertheless because we wanted
to study models that might be used by commercial practitioners in clinical or enterprise use cases,
and we think the permissive license combined with the seemingly competitive performance on public
benchmarks and its large context length makes it an ideal candidate for these use cases.</p>
      <p>From our results with our experiments with additional Wikipedia context in BioASQ we could see that
it had an impact on performance, but it was inconsistent across question batches. For some questions
in some subtasks it led to improvements, while for others the performance declined. We speculate
that this might be dependent on the relevant entities in the questions and the quality of the retrieved
Wikipedia context. Further experiments are needed to analyze the impact of this additional context.</p>
      <p>We also speculate that Wikipedia might not be a good proxy knowledge base for doing
domainspecific RAG for these models because they are probably already highly trained on Wikipedia data
and therefore the additional knowledge from this source might not tell the models much that they not
already know.</p>
      <p>Another reason for the inconsistent results with Wikipedia context could be that we only prepended
the context for the last prompt, and the preceding n-shot examples were not generated with taking this
context into account.</p>
      <p>Regarding fine-tuning, we had the impression that the commercial ofering from OpenAI was not
worth the cost. Even though it led to top results in some batches (batch 3 in Table 11) it also produced
models with worse results than their significantly cheaper non fine-tuned counterpart in others. Maybe
with the right training set and the right training run you might get a consistently superior model, but
then you also have to add the engineering cost compared to simple 10-shot learning to the already more
expensive usage and fine-tuning costs.</p>
      <p>We had a similar impression regarding the QLoRa fine-tuning that we explored for Mixtral 8x7B. For
example, the Mixtral model fine-tuned for list question answering was performing better than all other
systems in batch 1 of Phase A+ (see Table 10) but worse than its non fine-tuned counterpart in batch 2
of Phase B (see Table 13). Overall, adapter fine-tuning appears not to be straightforward and requires
more time for dataset creation, training and testing than simple few-shot learning. A more promising
research direction might be selecting optimal few-shot examples for a given task.</p>
      <p>It is important to note that most of the results we presented here are preliminary and might change
when the manual assessment of the system responses is completed by the BioASQ experts. But we
expect that yes/no answers will stay the same and factoid and list answers might just change slightly.</p>
      <p>The preliminary performance of our systems in the document retrieval stage (Phase A) was quite
poor compared to the other systems. We speculate that our approach of relying on TF_IDF-based
retrieval and adding a richer semantic representation to the keyword query instead of using embeddings
and vector search to add such information is not in line with the baseline system used to create the
preliminary gold set. If that is true, it might be possible that our retrieval performance is actually better
than we expect and the performance improves when the final results are out. But it could also just be
that the approach is inferior. The good performance of our systems in Phase A+ where the questions
had to be answered without gold snippets, might indicate that our retrieved snippets are not as useless
as the preliminary results from Phase A suggest.</p>
      <p>For future work, we would like to further explore optimal few-shot example selection [30][31][32],
as few-shot learning seems to ofer the best flexibility and requires less engineering efort compared to
ifne-tuning while being transferable between models. We also would like to revisit our knowledge base
context augmentation approach, beyond the BioASQ challenge. We think that on a more technical test
set where the relevant knowledge is highly unlikely to be present in the pre-training of these models,
this approach could have a bigger impact. We also would need to use a diferent knowledge source as
the English Wikipedia, which is part of most pre-training datasets [33].</p>
    </sec>
    <sec id="sec-6">
      <title>6. Ethical Considerations</title>
      <p>The current generation of LLMs still exhibit the phenomenon of so-called hallucinations [34] that is
they sometimes make up factually incorrect statements and even harmful misinformation. LLMs might
also reproduce myths or misinformation that they encountered during their open-domain training
or that might have been added to their input context. A recent prominent case was Google’s new AI
overview feature suggesting users should add glue to their pizza24.</p>
      <p>The hallucination and misinformation problems seem fundamental and dificult to solve as they
have been known for quite some time now and even Google, one of the most experienced AI research
companies, was unable to save itself from repeated embarrassment.</p>
      <p>Even though RAG has been shown to reduce hallucinations in some settings [21] occasional
hallucinations might still happen, which could be especially problematic in biomedical use cases and
might warrant additional manual fact checking [35] before using the output of LLM-based systems in
downstream tasks [36].</p>
      <p>Another issue to consider is data privacy. These models might repeat their training data if prompted
in a specific way [ 37]. That means training data has to be carefully anonymized before training or
ifne-tuning these models. The same problem arises for the few-shot examples, personal data should be
removed from all context that these models might repeat. They also might just make up facts about
people, which could put vendors and service providers at legal risk for defamation25.</p>
      <p>Another big ethical issue is job replacement and automation. Klarna, a financial service provider and
one of the early enterprise customers of OpenAI published a report in February 2024 stating that their
AI-powered customer support assistant handled one third of their customer support requests "doing
the equivalent work of 700 full-time agents"26. This automation trend could not only lead to societal
issues if more and more people are made redundant with LLM-powered systems, but also to quality
issues when humans are taken out of the loop and more and more users and companies are trusting
LLM-generated content without double-checking it2728.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>We showed that a downloadable open-source model (Mixtral 8x7B Instruct V0.1) was competitive with
some of the best available commercial models in a domain-specific biomedical RAG setting when used
with 10-shot learning. This opens up the possibility to have state-of-the-art performance in use cases
where using third-party APIs is not feasible because of the confidentiality of the data. The model used
via a commercial hosting service was also significantly faster than Claude 3 Opus while being at least
30x cheaper.</p>
      <p>We also observed that the zero-shot performance of this model was still lagging behind its commercial
competitors, making it even unusable in some settings where highly specific structured output is
required.</p>
      <p>We were unable to achieve consistent performance improvements from QLoRa fine-tuning Mixtral or
ifne-tuning gpt-3.5-turbo via the proprietary fine-tuning service of OpenAI. This might be an indication
that successfully fine-tuning these LLMs requires more engineering efort and costs and might not be
worth the efort in some use cases compared to few-shot learning.</p>
      <p>We tried to augment the context of these models with additional relevant knowledge from a knowledge
base (Wikipedia), but again could not see consistent performance improvements. We speculate that this
might be due to the knowledge in this setup being not novel enough for the models, or because of the
way we combined it with few-shot learning.</p>
      <p>For future work we want to verify our results with diferent, more domain-specific tasks where less
knowledge might have been present during the pre-training of these LLMs, and we also want to further
explore optimal selection of few-shot examples, as this seems to get the best performance out of these
models while being straightforward methodically.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>We want to thank the organizers of the BioASQ challenge for setting up this challenge and supporting us
during our participation. We are also grateful for the feedback and recommendations of the anonymous
reviewers.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Jiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ravaut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Joty</surname>
          </string-name>
          ,
          <article-title>Chatgpt's one-year anniversary: Are open-source large language models catching up</article-title>
          ?,
          <year>2024</year>
          . arXiv:
          <volume>2311</volume>
          .
          <fpage>16989</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piktus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Petroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Karpukhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Küttler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          , W.-t. Yih,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rocktäschel</surname>
          </string-name>
          , et al.,
          <article-title>Retrieval-augmented generation for knowledge-intensive nlp tasks</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>9459</fpage>
          -
          <lpage>9474</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Balaguer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Benara</surname>
          </string-name>
          , R. L.
          <string-name>
            <surname>de Freitas</surname>
            <given-names>Cunha</given-names>
          </string-name>
          ,
          <string-name>
            <surname>R. de M. Estevão Filho</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Hendry</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Holstein</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Marsman</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Mecklenburg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Malvar</surname>
            ,
            <given-names>L. O.</given-names>
          </string-name>
          <string-name>
            <surname>Nunes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Padilha</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sharp</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Aski</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Chandra</surname>
          </string-name>
          ,
          <article-title>RAG vs Fine-tuning:</article-title>
          <string-name>
            <surname>Pipelines</surname>
            ,
            <given-names>Tradeofs,</given-names>
          </string-name>
          <source>and a Case Study on Agriculture</source>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2401</volume>
          .
          <fpage>08406</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Katsimpras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lima-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré-Maduell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Loukachevitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Davydova</surname>
          </string-name>
          , E. Tutubalina, G. Paliouras,
          <source>Overview of BioASQ</source>
          <year>2024</year>
          :
          <article-title>The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering</article-title>
          , in: L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Quénot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Schwab</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Soulier</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Maria Di Nunzio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Galuščáková</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fifteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ),
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Katsimpras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Paliouras, Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . García Seco de Herrera (Eds.),
          <source>Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lima-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré-Maduell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rodríguez-Miret</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rodríguez-Ortega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lilli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lenkowicz</surname>
          </string-name>
          , G. Ceroni,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kossof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nentidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krithara</surname>
          </string-name>
          , G. Katsimpras, G. Paliouras,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Krallinger, Overview of MultiCardioNER task at BioASQ 2024 on Medical Speciality and Language Adaptation of Clinical NER Systems for Spanish, English and Italian</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . García Seco de Herrera (Eds.),
          <source>Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Davydova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Loukachevitch</surname>
          </string-name>
          , E. Tutubalina,
          <source>Overview of BioNNE Task on Biomedical Nested Named Entity Recognition at BioASQ</source>
          <year>2024</year>
          , in: CLEF Working Notes,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          , Attention is All You Need,
          <source>in: Proceedings of the 31st International Conference on Neural Information Processing Systems</source>
          , NIPS'17, Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <year>2017</year>
          , p.
          <fpage>6000</fpage>
          -
          <lpage>6010</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Narasimhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salimans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , et al.,
          <article-title>Improving language understanding by generative pre-training, preprint</article-title>
          ,
          <year>2018</year>
          . URL: https://web.archive.org/web/20240522131718/https: //cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>