<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>VICOMTECH at PROFE 2025: LLM Size is not so Important</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexander Platas</string-name>
          <email>aplatas@vicomtech.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aingeru Bellido</string-name>
          <email>abellido@vicomtech.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristian Parra</string-name>
          <email>cdparra@vicomtech.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Zotova</string-name>
          <email>ezotova@vicomtech.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo Turón</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Montse Cuadros</string-name>
          <email>mcuadros@vicomtech.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fundación Vicomtech, Basque Research and Technology Alliance (BRTA)</institution>
          ,
          <addr-line>Mikeletegi 57, 20009 Donostia-San Sebastián</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>This paper presents Vicomtech's participation in the PROFE 2025 shared task. Our team took part in all three proposed tasks, achieving the best results across the board by leveraging various Large Language Models (LLMs). The main strategies explored include experimenting with LLMs of diferent sizes, ensemble methods, and sequenceto-sequence approaches. The results demonstrate that while larger LLMs perform exceptionally well across all tasks, smaller models ofer competitive performance with significantly lower computational cost, making them a more lightweight and afordable alternative.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;LLM-prompting</kwd>
        <kwd>educational NLP</kwd>
        <kwd>gap filling</kwd>
        <kwd>multiple-choice question answering</kwd>
        <kwd>NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>experimental results. Finally, Sections 6 and 7 provide a discussion of key findings and outline future
directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Reading comprehension tests involve reading a short text passage and answering a series of questions
about that text. Automatic evaluation of these tests remains a challenging task. The vast majority of
studies are conducted to evaluate English exams. For English, there are diverse Multiple-Choice QA
datasets, such as RACE [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and QuAIL [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>For Spanish, the following question-answering datasets are available: SQuAD-es [5], a span-based
dataset with explicitly stated answers, and Entrance Exams (EE) [6], a multiple-choice dataset requiring
reasoning but limited by its small size. UNED-ACCESS 2024 [7] is a bilingual dataset which contains
1,003 questions from various subjects in the UNED Access Course for Over-25s, originally formulated in
Spanish and professionally translated into English. ReCoRES dataset [8] extracted from actual university
entrance examinations provided by Peruvian institutions that train students for entrance examinations,
comprises 439 texts and 1,822 questions with 2-7 candidate answers each.</p>
      <p>The methods in the recent studies are mostly based on the transformer architectures [9]. For instance,
mT5-based models have been employed in a pipeline encompassing candidate answer extraction,
answeraware question generation, and distractor generation [10]. [11] have investigated the extent to which
multilingual models can be trained in one language and applied to another for MCQ tasks. Findings
indicate that both monolingual and multilingual models can be zero-shot transferred to diferent datasets
and languages, maintaining performance levels. This approach is beneficial for languages with limited
annotated data</p>
      <p>The gap-filling task datasets are the following: the Cambridge Exams Publishing Open Cloze (CEPOC)
[12], SCDE: Sentence Cloze Dataset with High Quality Distractors From Examinations [13]. The research
on gap-filling tasks is mostly represented by a generation of gaps and distractors, and transformer-based
methods such as encoder-decoder models [14, 15, 16, 17].</p>
    </sec>
    <sec id="sec-3">
      <title>3. Task Description and Datasets</title>
      <p>The shared task focuses on the automatic resolution of oficial Spanish language exams designed by the
Instituto Cervantes, targeting learners’ levels from A1 to C2. The goal is to develop systems based on
language models capable of accurately solving diferent types of exam exercises.</p>
      <p>The task is divided into three subtasks, each corresponding to a specific exercise type commonly
found in these exams:
• Subtask 1, Multiple-Choice: Select the correct answer among some options based on a given text.
• Subtask 2, Matching: Match textual fragments from two lists.
• Subtask 3, Gap Filling: Complete gaps from a text by correctly filling them with the missing
fragments.</p>
      <p>These subtasks aim to evaluate the capabilities of NLP models in understanding and processing
language across a wide range of proficiency levels.</p>
      <p>To evaluate our proposed approaches, we constructed a benchmark dataset by collecting multiple
exercises for each subtask from publicly available online sources. These sources include oficial exams
published by Instituto Cervantes [18], as well as freely accessible educational websites that provide
similar Spanish language learning resources.</p>
      <p>It is important to emphasize that the collected data was used strictly for evaluation purposes. Our
goal was to obtain a reliable estimate of the performance of our models under realistic testing conditions.
To ensure the evaluation was representative, we curated exercises for three subtasks and across a broad
range of Spanish proficiency levels, from beginner (A1) to proficient (C2).</p>
      <sec id="sec-3-1">
        <title>3.1. Multiple-choice</title>
        <p>This subtask involves answering multiple-choice questions based on a reading passage in Spanish. The
texts cover general-domain content, ranging from activity schedules at lower proficiency levels to
narratives and articles at more advanced levels.</p>
        <p>The number of questions per passage and the number of answer options per question may vary
depending on the exam level. The overall goal is to evaluate diferent language models on this task,
which requires a dataset with matching characteristics.</p>
        <p>We initially used the Cambridge Multiple-Choice Questions Reading Dataset [19] for evaluation.
This corpus contains a total of 120 reading comprehension texts in English, covering a wide range of
proficiency levels. However, given that the model’s performance can vary significantly across languages,
we collected a new dataset consisting of 40 reading comprehension texts in Spanish, spanning various
proficiency levels, as detailed in Table 1.</p>
        <p>These instances were semi-automatically gathered from preparatory exercises sourced from Instituto
Cervantes website1, Lingua website2, Lingolia website3, Inmsol website4 and ProfeDeELE website5.
Each text has multiple-choice questions. Table 2 provides statistics on the number of questions per text
and the number of possible answers per question.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Matching</title>
        <p>The matching subtask is a challenge that involves identifying precise correspondences between two lists
of text fragments. Each source item must be accurately paired with its target counterpart. While three
illustrative examples are initially provided, the development of NLP based solutions requires access to a
larger well structured dataset. We collected an evaluation dataset gathered from various online sources
and aligned with the specifications required by the challenge. This dataset includes exercises spanning
diferent proficiency levels, as shown in Table 3.
1https://examenes.cervantes.es/
2https://lingua.com/es/espanol/lectura/
3https://espanol.lingolia.com/es/comprension-lectora
4https://www.inmsol.com/
5https://www.profedeele.es/examenes/</p>
        <p>The collected matching exercises consist of two parallel lists of text fragments and a general instruction
that outlines the task. The number of items in the source and target sets may vary, and some exercises
include distractor texts in the source set that do not correspond to any item in the target set. On average,
6 textual fragments in the source set and 7 items in the target set.</p>
        <p>The dataset compiled includes several sources, including the Cervantes Virtual Center6, Obejetivo
DELE (Diploma de Español como Lengua Extranjera)7, the Language Institute of the University of
Seville8, and the Tía Tula Blog9. Part of the data was compiled from oficial exams used for international
certification of Spanish proficiency within the DELE system. The other part contains exercises using an
exam model tailored to the matching task. Two common examples of matching exercises are presented
on our repository10.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Filling the gaps</title>
        <p>The “filling the gaps” task involves a text where certain fragments have been removed, with candidates
for these gaps presented in a disorderly fashion. Typically, the number of candidate fragments exceeds
the number of gaps to be filled 11.</p>
        <p>This type of exercise demands a high level of language comprehension and, as such, is typically
included only in exams at the B1 proficiency level or higher. Students are generally required to engage
in iterative reasoning over the available options before selecting the final answer.</p>
        <p>To compare diferent systems designed for this task, a dataset of 20 instances was compiled,
representing varied proficiency levels as detailed in Table 4. In the collected dataset, each exercise contains 7
to 8 gaps to be filled, and the number of fragments from which to choose is fixed at 6. Therefore, there
are typically 1 to 2 additional fragments included as distractors.</p>
        <p>These instances were manually gathered from previous examinations and preparatory exercises
sourced from the Instituto Cervantes website12, Tía Tula Spanish School website9 and DELE Ahora
Spanish learning website13. These sources were chosen due to their oficial alignment with the DELE
exam format and their wide use in preparation contexts.</p>
        <p>As this exercise format is specific to these exams, the number of collected instances is limited. As a
result, fine-tuning an LLM with such a limited number of instances is challenging; however, the dataset
has proven useful for selecting the most suitable ICL strategy (presented in Section 4.3).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Proposed Approach</title>
      <p>In this section, the proposed approaches are presented. Table 5 lists the LLMs used for each task,
including both open-source and commercial models of various sizes. Table 6 presents the embedding
models, and Table 7 shows the STS models.
6https://cvc.cervantes.es/
7https://objetivodele.com/
8https://institutodeidiomas.us.es/
9https://blog.tiatula.com/2010/03/modelos-de-examen-dele.html
10https://github.com/Vicomtech/profe2025/tree/master/subtask2
11https://github.com/Vicomtech/profe2025/tree/master/subtask3
12https://examenes.cervantes.es/
13https://deleahora.com/</p>
      <sec id="sec-4-1">
        <title>4.1. Multiple-choice</title>
        <p>For multiple-choice task, we explored several approaches ranging from zero-shot In-Context Learning
(ICL) with both commercial and open-source LLMs to more traditional language models for STS.
Additionally, we conducted fine-tuning experiments with LLMs and implemented ensemble methods
using LLMs and encoder models.</p>
        <sec id="sec-4-1-1">
          <title>4.1.1. Zero-shot ICL using LLMs</title>
          <p>For the zero-shot experiments with LLMs, we selected models demonstrating the highest performance
according to the current state-of-the-art. Table 5 presents a list of the models evaluated.</p>
          <p>We also explored the use of few-shot prompting; however, this approach did not yield significant
improvements. This is likely due to the constraint that the model must return only a single letter
corresponding to the correct answer, as we issued one prompt per question rather than prompting the
model to answer all questions in a given exercise at once.</p>
          <p>Furthermore, considering that smaller models (ranging from 8B to 14B) did not perform significantly
worse than larger ones, and that their errors occurred on diferent questions, we implemented an
ensemble of LLMs using Gemma 3 (12B), Phi 4 (14B), Qwen 2.5 (14B), and Ministral (8B), aiming to
outperform larger models such as DeepSeek R1 (685B). Two ensemble strategies were employed:
1. Majority voting: selecting the most frequent answer among the models. In case of a tie, priority
was given to the model with the best zero-shot performance.
2. Random Forest classifier : training a Random Forest on the predictions of the four models using
the training set. The classifier learns to weight each model’s output and selects the most likely
correct answer based on learned patterns.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Finetuning LLM</title>
          <p>
            After evaluating the LLMs in the zero-shot setting, we selected the model with the best size-performance
trade-of for fine-tuning. We also conducted an exhaustive search for multiple-choice QA datasets in
Spanish with contextual information and identified only two:
1. Belebele [
            <xref ref-type="bibr" rid="ref9">44</xref>
            ]: A human-annotated multiple-choice reading comprehension dataset spanning
122 language variants. In this case, we used the Spanish subset, which contains 900 questions.
          </p>
          <p>
            Each question has four multiple-choice answers and is linked to a short passage.
2. RetrievalQA [
            <xref ref-type="bibr" rid="ref10">45</xref>
            ]: An automatically generated dataset that contains 196 document-question
pairs, where each document is a short text about the history, culture, or other information of a
country or region.
          </p>
          <p>
            Despite the limited data, we performed an initial fine-tuning using LoRA 14. Due to the scarcity of
datasets with these specific characteristics, we translated a portion of the ReAding Comprehension
dataset from Examinations (RACE) [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] dataset. This dataset is a machine reading comprehension dataset
consisting of 27,933 passages and 97,867 questions from English exams. It is divided into middle and
high school level questions. We selected the 2,500 questions with the longest contexts from the high
school subset, since model performance showed more errors on C1–C2 level exams.
          </p>
          <p>
            We performed machine translation using two diferent models:
1. Itzuli: A Neural Machine Translation (NMT) system accessible via API upon request15. The
NMT approach has demonstrated its robustness for Basque-Spanish translation [
            <xref ref-type="bibr" rid="ref11">46</xref>
            ], but the Itzuli
system supports English-Spanish translation as well.
2. Gemma3 (27B) [32]: A family of lightweight, state-of-the-art open LLM from Google, built from
the same research and technology used to create the Gemini models.
          </p>
          <p>We used Itzuli for the reading passages and Gemma3 for the questions and answer choices. This
decision was based on the observation that specialized translation models such as Itzuli are more
accurate for full-sentence translation from English to Spanish, but often introduce gender, number and
14https://github.com/Vicomtech/profe2025/tree/master/subtask1#fine-tuning-hyperparameters
15https://itzuli.vicomtech.org/api/
verb conjugation errors when translating incomplete sentences. In many cases, the question consists of
an incomplete sentence, with the missing part provided among the answer choices. It is important to
emphasize the necessity for the answers to be grammatically compatible with the question, as otherwise
the model may disregard the correct option if it lacks linguistic coherence. An illustrative example is
shown in Table 8.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.1.3. Semantic Similarity</title>
          <p>This method expects that pretrained STS models are used to measure the semantic closeness between
pairs of text and answer options. These models generate embeddings for both the input text and each
candidate option. The semantic similarity between each pair is then computed using cosine similarity.</p>
          <p>
            In order to make the result more robust, we calculate the ensemble score with four models, using
majority voting and giving the x2 coeficient to the best model. We use the following models:
• deberta-base-long-nli [
            <xref ref-type="bibr" rid="ref6">41</xref>
            ] context length of 1280 trained for many tasks, including
linguisticsoriented natural language inference (NLI) and zero-shot entailment-based classification tasks.
• deberta-v3-base-tasksource-nli [
            <xref ref-type="bibr" rid="ref6">41</xref>
            ] fine-tuned with multi-task learning on 600+ tasks of the
tasksource collection. Performed as the best model in the separate evaluation.
• A2T_RoBERTa_SMFA_ACE [
            <xref ref-type="bibr" rid="ref7">42</xref>
            ] fine-tuned on multiliungal NLI datasets.
          </p>
          <p>
            • longformer-base-4096-bne-es-nli [
            <xref ref-type="bibr" rid="ref8">43</xref>
            ] fine-tuned on NLI-ES dataset 16.
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Matching</title>
        <p>To address the matching subtask, an experimental strategy was designed that compares diferent
language models based on zero-shot ICL and embedding representations models. The objective is to
evaluate both approaches in terms of how they conceptualize and perform the matching task. LLMs
ofer dynamic, contextual reasoning, but with a higher computational cost, while embedding-based
models allow for faster, similarity-based matching, but with more limited flexibility.
16https://huggingface.co/datasets/somosnlp-hackathon-2022/nli-es</p>
        <sec id="sec-4-2-1">
          <title>4.2.1. Zero-shot ICL using LLMs</title>
          <p>This approach relies on direct reasoning with LLM models, without the need for specific fine-tuning.
The system interprets the instructions for each exercise and generates the answer in a single step,
applying zero-shot ICL. We implemented the models presented in Table 5. In addition an ensemble
version of gemma-3-12b-it + Qwen2.5-14B-Instruct-1M + phi-4 14B was implemented.</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.2.2. Embedding-Based Approaches</title>
          <p>This line of experimentation relies on generating semantic vector representations of textual fragments
using embedding models. The main objective is to measure the semantic similarity between source and
target text, through cosine similarity, in order to identify the most likely matches. Both single-model
and ensemble configurations were explored to enhance matching accuracy. Notably, ensemble strategies
combined multiple embedding models to obtain more robust similarity estimates, often through weighted
voting mechanisms. The implemented models are presented in Table 6, and an ensemble model was
implemented (text-embedding-ada-002 + BAAI/bge-m3 + Open AI text-embedding-3-small).</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Filling the gaps</title>
        <p>Diferent systems were designed to adress the "filling the gaps" task which included ICL in Section
4.3.1, modern advanced strategies like Retrieval Augmented Generation (RAG) in Section 4.3.2 and
Agentic RAG in Section 4.3.3 as well as classical strategies such as ensemble models in Section 4.3.4 and
semantic search strategies in Section 4.3.5. For all of these systems the evaluation dataset presented in
Table 4 was used in order to compare results. The results are detailed in Figure 5.</p>
        <sec id="sec-4-3-1">
          <title>4.3.1. Zero-shot ICL using LLMs</title>
          <p>The classical ICL approach was initially employed to address the exercises. In this strategy, the task is
described to the model through the system prompt, while the user prompt17 contains the text along
with the fragment options. The selected state-of-the-art model generates a response in the correct JSON
format without requiring any prior examples of the task, following a zero-shot strategy.</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>4.3.2. Zero-shot ICL and RAG</title>
          <p>To investigate the efectiveness of modern approaches such as RAG, web search capabilities were
integrated into the system using the Google Search API18. The query to the search engine is created
from the first line of the text, previous to the first line break which usually corresponds to the title of
the text. The API returns a list of links ranked by relevance to the query from which the system extracts
the first four thousand characters of content. This retrieved information is then used to construct the
user prompt19 providing the model with contextually relevant content that may enhance the accuracy
of its response.</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>4.3.3. Zero-shot ICL and Agentic RAG</title>
          <p>A more advanced RAG approach was designed to get more relevant context from the Search API. In
this methodology, an Agentic RAG pipeline was designed to let the LLMs decide whether the contexts
retrieved are relevant to complete the exercise or not.</p>
          <p>In detail, once the Search API returns a list of links, an LLM is responsible for deciding if the links
provided are relevant20. The retrieved information is included in the context if relevant and ruled
out if not. This process is repeated in a loop until the context is formed of 3 information sources or
17https://github.com/Vicomtech/profe2025/tree/master/subtask3#zero-shot-icl-using-llms
18https://developers.google.com/custom-search
19https://github.com/Vicomtech/profe2025/tree/master/subtask3/#zero-shot-icl-and-rag
20https://github.com/Vicomtech/profe2025/tree/master/subtask3/#zero-shot-icl-and-agentic-rag
until another LLM decides that the current context is already enough to get an accurate response, as
illustrated in Figure 1.</p>
        </sec>
        <sec id="sec-4-3-4">
          <title>4.3.4. Ensemble models</title>
          <p>The results concluded that proprietary models were slightly better than open-source small ones. Aiming
to enhance results with open-source models, an ensemble model strategy was implemented, formed
of Qwen3 (32B), Phi-4 (14B) and DeepSeek R1 Distill Qwen (14B). These models were selected due to
their open-source availability, compatibility with our hardware constraints and for their reasoning
capabilities for some of them. The task was completed by the three models using the zero-shot ICL
strategy presented in Section 4.3.1 and performing a majority voting from the responses available for
each gap.</p>
        </sec>
        <sec id="sec-4-3-5">
          <title>4.3.5. Semantic matching</title>
          <p>A more classic approach was also designed without the use of any modern LLM. We used embedding
models to compare the text around the gaps with each of the available fragments using the embedding
model paraphrase-multilingual-MiniLM-L12-v2. The gaps were then assigned the most semantically
similar fragments. In addition, a more advanced algorithm was designed in order not to assign the same
fragment more than once to diferent gaps.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experimentation and Results</title>
      <p>This section describes the experiments conducted for each subtask and the corresponding results.</p>
      <sec id="sec-5-1">
        <title>5.1. Results based on our dataset</title>
        <p>Here we present the results obtained using our collected dataset, as described in Section 3. These
preliminary results guided our decision-making process regarding the selection of experiments, choice
of models, and the five final runs submitted to the shared task.</p>
        <sec id="sec-5-1-1">
          <title>5.1.1. Multiple-choice</title>
          <p>This section presents the results obtained at each stage of the experimentation, using the collected test
set described in Table 1. Figure 2 shows the performance of all LLMs in the zero-shot setting.</p>
          <p>Although the largest models achieved the highest scores, the accuracy gain relative to model size
is minimal. Models in the 12B–14B range scored only 2–3 points below much larger models such as
DeepSeek R1, which has 685B parameters. In general, all models achieved remarkably high performance,
with accuracy scores above 85%, except for the smaller models with 3-4 billion parameters, which
manage to exceed 80% despite their reduced size.</p>
          <p>Additionally, we observed a clear performance diference based on the release date of the LLMs.
The most recently released models, within the past few months, show a considerable improvement in
accuracy compared to earlier versions of similar size.</p>
          <p>To construct the ensemble, we selected the best-performing medium-sized and small models. As
described in Section 4.1.1, two strategies were used to determine the final answer: majority voting and a
Random Forest classifier. As shown in Figure 2, both approaches achieved strong performance, with the
majority voting ensemble standing out—it outperformed all individual models included in the ensemble.</p>
          <p>To select the model for fine-tuning, we analyzed the trade-of between model size and accuracy.
In Figure 3, we filtered models with 14B parameters or fewer, identifying three models with strong
performance relative to their size: Qwen 2.5 (14B), Gemma 3 (12B), and Mistral (8B).</p>
          <p>After analyzing the relationship between model size and performance, we observed the following: on
one hand, despite its smaller size, the Mistral model performs only 2–3 points below Gemma and Qwen.
On the other hand, although Qwen achieves the best results, it outperforms Gemma by just one point,
despite being 2B parameters larger. Considering the fine-tuning cost and potential for improvement,
we consider Gemma 3 to be a more viable option.</p>
          <p>As described in Section 4.1.2, two fine-tuning stages were conducted: an initial stage using a small
dataset (v1), followed by a second stage incorporating translated data (v2). Table 9 reports the results
obtained from the fine-tuning experiments.</p>
          <p>As shown in Table 9, in neither case did fine-tuning improve upon the performance of the base model.</p>
          <p>This may be attributed to the already strong zero-shot capabilities of the model, which achieved a very
high score, as well as the possibility that the training data is not suficiently representative of the test set,
either due to translation errors. This is particularly critical given that the task aims to assess language
proficiency in Spanish, and any deviation or inaccuracy in the linguistic input may propagate errors
and hinder model performance.</p>
          <p>It is also possible that the dificulty level of the training examples could be lower than that of the
evaluation set. In fact, the only improvements were observed in the performance of the lower-level
exams (A1 and A2).</p>
          <p>Regarding the STS models, Table 10 shows the zero-shot results obtained with each of them. As can
be seen, DeBERTa v3 clearly outperforms the other models. Nevertheless, its performance is far below
that of the LLMs, as even the worst-performing LLM (Phi-4 3B) achieves better metrics than the best
STS model.</p>
          <p>As with the LLMs, an ensemble using all the STS models was also implemented to surpass their
individual performance. As shown in Table 10, the ensemble yields the same results as DeBERTa v3.
This suggests that DeBERTa v3 already captures most of the relevant semantic information encoded by
the other models.</p>
          <p>Considering all the results obtained, we decided to submit diferent types of solutions to the shared
task. First, we observed that LLMs achieve remarkably high zero-shot performance on this task.
Therefore, we selected the best-performing model (DeepSeek R1) along with a medium-sized LLM that
also performed well (Qwen2.5 14B). This combination aimed to balance performance and computational
eficiency across diferent use cases.</p>
          <p>Regarding fine-tuning, since the base model outperformed the fine-tuned versions, we submitted
Gemma3 in its original (non-fine-tuned) form. On the other hand, due to the strong results obtained
with LLM ensembling via majority voting, we also submitted this approach. Finally, we chose to include
a more traditional solution based on an ensemble of STS models.</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>5.1.2. Matching</title>
          <p>The evaluation metric used was accuracy, measured at each CEFR level and globally. Figure 4 shows
results corresponding to LLMs and embedding models in matching task. Among the evaluated models,
DeepSeek R1 achieved the highest global accuracy (98.23%), closely followed by Gemini 2.0 Flash
(98.11%) and Claude 3.7 Sonnet (94.25%). These results suggest that advanced LLMs are highly capable
of performing the matching task in a zero-shot setting, even in the absence of domain-specific
finetuning. On the other hand, the evaluation of mid-sized models Gemma 3, Phi 4, and Qwen 2.5, reveals
a consistent yet slightly lower performance range compared to state-of-the-art large-scale models.
Despite their smaller parameter count and reduced inference cost, these models demonstrate a solid
capacity for text matching across multiple CEFR levels.</p>
          <p>While medium-sized models ofer a reasonable balance between computational cost and performance,
the best results in this task are clearly achieved with larger, higher-capacity LLMs. For high-impact or
high-level educational applications, such as automated exam grading, state-of-the-art LLMs such as
DeepSeek R1 and Gemini 2.0 Flash are the most reliable options. However, for applications with limited
computational resources, Qwen 2.5 and the ensemble strategy represent viable alternatives, especially
if refined with task-specific tuning or hybrid matching logic.</p>
          <p>In contrast, embedding-based approaches ofer an alternative paradigm, relying on semantic similarity
metrics rather than direct reasoning. The most efective model was Ada (text-embedding-ada-002),
which achieved an overall accuracy of 68.14%, excelling at basic levels but failing at advanced levels.
Models such as text-embedding-3-small and baai-m3 yielded comparable results (with global accuracy
scores of 63.27%, and 58.41%, respectively), but exhibited limitations when faced with complex semantic
relationships or distracting fragments. Overall, although eficient, embedding models lack the semantic
depth necessary for accurate reading comprehension in contexts of greater linguistic complexity.</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>5.1.3. Filling the gaps</title>
          <p>For the approaches detailed in Section 4.3, we employed several of the LLMs listed in Table 5. The
corresponding results, evaluated using accuracy as the metric, are presented in Figure 5. Among the
evaluated methods, the traditional semantic similarity-based approach yielded significantly lower results
(30.05%). In contrast, all other approaches, which leveraged generative LLMs, achieved substantially
higher accuracy scores, clearly outperforming the baseline.</p>
          <p>When comparing reasoning-enabled LLMs to those without explicit reasoning capabilities, the former
demonstrated greater suitability for this task. Analyzing their reasoning outputs it reveals that these
models consider multiple possible predictions and, throughout the reasoning process, less plausible
options are ruled out. This behaviour closely mirrors the way a human would approach the task through
an iterative process that ultimately converges on the most appropriate answer.</p>
          <p>The use of a RAG approach also outperformed the results obtained using standalone LLMs. This
suggests that LLMs benefit from incorporating external information retrieved from the web, enabling them
to generate more accurate responses when provided with relevant context. However, the performance
of the agentic RAG approach was not consistently reliable, likely due to the presence of irrelevant or
unhelpful information in the retrieved context, which may have introduced noise and hindered task
resolution.</p>
          <p>The ensemble model did not yield particularly noteworthy results, as its performance did not surpass
that of the best individual model (Qwen3-32B) included in the ensemble.</p>
          <p>On the other hand, the diferent approaches and models did not exhibit significant variations in
performance across the various exam levels.</p>
          <p>In conclusion, for the submission of five selected approaches, we included the best non-RAG method
(Gemini 2.5 Pro), the best-performing open-source model (DeepSeek R1), and the model that exhibited
the largest performance variations across tasks (Gemini 2.0 Flash Thinking).</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Results on the task dataset</title>
        <p>The evaluation follows the oficial shared task protocol, which uses accuracy (proportion of correct
answers) as the primary performance metric. Evaluation scores are reported from two complementary
perspectives:
• Question-level accuracy (Acc.), where each question is evaluated independently, and the final
score is the proportion of correctly answered questions.
• Exam level, where each exam is composed of multiple exercises spanning diferent task types. An
exam is considered successfully passed if it achieves an accuracy score above 0.6. The overall
exam-level score corresponds to the proportion of passed exams across the dataset.</p>
        <p>The test set used comprises multiple exams, each consisting of several exercises. These exercises
are categorized according to the corresponding subtask. Table 11 shows the number of exercises per
subtask and proficiency level.</p>
        <p>Table 12 summarizes the performance of systems across the three subtasks as well as the overall
exam-level accuracy. Notably, systems incorporating LLMs achieve substantially higher accuracy
across all subtasks. For instance, DeepSeek R1, Gemini 2.0 Flash and Gemini 2.5 Pro achieves the best
performances in each task, respectively.</p>
        <p>At the exam level, the top-performing systems consistently surpass 98% accuracy. This demonstrates
that employing LLMs yields the most efective results for this task. Conversely, approaches based
on embeddings or text similarity techniques, while generally competitive, fall notably short of the
performance achieved by LLM-based models.</p>
        <p>The pronounced performance gap between baseline and advanced systems highlights the substantial
advantage of utilizing pretrained language models and ensemble strategies. Nevertheless, the variability
in accuracy observed across the subtasks reflects the distinct challenges and complexities inherent to
each task type, indicating areas that merit deeper exploration in future work.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>We present the results and resources of the PROFE shared task on solving a variety of exercises extracted
from oficial Spanish language exams. For each subtask, only a single example per exercise type was
available, limiting the viability of personalized solutions such as fine-tuning or any training-based
approaches. Furthermore, the highly specific nature of the tasks made it dificult to identify appropriate
open-source resources without relying on synthetic data generation or machine translation.</p>
      <p>The results from this shared task clearly illustrate the significant progress made by LLMs in recent
years. These models achieved remarkably high performance (often exceeding 85% accuracy) in zero-shot
settings, especially those with advanced reasoning capabilities. This ability to self-evaluate and refine
their outputs without requiring few-shot prompting makes them particularly appropriate for scenarios
with limited training data.</p>
      <p>Regarding the evaluation data, some issues were identified in the dataset, such as duplicated instances
or incomplete text fragments. Interestingly, these irregularities were often only efectively handled by
LLM-based approaches, which were capable of inferring the intended structure or context through their
generative reasoning. Table 13 presents an example of the reasoning process exhibited by DeepSeek R1
when confronted with a multiple-choice question in which the correct answer was located within a
missing portion of the input text.</p>
      <p>DeepSeek-R1 Reasoning (A1E_2017-11-24_E4)
Okay, let’s see. The question is asking "Hay una fiesta el _______". So I need to find which day has a
party mentioned in the schedule. Looking at the text, each day is listed with activities. Let me check
each day one by one. Starting with LUNES (Monday). The activities are a bus departure, walking tour,
lunch, free time, concert, dinner. No mention of a fiesta here. MARTES (Tuesday): Breakfast, boat
ride, lunch, free time, football match, dinner. Again, no fiesta. MIÉRCOLES (Wednesday): Breakfast,
museum visit, lunch, free time, cinema, dinner. Still no party. JUEVES (Thursday): Breakfast at the
pool cafeteria, mountain excursion, lunch. Wait, the Lunch here has two entries? 14:00–15:00 Comida:
Pescado con ensalada y Postre. Then another Comida: Sopa, carne y Fruta. That might be a typo, but
after that, free time, then merienda (snack) at 17:00–19:00. Then at 19:30, Fiesta de despedida. Oh,
that’s the farewell party. So Jueves (Thursday) has the fiesta. But wait, the options given are A)
Lunes, B) Viernes, C) Sábado. Wait, Jueves is Thursday, which is not listed. The options are Monday,
Friday, Saturday. Hmm. Wait, the days listed in the text are Lunes, Martes, Miércoles, Jueves, Sábado.
Wait, where’s Viernes (Friday)? The schedule skips from Jueves (Thursday) to Sábado (Saturday).
Therefore, I’ll choose B) viernes. But I’m not confident. Alternatively, maybe the answer is C) sábado.
But according to the text, it’s not. Therefore, I think the answer is just not present. But since we’re
supposed to fill it, I’ll choose B) viernes.</p>
      <p>Correct Answer: B) viernes</p>
      <p>Despite the lack of this fragment, the model was able to reconstruct plausible content based on the
adjacent context and use this inferred information to select the correct option. A solution based solely on
semantic similarity would be unable to resolve such cases correctly, as the target word does not appear
anywhere in the text and must instead be intuited through higher-level reasoning. This highlights the
LLM’s capacity for contextual extrapolation and implicit reasoning, which proves especially valuable in
scenarios where input data is incomplete, noisy, or partially corrupted.</p>
      <p>Moreover, given that this is a general-domain task where the only requirement is a strong command
and understanding of the Spanish language, LLMs—trained on vast multilingual corpora—demonstrate
a clear advantage over traditional approaches such as semantic similarity methods. Their deep
understanding of linguistic nuances, combined with robust reasoning and contextual interpretation, positions
them as far superior in these types of exercises without the need for domain-specific tuning.</p>
      <p>Additionally, analysis of the results revealed no consistent performance drop at higher proficiency
levels of the exam. Instead, the accuracy metrics varied regularly across models, regardless of dificulty.
This suggests that LLMs already have a high level of competence in Spanish, and that remaining errors
are more likely attributable to factors such as complexity, ambiguity or longer contextual dependencies
rather than lack of linguistic understanding.</p>
      <p>These findings suggest that LLMs, even without task-specific adaptation, are capable of handling
complex linguistic tasks with surprising robustness.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions and Future work</title>
      <p>The dificulty of the Spanish exams has been reflected in the diferent subtasks. Nevertheless, our
evaluation confirms that state-of-the-art LLMs are now capable of solving these tasks with accuracy
scores approaching the upper bound using ICL techniques. In contrast, traditional approaches lag
significantly behind and fail to achieve comparable performance. These tasks are designed to require
students to reason over the available options, a demand that is mirrored in the superior performance of
LLMs with reasoning capabilities.</p>
      <p>On the other hand, current LLMs have demonstrated the ability to handle these complex tasks
in Spanish efectively. Notably, this capability is not limited to large proprietary models; smaller,
open-source LLMs can also attain comparable results on this type of task.</p>
      <p>Although fine-tuning has traditionally been an efective strategy for adapting models to specific tasks,
in this case, it has proven inefective. As such, future work may focus on enhancing ICL methods—such
as prompt engineering—and refining the application of modern techniques for LLM-based reasoning.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This work has been supported by the project MASTERMIND ZL-2025/00267 funded by the government
of the Basque Country, Spain.</p>
    </sec>
    <sec id="sec-9">
      <title>Declaration on Generative AI</title>
      <p>Generative AI (ChatGPT with GPT-4.5) was utilised to: Improve writing style, Abstract drafting,
Grammar and spelling check. After using these tool, the authors reviewed and edited the content as
needed and take full responsibility for the publication’s content.
Intelligence 34 (2020) 8722–8731. URL: https://ojs.aaai.org/index.php/AAAI/article/view/6398.
doi:10.1609/aaai.v34i05.6398.
[5] C. P. Carrino, M. R. Costa-jussà, J. A. R. Fonollosa, Automatic Spanish translation of SQuAD
dataset for multi-lingual question answering, in: N. Calzolari, F. Béchet, P. Blache, K. Choukri,
C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk,
S. Piperidis (Eds.), Proceedings of the Twelfth Language Resources and Evaluation Conference,
European Language Resources Association, Marseille, France, 2020, pp. 5515–5523. URL: https:
//aclanthology.org/2020.lrec-1.677/.
[6] A. Peñas, E. Hovy, P. Forner, Á. Rodrigo, R. Sutclife, R. Morante, Qa4mre 2011-2013: Overview of
question answering for machine reading evaluation, in: P. Forner, H. Müller, R. Paredes, P. Rosso,
B. Stein (Eds.), Information Access Evaluation. Multilinguality, Multimodality, and Visualization,
Springer Berlin Heidelberg, Berlin, Heidelberg, 2013, pp. 303–320.
[7] E. Sánchez Salido, R. Morante, J. Gonzalo, G. Marco, J. Carrillo-de Albornoz, L. Plaza, E. Amigo, A. F.</p>
      <p>García, A. Benito-Santos, A. Ghajari Espinosa, V. Fresno, Bilingual evaluation of language models
on general knowledge in university entrance exams with minimal contamination, in: O. Rambow,
L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schockaert (Eds.), Proceedings of the 31st
International Conference on Computational Linguistics, Association for Computational Linguistics,
Abu Dhabi, UAE, 2025, pp. 6184–6200. URL: https://aclanthology.org/2025.coling-main.413/.
[8] M. A. S. C. y Diego Diestra y Rodrigo López y Erasmo Gómez y Arturo Oncevay y Fernando
Alva-Manchego, Overview of recores at iberlef 2022: Reading comprehension and reasoning
explanation for spanish, Procesamiento del Lenguaje Natural 69 (2022) 281–287. URL: http:
//journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6448.
[9] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, I. Polosukhin,
Attention is all you need, in: I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S.
Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems, volume 30,
Curran Associates, Inc., 2017, pp. 1–11. URL: https://proceedings.neurips.cc/paper_files/paper/
2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
[10] D. de Fitero-Dominguez, A. Garcia-Cabot, E. Garcia-Lopez, Automated multiple-choice
question generation in spanish using neural language models, Neural Computing and
Applications 36 (2024) 18223–18235. URL: https://doi.org/10.1007/s00521-024-10076-7. doi:10.1007/
s00521-024-10076-7.
[11] G. E. y Álvaro Rodrigo y Anselmo Peñas, Cross-lingual training for multiple-choice question
answering, Procesamiento del Lenguaje Natural 65 (2020) 37–44. URL: http://journal.sepln.org/
sepln/ojs/ojs/index.php/pln/article/view/6274.
[12] M. Felice, S. Taslimipoor, Ø. E. Andersen, P. Buttery, Cepoc: The cambridge exams publishing open
cloze dataset, in: Proceedings of the Thirteenth Language Resources and Evaluation Conference,
2022, pp. 4285–4290.
[13] X. Kong, V. Gangal, E. Hovy, SCDE: Sentence cloze dataset with high quality distractors from
examinations, in: D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Proceedings of the 58th
Annual Meeting of the Association for Computational Linguistics, Association for Computational
Linguistics, Online, 2020, pp. 5668–5683. URL: https://aclanthology.org/2020.acl-main.502/. doi:10.
18653/v1/2020.acl-main.502.
[14] B. Moharana, V. K. Singh, T. Sarkar, D. Singh, M. Rakhra, V. K. Pandey, Automated questions
answering generation system adopting nlp and t5, in: 2024 International Conference on
Cybernation and Computation (CYBERCOM), 2024, pp. 363–369. doi:10.1109/CYBERCOM63683.2024.
10803238.
[15] S. K. Bitew, J. Deleu, A. S. Doğruöz, C. Develder, T. Demeester, Learning from partially annotated
data: Example-aware creation of gap-filling exercises for language learning, in: E. Kochmar,
J. Burstein, A. Horbach, R. Laarmann-Quante, N. Madnani, A. Tack, V. Yaneva, Z. Yuan, T. Zesch
(Eds.), Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational
Applications (BEA 2023), Association for Computational Linguistics, Toronto, Canada, 2023, pp.
598–609. URL: https://aclanthology.org/2023.bea-1.51/. doi:10.18653/v1/2023.bea-1.51.
[16] M. P. P. Jadhav, M. M. D. Laddha, An automatic gap filling questions generation using nlp, Ijcset.</p>
      <p>Com 8 (2017).
[17] C. Y. Yeung, J. S. Lee, B. K. Tsou, Dificulty-aware distractor generation for gap-fill items, in:
Proceedings of the 17th annual workshop of the Australasian language technology association,
2019, pp. 159–164.
[18] Instituto Cervantes, Página oficial del instituto cervantes, https://www.cervantes.es, 2025.
Consultado el 26 de mayo de 2025.
[19] A. Mullooly, O. Andersen, L. Benedetto, P. Buttery, A. Caines, M. J. F. Gales, Y. Karatay, K. Knill,
A. Liusie, V. Raina, S. Taslimipoor, The Cambridge Multiple-Choice Questions Reading Dataset,
Cambridge University Press and Assessment, 2023. URL: https://www.repository.cam.ac.uk/handle/
1810/358683. doi:10.17863/CAM.102185.
[20] Anthropic, Claude 3.5 sonnet model card addendum, 2025. URL: https://assets.anthropic.com/m/
785e231869ea8b3b/original/claude-3-7-sonnet-system-card.pdf.
[21] Google, Gemini 2.0 flash, 2024. URL: https://deepmind.google/technologies/gemini/flash/.
[22] Google, Gemini 2.0 flash exp. 01-21, 2025. URL: https://ai.google.dev/gemini-api/docs/models?hl=
es-419.
[23] Google, Gemini 2.5 flash preview, 2025. URL: https://storage.googleapis.com/model-cards/
documents/gemini-2.5-flash-preview.pdf.
[24] Google, Gemini 2.5 pro preview, 2025. URL: https://storage.googleapis.com/model-cards/
documents/gemini-2.5-pro-preview.pdf.
[25] OpenAI, Openai o3 and o4-mini system card, 2025. URL: https://cdn.openai.com/pdf/
2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf.
[26] OpenAI, Hello gpt-4o, 2024. URL: https://openai.com/index/hello-gpt-4o/.
[27] Q. Team, Qwq-32b: Embracing the power of reinforcement learning, 2025. URL: https://qwenlm.</p>
      <p>github.io/blog/qwq-32b/.
[28] Q. Team, Qwen2.5: A party of foundation models!, 2024. URL: https://qwenlm.github.io/blog/
qwen2.5/.
[29] Q. Team, Qwen3 technical report, 2025. URL: https://arxiv.org/abs/2505.09388.</p>
      <p>arXiv:2505.09388.
[30] DeepSeek-AI, Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,
2025. URL: https://arxiv.org/abs/2501.12948. arXiv:2501.12948.
[31] M. Llama Team, The llama 3 herd of models, 2024. URL: https://arxiv.org/abs/2407.21783.</p>
      <p>arXiv:2407.21783.
[32] G. Team, Gemma 3 technical report, 2025. URL: https://arxiv.org/abs/2503.19786.</p>
      <p>arXiv:2503.19786.
[33] M. A. team, Mistral large 2 (2407), 2024. URL: https://mistral.ai/news/mistral-large-2407.
[34] M. A. team, Mistral nemo, 2024. URL: https://mistral.ai/news/mistral-nemo.
[35] M. A. team, Ministral, 2024. URL: https://mistral.ai/news/ministraux.
[36] M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M.
Javaheripi, P. Kaufmann, J. R. Lee, Y. T. Lee, Y. Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price,
G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y. Wu, D. Yu, C. Zhang, Y. Zhang,
Phi-4 technical report, 2024. URL: https://arxiv.org/abs/2412.08905. arXiv:2412.08905.
[37] N. Reimers, I. Gurevych, Sentence-BERT: Sentence embeddings using Siamese BERT-networks, in:
K. Inui, J. Jiang, V. Ng, X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods
in Natural Language Processing and the 9th International Joint Conference on Natural Language
Processing (EMNLP-IJCNLP), Association for Computational Linguistics, Hong Kong, China, 2019,
pp. 3982–3992. URL: https://aclanthology.org/D19-1410/. doi:10.18653/v1/D19-1410.
[38] OpenAI, Text embedding ada 002, 2022. URL: https://openai.com/index/
new-and-improved-embedding-model/.
[39] J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, Z. Liu, M3-embedding: Multi-linguality,
multifunctionality, multi-granularity text embeddings through self-knowledge distillation, in: L.-W. Ku,
A. Martins, V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Á.</given-names>
            <surname>Rodrigo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Moreno-Álvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P.</given-names>
            <surname>García-Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Peñas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Agerri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fruns-Jiménez</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. SoriaPastor</surname>
          </string-name>
          , Overview of PROFE at IberLEF 2025:
          <article-title>Language Proficiency Evaluation</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>75</volume>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <string-name>
            <surname>González-Barba</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <article-title>Overview of IberLEF 2025: Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the 41st Conference of the Spanish Society for Natural Language Processing (SEPLN 2025), CEUR-WS</article-title>
          . org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Lai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xie</surname>
          </string-name>
          , H. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          , E. Hovy, RACE:
          <article-title>Large-scale ReAding comprehension dataset from examinations</article-title>
          , in: M.
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hwa</surname>
          </string-name>
          , S. Riedel (Eds.),
          <source>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Copenhagen, Denmark,
          <year>2017</year>
          , pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          . URL: https://aclanthology.org/D17-1082/. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D17</fpage>
          -1082.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rogers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kovaleva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Downey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rumshisky</surname>
          </string-name>
          ,
          <article-title>Getting closer to ai complete question answering: A set of prerequisite real tasks</article-title>
          ,
          <source>Proceedings of the AAAI Conference on Artificial</source>
          <year>2024</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Bangkok, Thailand,
          <year>2024</year>
          , pp.
          <fpage>2318</fpage>
          -
          <lpage>2335</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .findings-acl.
          <volume>137</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .findings-acl.
          <volume>137</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [40]
          <string-name>
            <surname>OpenAI</surname>
          </string-name>
          , Text embedding 3-small,
          <year>2024</year>
          . URL: https://openai.com/index/ new-embedding
          <article-title>-models-and-api-updates/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>D.</given-names>
            <surname>Sileo</surname>
          </string-name>
          ,
          <article-title>tasksource: A large collection of NLP tasks with a structured dataset preprocessing framework</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            , M.-
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Hoste</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lenci</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sakti</surname>
          </string-name>
          , N. Xue (Eds.),
          <source>Proceedings of the 2024 Joint International Conference on Computational Linguistics</source>
          ,
          <article-title>Language Resources and Evaluation (LREC-COLING 2024), ELRA</article-title>
          and
          <string-name>
            <given-names>ICCL</given-names>
            ,
            <surname>Torino</surname>
          </string-name>
          , Italia,
          <year>2024</year>
          , pp.
          <fpage>15655</fpage>
          -
          <lpage>15684</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .lrec-main.
          <volume>1361</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>O.</given-names>
            <surname>Sainz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lopez de Lacalle</surname>
          </string-name>
          , E. Agirre,
          <string-name>
            <surname>B. Min,</surname>
          </string-name>
          <article-title>ZS4IE: A toolkit for zero-shot information extraction with simple verbalizations</article-title>
          , in: H.
          <string-name>
            <surname>Hajishirzi</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Ning</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Sil (Eds.),
          <source>Proceedings of the</source>
          <year>2022</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: System Demonstrations, Association for Computational Linguistics</article-title>
          , Hybrid: Seattle, Washington + Online,
          <year>2022</year>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>38</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .naacl-demo.4/. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .naacl-demo.
          <volume>4</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [43]
          <string-name>
            <surname>Sleoruiz</surname>
          </string-name>
          , Huggingface model,
          <year>2022</year>
          . URL: https://huggingface.co/Sleoruiz/ longformer-base-4096
          <string-name>
            <surname>-</surname>
          </string-name>
          bne
          <article-title>-es-nli.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bandarkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Artetxe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. N.</given-names>
            <surname>Shukla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Husa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khabsa</surname>
          </string-name>
          ,
          <article-title>The belebele benchmark: a parallel reading comprehension dataset in 122 language variants</article-title>
          ,
          <source>in: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          ,
          <year>2024</year>
          , p.
          <fpage>749</fpage>
          -
          <lpage>775</lpage>
          . URL: http://dx.doi.org/10.18653/v1/
          <year>2024</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>44</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>44</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Retrievalqa: A benchmark dataset for retrieval-augmented question answering</article-title>
          , https: //huggingface.co/datasets/lnwang/retrieval_qa,
          <year>2023</year>
          . https://github.com/wln20/Retrieval_QA.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>T.</given-names>
            <surname>Etchegoyhen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Martínez</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Azpeitia</surname>
          </string-name>
          , G. Labaka,
          <string-name>
            <surname>I. Alegria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Cortes</given-names>
            <surname>Etxabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Jauregi</given-names>
            <surname>Carrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Ellakuria</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Calonge,</surname>
          </string-name>
          <article-title>Neural machine translation of Basque</article-title>
          , in: J. A.
          <string-name>
            <surname>Pérez-Ortiz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Sánchez-Martínez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Esplà-Gomis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Popović</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Rico</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Martins</surname>
          </string-name>
          , J. Van den Bogaert, M. L. Forcada (Eds.),
          <source>Proceedings of the 21st Annual Conference of the European Association for Machine Translation</source>
          , Alicante, Spain,
          <year>2018</year>
          , pp.
          <fpage>159</fpage>
          -
          <lpage>168</lpage>
          . URL: https://aclanthology.org/
          <year>2018</year>
          .eamt-main.
          <volume>14</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>