<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Team LLMinds Submission for ELOQUENT Sensemaking Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anna Sajdoková</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matěj Macek</string-name>
          <email>matej.macek@email.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ondřej Hlava</string-name>
          <email>ohlava@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marek Štefanec</string-name>
          <email>marek.stefanec@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonín Kříž</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jakub Kučera</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ondřej Bojar</string-name>
          <email>bojar@ufal.mf.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics</institution>
          ,
          <addr-line>Malostranské náměstí 25, 118 00 Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Czech Technical University in Prague, Faculty of Information Technology</institution>
          ,
          <addr-line>Thákurova 9, 160 00 Prague 6</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents our system submission for the CLEF 2025 ELOQUENT Sensemaking Task, which focuses on generating and answering questions based solely on provided textual materials, such as lecture transcripts and textbooks. Our approach combines open-source large language models (LLMs) in a two-stage pipeline: a Teacher component for question generation, a Student for grounded answering using Retrieval-Augmented Generation (RAG). Ideally, the third stage would be an Evaluator for an assessment of response quality but we leave this part for the future. To ensure transparency, privacy, and accessibility, we prioritized using local models such as LLaMA 7B and DistilQwen 1.5B, avoiding reliance on proprietary APIs. For question generation, we implemented both a simple single-question prompting method and an enhanced generator-discriminator pipeline that scores candidates on answerability, context relevance, and diversity. The Student model leverages a RAG system with FAISS to extract relevant context chunks and generate concise answers. We evaluated several configurations, including standalone model outputs and hybrid combinations (llama + distilqwen), with the final answers polished by OpenAI under strict token and editing constraints. Our findings demonstrate that lightweight local models, when properly orchestrated, can produce competitive results in question answering tasks while respecting data sovereignty and resource limitations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large Language Models</kwd>
        <kwd>Question Generation</kwd>
        <kwd>Question Answering</kwd>
        <kwd>Retrieval-Augmented Generation</kwd>
        <kwd>GeneratorDiscriminator Pipeline</kwd>
        <kwd>Open-Source NLP</kwd>
        <kwd>Educational NLP</kwd>
        <kwd>Local Inference</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The ELOQUENT Sensemaking task, part of the CLEF 2025 evaluation lab, challenges participants
to develop systems capable of generating and answering questions based solely on provided textual
materials, such as lecture transcripts or textbook excerpts. The primary objective is to assess the ability
of large language models (LLMs) to comprehend and reason over specific content provided as input and
without being overly influenced by the knowledge accumulated in the model from the training data.</p>
      <p>
        Šindelář and Bojar published the process of making this task as well as its evaluation in the article
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Our project, LLMinds, originally aimed to construct a pipeline comprising three components: a
Teacher that generates questions from the input materials, a Student that answers these questions
using only the given context, and an Evaluator that assesses the quality of the generated
questionanswer pairs. Eventually, we skipped the Evaluator component due to time constraints.</p>
      <p>A significant motivation for our approach is the desire to utilize open-source, locally deployable
models instead of proprietary solutions like OpenAI’s GPT series. Proprietary models mostly require
sending data to external servers, raising concerns about data privacy and compliance. This problem can
be even more pressing in educational settings where the school is obliged to protect the privacy of its
student’s data. Additionally, such models may demand substantial computational resources and incur
usage costs, making them less accessible for institutions with limited budgets.</p>
      <p>By leveraging local models, we aim to:
• Ensure data privacy by processing all information on-premises.
• Reduce dependency on external services and associated costs.</p>
      <p>• Enable deployment on hardware with limited computational capabilities.</p>
      <p>This approach aligns with the goals of the ELOQUENT lab to explore trustworthy, controllable, and
accessible LLM applications in real-world scenarios.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        The field of question generation and answering has seen significant advancements with the emergence
of large language models (LLMs). Transformer-based architectures such as BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and GPT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] have
demonstrated strong capabilities in processing and generating natural language. Retrieval-Augmented
Generation (RAG) [4] further enhances performance by combining dense retrieval mechanisms with
generative models, allowing systems to ground their outputs in external knowledge.
      </p>
      <p>In our Student setup, we build on RAG, optimizing both retrieval relevance and answer quality. In
the Teacher submission, we try generator-discriminator pipeline much like in generative adversarial
networks (GANs).</p>
    </sec>
    <sec id="sec-3">
      <title>3. The System</title>
      <sec id="sec-3-1">
        <title>3.1. Teacher Submission</title>
        <p>The Teacher component is tasked with creating questions from the provided content, be it lectures
or books, for the Student component to answer. We implemented the Teacher in two ways. First by
using small local models to generate only one single question and then using larger models available
through APIs to generate multiple questions at the same time with large context windows and focus on
uniqueness.</p>
        <p>Our second approach to Teacher task uses generator-discriminator pipeline further described in
Section 3.1.2.</p>
        <sec id="sec-3-1-1">
          <title>3.1.1. Small Local Models Approach</title>
          <p>We started with multiple smaller local models, ranging from Llama-3.2-1B and Llama-3.2-3B-Instruct to
larger “thinking” models like DeepSeek-R1-Distill-Qwen-14B-GGUF.</p>
          <p>To obtain questions, we prompted the LLM with the provided input document and a request to
generate a question. To battle the repetitiveness of the questions generated by the models, we set the
random seed of each run randomly. Since the provided input material can be quite large, we used only
the last 16384 characters of each document to fit them in the available VRAM.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.1.2. Discriminator and Generator Approach</title>
          <p>In our second approach to the Teacher task, we use an LLM with diferent seeds to generate 5 questions
per entry. Our question generation system includes a simple discriminator component that evaluates
and selects high-quality questions from a larger pool of generated candidates.</p>
          <p>We aim to ensure that the selected questions are diverse, relevant, and of appropriate dificulty. The
discriminator uses the following metrics:
1. Answerability By answerability, we aim to verify that a question can be completely and accurately
answered using only the information provided in the document excerpt. We estimate answerability
with the same large language model (LLM) used for generation, by applying the following prompt:
Prompt:
Can the question be answered completely and accurately using only the information
provided in the document excerpt?</p>
          <p>Answer with just Yes or No.</p>
          <p>We automatically filter out all questions where the LLM returns No.</p>
          <p>We optionally use a more advanced version of the prompt to improve the estimate: First, the LLM
is asked to select the most relevant passage from the context that supports the answer. Then, in a
second step, we prompt the LLM without any further context to decide whether that selected passage is
suficient to answer the question. This two-step process acts as a confidence check and helps avoid
mistakenly labeling unanswerable questions as valid.
2. Context Relevance Context relevance numerically estimates how much the question is related to
the provided context. We use lexical overlap between the question and the context:
• Common stop words are removed from the question.
• We compute what percentage of meaning-bearing words from the question (excluding
interrogative words like what, where, etc., unless contentful) also appear in the context.
• The score ranges from 0.0 to 1.0, where a higher score indicates that the question is more directly
grounded in the given context.
3. Diversity Score We assign a diversity score to the entire pool of questions generated for a single
input text. This score reflects how much each question difers from the others in both form and content.</p>
          <p>Technically, we estimate diversity by:
• Comparing each question with all others in the same pool.
• Calculating text similarity as the word overlap between questions.
• Calculating context similarity as the word overlap between the associated contexts.
• Computing the final diversity score as:</p>
          <p>Diversity = 1.0 −
︂( avg_text_sim + avg_context_sim )︂
2.0
• Higher scores (closer to 1.0) indicate more unique and diverse questions.</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Ranking and Selection Process</title>
          <p>process:
1. Sort all candidate questions by:
• Answerability: answerable questions first
• Diversity score: higher score first
• Context relevance: higher score first</p>
          <p>Overall, the final set of questions is selected using the following
2. Ensure question variety by organizing questions into a 3x3 matrix:
• Three dificulty levels: easy, medium, hard
• Three question types: factual, inferential, analytical
We label each question using an additional classification step with an LLM that assigns dificulty
and type labels based on custom prompts. We ensure that at least one question is selected from
each matrix cell.
3. Fill remaining slots with top-ranked questions not yet selected, ensuring the overall number of
questions per document stays within the desired limit.</p>
          <p>Handling Large Documents For long documents that exceed input limits, we apply a chunking
strategy:
• The document is split into overlapping chunks, typically at sentence or paragraph boundaries.</p>
          <p>When natural breaks are unavailable, we fall back on token-based limits with overlaps to ensure
context continuity.
• We generate a set of candidate questions for each chunk individually and collect all candidates
into a common pool.
• The full candidate pool from all chunks is combined and passed through the same ranking and
selection pipeline described above.</p>
          <p>This method ensures that all parts of a long document are represented, enlarges the initial question
pool, and ultimately helps maintain a fixed-size, high-quality final question set.</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>3.1.3. Future Improvements of Generator Discriminator Pipeline</title>
          <p>In the future, we would like to improve method at least in the following ways:
• Use semantic similarity measures (e.g., cosine similarity with Sentence-BERT embeddings) to
enhance diversity beyond simple word overlap.
• Score questions on fluency, relevance, and dificulty using prompt engineering.</p>
          <p>• Employ better models for question generation and evaluation.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Student</title>
        <p>Our Student component is designed to generate concise answers based solely on the context retrieved
from relevant documents.</p>
        <p>Model Selection and Loading We dynamically load language models from Hugging Face’s
Transformers library, specifically GPT-Neo-125M, LLaMA-7B, and Mixtral-8x7B, using PyTorch. Models are
loaded onto GPUs with half-precision floating points (float16) when CUDA is available, optimizing
inference speed and reducing memory footprint.</p>
        <p>Document Processing and Retrieval We preprocess input documents by splitting them into
chunks of 100 words each with an overlap of 20 words to ensure continuity. This is accomplished
using spaCy’s sentence segmentation. Each chunk is then transformed into semantic embeddings
via the BAAI/bge-base-en-v1.5 SentenceTransformer model. These embeddings are normalized
and indexed with FAISS’s IndexFlatIP for fast inner product similarity search, retrieving the top 20
relevant document chunks per query.</p>
        <p>Prompt Construction and Generation For each chunk collected by FAISS, we construct prompts
that explicitly include instructions (“Answer the following QUESTION based only on the provided
CONTEXT.”), context segments retrieved by the FAISS-based search, and the specific question. Context size
management dynamically calculates maximum allowable tokens by considering the model’s maximum
position embeddings minus reserved tokens (set at 256 tokens) and tokens used by instructional text.
Responses are generated using GPU acceleration with PyTorch, sampling from a probability distribution
at a temperature of 0.7 to maintain a balance between answer variability and consistency.
Advanced Generation Control Generation employs a custom stopping criterion that halts the
generation process upon encountering a newline character, thus ensuring concise and coherent
singlesentence responses. The system dynamically manages token budgets strictly to avoid overflow and
eficiently uses available computational resources.</p>
        <p>This detailed, structured approach ensures a robust and eficient Student component, significantly
enhancing the overall efectiveness and reliability of our RAG system.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. Teacher Submission Results</title>
        <p>• Deepsek r1 DistilQwen 14B
• Deepsek r1 14B
• LLaMA 3.2 3B Instruct
In the Teacher system, we provided results generated using pure LLMs:</p>
        <p>Additionally, we applied the generator-discriminator approach. Our primary submission is based on
the generator-discriminator pipeline run with the LLaMA 7B model.</p>
        <p>Teacher Primary Submission Example Output Below is a sample of the generated questions and
reference answers:
"nmtclass/lecture01-eval/screen80-slide80/text.en.txt": [
{
"question": "What units can be used instead of exact word forms for evaluating a
˓→ fundamental problem?",
"reference-answers": [
"So to fix this fundamental problem we can evaluate coarser units, so not exact word
˓→ forms, but lemmas or deep-lemmas."
},
{
},
{
},
{</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Student Results</title>
        <p>In the Student system, we used LLaMA 7B and DeepSeek DistilQwen 1.5B, both running on messages
augmented with RAG-retrieved results. DistilQwen, being a smaller model, tended to hallucinate
less than LLaMA 7B and primarily focused on polishing the retrieved content. In contrast, LLaMA
7B occasionally introduced information not present in the retrieval context—for example, generating
equations in the neural machine translation domain.</p>
        <p>Our final submission is a combination of Qwen and LLaMA outputs. In cases where LLaMA failed
(e.g., due to crashing or missing answers), we defaulted to the DistilQwen response. This combination
was then cleaned using OpenAI: we instructed the model to polish the answers without adding new
content. This post-processing step was particularly useful because we imposed a token generation limit
of 300 tokens to avoid excessively long or of-topic answers and then answers were cut short. The
post-processing helped to at least discard the sentence end that did not make any sense, or keep it there
if it brings information.</p>
        <p>We submitted the following runs:
• LLaMA 7B only
• DeepSeek DistilQwen 1.5B only
• LLaMA + DistilQwen combination
• Primary submission: LLaMA + DistilQwen cleaned by OpenAI
Student Primary Submission Example Output
{
"question": "What is the purpose of each decoder in a transformer model?",
"answer": "The purpose of each decoder in a transformer model is to reconsider and now in
˓→ this second step deliberation, it is following the dependencies that humans found
˓→ important because this upper part is made to mimic the syntactic structure of the
˓→ sentence."
},
{
},
{
},
{
}
"question": "Why are residual connections important in a deep network like a
˓→ transformer?",
"answer": "Residual connections allow the network to reconsider and deliberate on the
˓→ dependencies that humans found important because this upper part is made to mimic the
˓→ syntactic structure of the sentence."
"question": "Why do different types of masks (homemade, surgical, FFP) vary in their
˓→ impact on airflow and droplet containment?",
"answer": "Different types of masks vary in their impact on airflow and droplet
˓→ containment due to differences in their design, materials, and construction. Homemade
˓→ masks, such as cloth masks, may not provide adequate filtration or protection against
˓→ airborne pathogens, while surgical masks are designed to filter out larger droplets and
˓→ may not provide complete protection against smaller droplets or aerosols."
"question": "What are the differences in protective capabilities among FFP-1, FFP-2, and
˓→ FFP-3 masks, and what applications are appropriate for each?",
"answer": "The differences in protective capabilities among FFP-1, FFP-2, and FFP-3 masks
˓→ are based on their respective filtration efficiencies against airborne particles. FFP-1
˓→ masks offer the lowest level of protection, filtering out at least..."</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Future Work and Improvements</title>
      <p>While our current system successfully implements a Teacher-Student pipeline, several opportunities for
enhancement remain. Notably, the Retrieval-Augmented Generation (RAG) paradigm—currently used
only in the Student component—could also be incorporated into the Teacher module. By grounding
question generation in retrieved contextual segments, RAG-enabled Teachers may produce more focused
and relevant questions, particularly for long or complex source documents.</p>
      <p>Another promising direction involves the implementation of an automated Evaluator. Although
not completed in this iteration, we envision two evaluation strategies: (1) lightweight LLMs used as
automatic graders to assess fluency, relevance, and factual accuracy of Student answers; and (2) a
pairwise comparison framework, where multiple generated answers to the same question are evaluated
head-to-head, with a scoring model selecting the superior response. These approaches could provide
meaningful, scalable alternatives to manual annotation.</p>
      <p>Future improvements could also explore:
• Enhanced answer evaluation using semantic similarity metrics, such as embedding-based
comparison (e.g., cosine similarity using Sentence-BERT).
• Incorporation of explainable evaluation criteria (e.g., highlighting which context passages support
an answer).
• Expansion to multilingual materials and models for broader applicability in international
educational contexts.</p>
      <p>Overall, the integration of retrieval in both generation and evaluation stages ofers a compelling path
toward more interpretable, accurate, and resource-eficient systems.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this paper, we presented LLMinds, our submission for the ELOQUENT Sensemaking Task, which
combines open-source large language models in a pipeline designed to generate and answer questions
from instructional materials without relying on external knowledge sources. We successfully
implemented two core components: the Teacher, responsible for generating diverse and context-relevant
questions using a generator-discriminator framework; and the Student, which answers these questions
based solely on RAG-retrieved content using local models and semantic search.</p>
      <p>Our experiments demonstrated that smaller, locally deployed models like LLaMA 7B and DistilQwen
1.5B can be orchestrated to produce coherent and useful results with a minimal computational footprint.
By refining outputs with OpenAI while preserving the original content, we struck a balance between
lfuency and factual grounding.</p>
      <p>Although we did not implement the Evaluator module in this iteration, our framework is designed to
support its future integration. We see great potential in extending this pipeline with automatic answer
scoring and improved semantic evaluation. This work highlights the feasibility of privacy-respecting,
cost-efective alternatives to cloud-based LLMs for educational and research tasks requiring controlled
generation and comprehension.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>LLMs (as a particular type of generative AI tool) are the central element in this study (they create
questions, answer them, and evaluate the answers). They were also used in the production of this
text. In particular we used ChatGPT to rephrase and polish sentences. We also used Overleaf’s
suggestions tool to rephrase our research. After using this tool and service, the authors reviewed and
edited the content as needed and take full responsibility for the publication’s content.
[4] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, A. Fan, V. Chaudhary, T. Rocktäschel,
D. Kiela, et al., Retrieval-augmented generation for knowledge-intensive nlp tasks, in: Advances in
Neural Information Processing Systems, volume 33, 2020, pp. 9459–9474.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Šindelář</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bojar</surname>
          </string-name>
          ,
          <article-title>Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , D. Spina (Eds.), Working Notes of CLEF 2025 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-</article-title>
          <string-name>
            <surname>WS</surname>
          </string-name>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          , in: J.
          <string-name>
            <surname>Burstein</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Doran</surname>
          </string-name>
          , T. Solorio (Eds.),
          <source>Proceedings of the</source>
          <year>2019</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis</article-title>
          , MN, USA, June 2-7,
          <year>2019</year>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <source>Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . URL: https://doi.org/10.18653/v1/n19-
          <fpage>1423</fpage>
          . doi:
          <volume>10</volume>
          .18653/V1/N19-1423.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Language models are unsupervised multitask learners</article-title>
          ,
          <source>OpenAI Blog 1</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>