<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the CLEF 2025 SimpleText Task 2: Identify and Avoid Hallucination</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benjamin Vendeville</string-name>
          <email>benjamin.vendeville@univ-brest.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Bakker</string-name>
          <email>j.bakker@uva.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hosein Azarbonyad</string-name>
          <email>h.azarbonyad@elsevier.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liana Ermakova</string-name>
          <email>liana.ermakova@univ-brest.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaap Kamps</string-name>
          <email>kamps@uva.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Elsevier</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lab-STICC (UMR CNRS 6285)</institution>
          ,
          <addr-line>Brest</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Université de Bretagne Occidentale, HCTI</institution>
          ,
          <addr-line>Brest</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Amsterdam</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents an overview of the CLEF 2025 SimpleText Task 2 on Controlled Creativity. The task aims to identify and avoid hallucination. We discuss the data and benchmarks provided for these tasks, along with preliminary insights and anticipated challenges. Our main findings are the following. First, we used aligned sources, predictions, and references in text simplification to detect and quantify hallucinations-spurious content introduced by generative models-highlighting a critical limitation of current evaluation metrics. Second, we found that overgeneration and information distortion in model outputs can be detected with high accuracy, even without access to the original source text, suggesting that automatic detection is a promising strategy. Third, while automatic methods show promise, the detailed classification of distortions remains dificult to replicate without human expertise, underscoring the continued importance of expert human evaluation and the research challenge of building efective classification models to match this. More generally, we hope and expect that the constructed corpora and evaluation data will be used by researchers to further advance information distortion detection and classification approaches, both in general and specifically for scientific text simplification models.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Scientific text simplification</kwd>
        <kwd>Biomedical AI</kwd>
        <kwd>Generative AI</kwd>
        <kwd>Information access</kwd>
        <kwd>Natural language processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Becoming science-literate is more important than ever before. Objective scientific information helps any
user navigate a world where misinformation, disinformation, or generated and unfounded information
is only a single mouse click away. Everyone acknowledges the importance of objective scientific
information, but the general public seldom consults scientific sources. The value of objective scientific
information cannot be overstated. Biomedical research can directly impact people’s decisions about
health. However, the most reliable and up-to-date sources in biomedicine contain complex language
and assume a high degree of background knowledge, making them dificult for the general public to
understand.</p>
      <p>
        To address these challenges, the CLEF 2025 Simple Track has three aims. First, we push the research
frontier in text simplification by further expanding the scientific text simplification corpora, focusing
on true paragraph-level and document-level simplification with greater variation, and considering the
complex discourse structure. This setup fits current models, such as LLMs, that operate on a long
input. This new biomedical corpus is constructed from aligned Cochrane abstracts and plain language
summaries [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Second, we exploit the text simplification setup with aligned sources, references,
and the output of generative models to detect, quantify, and avoid spurious information introduced
gratuitously by the generative model. This is what is informally referred to as “hallucinations,” addressing
the remaining limitations of large generative models is crucial for the scientific use case, as current
evaluation measures are “blind” and don’t punish the unwarranted generation of additional content.
This task addresses one of the main challenges in the Track, CLEF, and the fields of NLP and IR in
general. Third, by popular demand, we will revisit and rerun some earlier tasks to ensure that the
transition to the new track setup will retain the active track participants of earlier years.
      </p>
      <p>Hence, the CLEF 2025 SimpleText track is based on three interrelated tasks:
• Task 1: Text Simplification simplify scientific text.
• Task 2: Controlled Creativity identify and avoid hallucination.</p>
      <p>• Task 3: SimpleText 2024 Revisited selected tasks by popular request.</p>
      <p>
        This paper gives an overview of the CLEF 2025 SimpleText Task 2 on Controlled Creativity, which aims
to identify and avoid hallucination. Further detail on the entire track is in the CLEF 2025 SimpleText
Track Overview [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Additional details on Task 1 on Text Simplification are in a companion Task 1
overview paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We also refer to the respective participants’ papers for further details.
      </p>
      <p>A total of 74 teams registered for our SimpleText track at CLEF 2025. A total of 18 teams submitted
198 runs in total for Tasks 1 and 2. The statistics for these runs submitted are presented in Table 1.1
However, some runs had problems that we could not resolve. We do not detail them in the rest of the
paper and leave out the 0-scoring runs. More details about individual runs and experiments can be
found in the participants’ papers, also shown in Table 1.</p>
      <p>The rest of this paper is structured in the following way. Section 2 describes the task, the data,
the format, and the evaluation measures. Section 3 describes the participants’ approaches. Section 4
provides detailed results for the task. Section 5 provides further analysis of the results. We end with a
discussion and conclusions in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task 2: Identify and Avoid Hallucination</title>
      <p>This section details Task 2: Controlled Creativity on identify and avoid hallucination.
1The table includes submissions in the Tasks 1 and 2 Codabench evaluation platform, where we were privileged to have 29
(Task 1) and 13 (Task 2) participants.
As various kinds of output devices emerged , such as highresolution printers or a display of PDA ( Personal Digital
Assistant ) , the . The importance of high-quality resolution conversion has been increasing . |This paper proposes a
new method for enlarging an image with high quality . It will involve using a combination of high-speed imaging
and high-resolution video . |One of the largest biggest problems on image enlargement is the exaggeration of the
jaggy edges . This is especially true when the image is enlarged , as in this case . |To remedy this problem , we propose
a new interpolation method , which . This method uses artificial neural network to determine the optimal values of
interpolated pixels . |The experimental results are shown and evaluated . The results are compared to other studies
and found to be inconclusive . |The efectiveness of our methods is discussed by comparing with the conventional
methods . Our methods are designed to help people with mental health problems , not just as a way to cure them . |</p>
      <sec id="sec-2-1">
        <title>2.1. Description</title>
        <p>The Controlled Creativity task aims to identify and avoid hallucination. To our own surprise, the
SimpleText track has collected a massive collection of spurious or overgeneration content from its
participants in earlier years of the track. Table 3 shows an example output simplification of one of the
participating teams. For the CLEF 2024 task on text simplification, a total of 17 out of 36 submissions
(47%) contain spurious whole sentences in at least 10% of the input sentences. In fact, 14 submissions
(39%) have spurious sentences in at least 20% of the input, while 7 submissions (19%) have them in
at least 50% of the input sentences [19]. Our text simplification setup has sources, predictions, and
references that are closely aligned and in the same language. This design allows us to study source
attribution and creative variation while also identifying and avoiding what is informally referred to as
"hallucinations." This task builds on earlier manual analysis of information distortion in our track since
2022 [20, 21, 19], and similar work by others [22].</p>
        <p>Task 2.1 is to identify creative generation, at the abstract or document level. We will provide realistic
system outputs from participants in previous years, along with some intentionally generated outputs
from known models. The task is to identify which sentences are fully grounded in the source input: (a)
without access to the source sentences and (b) with access to them. This also includes labeling sentences
that introduce significant new content. Task 2.1 can be seen as a post-hoc identification task.
Task 2.2 focuses on detecting and classifying information distortion in simplified sentences.
Specifically, it is a multi-label text classification task in which participants are asked to identify the types of
information distortion issues based on the annotation scheme introduced by Vendeville et al. [23]. This
scheme discerns four broad categories of information distortion:
A. Fluency Is the answer provided in a correct form that a fluent speaker would speak?
B. Alignment Is the format of the answer correct?
C. Information Is the information provided accurate and relevant to the input?
D. Simplification Does the response focus on simplification?
Each group contains several fine-grained error types, for a total of 14 classes. 2 The test set is based on
manual annotations, while the training set consists of synthetically generated simplifications containing
targeted errors. Both datasets were constructed using runs submitted to previous editions of the
SimpleText track.
2Our annotation scheme focuses on content and meaning preservation. Following [22], we use the word “error” as a general
term for annotated issues. The term error is used for brevity, acknowledging that some cases can be considered acceptable in
a text simplification context.
Summary: We propose the in vivo/vitro use of prokaryotic adaptive immune systems for distributed learning. In
the coming years synthetic biologists will learn to control, program, and modify such systems. We design an
enhancement to CRISPR-Cas immune systems and demonstrate the learning potential of the modified system
by showing it can approximate solutions to a computationally hard problem. To our knowledge this is the first
proposed use of CRISPR-Cas systems for computational purposes.</p>
        <sec id="sec-2-1-1">
          <title>Generated Simplification</title>
          <p>Summary : We propose the suggest that in vivo/vitro use of prokaryotic adaptive immune systems for distributed
learning . |In the coming years synthetic biologists will learn to control , program , and modify such systems . |We
design an enhancement to found that CRISPR-Cas immune systems and demonstrate the learning potential
of the modified system by showing it can approximate solutions to a computationally hard problem . |To our
knowledge this is the first proposed use of CRISPR-Cas systems for computational purposes .
Task 2.3 Finally, we have a text alignment on avoiding creative generation and performing grounded
generation by design. This task mirrors Task 1 on text simplification and requires the submission of
pairs of runs, both with and without source grounding or source attribution by design.
2.2. Data
In running the SimpleText track over the last three years, we have collected an extensive set of realistic
and representative predictions in the run submissions. For Tasks 2.1 and 2.2, we use this corpus of
realistic generations to build datasets with appropriate labels. Task 2.1 focuses on identifying whether a
generated sentence in the prediction is spurious. This is essentially a sentence label task, and the data was
provided accordingly. We constructed the data using simplifications submitted in previous SimpleText
labs. These simplifications were evaluated based on token-level alignment with the source document.
If more than 10% of the generated tokens could not be aligned with the source, the corresponding
sentence was labeled as "spurious".</p>
          <p>From this, we identify 2 cases:
• "sourced": participants are tasked with labeling the generation with access the source
• "posthoc": participants are tasked with labeling the generation without access the source
We provide both test and train datasets, as well as a source dataset containing the abstracts for the
sourced runs.</p>
          <p>Task 2.2 mimics the human annotation of information distortion as done in earlier years of the track.
Specifically, each simplification was labeled according to the annotation scheme of [23].</p>
          <p>Finally, for Task 2.3, we use the same data as for tasks 1.1 and 1.2, and expect the same format of
output.</p>
          <p>Train data For Task 2.1, we selected 782 abstracts used at CLEF 2024 SimpleText Task 3 on Text
Simplification. The Task 2.1 train data with sentence labels consisted of 13,341 sentences (posthoc)
and 13,514 sentences (sourced). The prevalence was very high: 11,991 (89.9%) sentences were
labeled spurious for posthoc and 12,115 (89.6%) sentences for sourced.</p>
          <p>For Task 2.2, the train data is based on a synthetic dataset starting from simplifications previously
annotated as error-free. We started with submissions from past years that we annotated as
error-free. Then, we used a combination of formal algorithms and large language models (LLMs)
for each error class in the taxonomy to generate variants of the simplification containing the
targeted error class. This approach enabled us to create a large-scale training dataset without
relying on time-expensive manual annotation. The set contained 42,392 sentences with a detailed
information distortion label.</p>
          <p>Test data For Task 2.1, the test data with sentence labels consisted of 3,336 sentences (posthoc) and
3,379 sentences (sourced). The prevalence was very high: 3,006 (90.1%) sentences were labeled as
spurious for posthoc, and 3,033 (89.8%) sentences for sourced.</p>
          <p>The test data for Task 2.2 consists of 2,659 sentences produced by participants in previous years of
the SimpleText challenge. We manually annotated 2,659 sentences with an information distortion
taxonomy of [23]. The submission will be evaluated against these manually annotated sentences.
A total of 820 (30.1%) of sentences had no errors, and 1,839 were classified into four categories
(Fluency, Alignment, Information, and Simplification issues) and 14 detailed types.</p>
          <p>
            Task 2.3 follows the setup of Task 1: Text Simplification in [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ], but requested paired runs with and
without the special processing to avoid ungrounded generation.
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.3. Formats</title>
        <p>2.3.1. Train data
Task 2.1 The two instances—posthoc and sourced—difer mainly in the fields</p>
        <p>For sourced, we construct the gen_id as:
"&lt;anonymised_run_id&gt;//&lt;abs_id&gt;//&lt;sentence_id&gt;"
gen_id and anon_gen_id.
{
}
{
}
Example format for Task 2.1 (sourced):
"abs_id": "G10.1_2010209632",
"sentence": "system and present our results.",
"is_spurious": true,
"gen_id": "35623979//G10.1_2010209632//7"
We also provide the abs_id as a separate field to facilitate joining with the abstracts. For posthoc,
we provide anon_gen_id, formatted similarly, but all components are anonymized:
In both cases, the data is provided in JSON Lines (jsonl) format.</p>
        <p>Example format for Task 2.1 (posthoc):
"sentence": "Here's the simplified sentence:\n\n'Sometimes, when you're playing on a computer
˓→ or tablet, special tiny helpers called 'cookies' can follow you around.",
"is_spurious": true,
"anon_gen_id": "74704850//98491492//4"
Task 2.2 For this task, we provide a jsonl file containing the synthetically generated simplifications,
annotated with error types and identifiers. Each entry includes:
• snt_id: Represents the source document and sentence.
• simp_id: Identifies the simplification and the error-generation algorithm; both are
anonymized.</p>
        <p>Example format for Task 2.2:
{
"source sentence": "Compliance to the GDPR is a problem for organizations, it imposes strict
˓→ constraints whenever they deal with personal data and, in case of infringement, it
˓→ specifies severe consequences such as legal and monetary penalties.",
"simplified sentence": "Organizations face challenges in complying with the GDPR, which sets
˓→ strict rules for handling personal data and imposes penalties for violations.",
"snt_id": "G15.3_2766353613_2",
"simp_id": "429978-180325",
"No error": false,
"A1. Random generation": false,
"A2. Syntax error": false,
"A3. Contradiction": false,
"A4. Simple punctuation / grammar errors": false,
"A5. Redundancy": false,
"B1. Format misalignment": false,
"B2. Prompt misalignment": false,
"C1. Factuality hallucination": false,
"C2. Faithfulness hallucination": false,
"C3. Topic shift": false,
"D1.1. Overgeneralization": true,
"D1.2. Overspecification of Concepts": false,
"D2.1. Loss of Informative Content": false,
"D2.2. Out-of-Scope Generation": false
Task 2.3 follows the same format as tasks 1.1 and 1.2. Example format for Task 2.3 sentence-level:
2.3.2. Test Data
Task 2.1 The format is the same as for training but without the is_spurious field. In both cases, the
data is provided in JSON Lines (jsonl) format.</p>
        <p>Example format for Task 2.1 (posthoc):
"sentence": "I explained the complex terms directly within the simplified sentence: *
˓→ 'Next-generation model' means a new and improved plan.",
"anon_gen_id": "74704850//66348262//3"
Example format for Task 2.1 (sourced):
"abs_id": "G01.1_1570837852",
"sentence": "In this paper, we share our findings on how evolutionary algorithms and
˓→ multi-agent systems can be used to understand a user's preferences while they interact
˓→ with a digital assistant.",
"gen_id": "11102757//G01.1_1570837852//1"
"pair_id": "CD009102",
"complex": "However, the evidence is very uncertain.",
"simple": "['As a result, we have little confidence in the evidence and the results of this
˓→ outcome should be interpreted with caution.']"
Example format for Task 2.3 abstract-level:
"pair_id": "CD008996",
"complex": "A total of 1437 adult patients participated in the five randomized parallel
˓→ group studies, with treatment durations ranging from 8 to 16 weeks. The daily doses of
˓→ eplerenone ranged from 25 mg to 400 mg daily. Meta-analysis of these studies showed
˓→ [...]",
"simple": "These studies followed patients for 8 to 16 weeks while on therapy. The doses of
˓→ eplerenone used in these studies ranged from 25 mg to 400 mg daily. None of the studies
˓→ reported on the clinically meaningful outcomes of eplerenone, such as whether
˓→ eplerenone can reduce [...]"
Task 2.2 The format is the same as for training but without the labels in error field. The data is provided
in JSON Lines (jsonl) format.
Task 2.3 follows the same format as tasks 1.1 and 1.2. Example format for Task 2.3 sentence-level:
"pair_id": "CD012520",
"para_id": 0,
"sent_id": 0,
"complex": "We included seven cluster-randomised trials with 42,489 patient participants
˓→ from 129 hospitals, conducted in Australia, the UK, China, and the Netherlands."
Example format for Task 2.3 abstract-level:
"pair_id": "CD012520",
"source": "Cochrane",
"complex": "We included seven cluster-randomised trials with 42,489 patient participants
˓→ from 129 hospitals, conducted in Australia, the UK, China, and the Netherlands. Health
˓→ professional participants (numbers not specified) included nursing, medical and allied
˓→ health professionals. Interventions in all studies included [...]"
2.3.3. Sources
For Task 2.1 sourced, we provide a source file containing all abstracts used. This file is also in jsonl
format and includes the following fields:
{
}
"query_id": "G07.1",
"query_text": "misinformation",
"doc_id": 2100028027,
"abs_id": "G07.1_2100028027",
"abs_source": "Inaccurate information, in the field of library and information science, is often
˓→ regarded as a problem that needs to be corrected or simply understood as either
˓→ misinformation or disinformation without [...]"
2.3.4. Predictions
In all cases, we asked for JSON submissions, but during evaluation we tried to parse JSON, JSONL, CSV
and TSV formats to fix any error by the participants. We also expected a run_id of the format:
&lt;team-name&gt;_&lt;task-name&gt;_&lt;method-used&gt;</p>
        <p>In practice, some participants used "_" in the method names so we parsed everything after the task
name into the method used.</p>
        <p>Task 2.1 For this task, we expected a JSON file containing the sentence, identifier (either
anon_gen_id), is_spurious label, and the run_id
gen_id or
Example format for Task 2.1 sourced
"sentence":"In this paper, we share our findings on how evolutionary algorithms and
˓→ multi-agent systems can be used to understand a user's preferences while they interact
˓→ with a digital assistant.",
"gen_id":"11102757//G01.1_1570837852//1",
"is_spurious":false,
"run_id":"UBOnlp_task21sourced_gpt4o"
Task 2.2 For this task, we expected a JSON file containing the source and simplified sentences, the
snt_id and simp_id identifiers, the the run_id, and a label for each error class.</p>
        <p>Example format for Task 2.1 sourced
Example format for Task 2.1 posthoc
"sentence":"I explained the complex terms directly within the simplified sentence:\n\n*
˓→ \"Next-generation model\" means a new and improved plan.",
"anon_gen_id":"74704850//66348262//3",
"is_spurious":false,
"run_id":"UBOnlp_task21posthoc_gpt4o"
Task 2.3 Example format for Task 2.3 Abstract level
"pair_id": "CD012520",
"source": "Cochrane",
"complex": "We included seven cluster-randomised trials with 42,489 patient participants
˓→ from 129 hospitals, conducted in Australia, the UK, China, and the Netherlands. Health
˓→ professional participants (numbers not specified) included [...]",
"run_id": "AIIRLab_task12_Mistral_7b_base_grounded"
},
Example format for Task 2.3 sentence level
"pair_id":"CD012520","para_id":0,"sent_id":0,"complex":"We included seven
˓→ cluster-randomised trials with 42,489 patient participants from 129 hospitals,
˓→ conducted in Australia, the UK, China, and the Netherlands.",
"prediction":"We studied seven trials with 42,489 patients from 129 hospitals in four
˓→ countries.",
"run_id":"dsgt_Task11_plan_guided_llama_grounded"</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.4. Codabench</title>
        <p>Submissions were made through Codabench.3 Due to the diferences in the setup, each task had a
designated separate competition on Codabench. The Task 1 runs were submitted at: https://www.
codabench.org/competitions/8400/. The Task 2 runs were submitted at: https://www.codabench.org/
competitions/8327/ (shown in Figure 1). The Codabench greatly facilitated running the track in 2025
and provided active participants (who had also registered at the Codabench) with full access to the
competition, including the submission and leaderboard pages.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.5. Evaluation</title>
        <p>
          Task 2.1 is essentially a sentence label task, evaluated in the standard way with Precision, Recall, F1,
and AUC. Task 2.2 is a multi-label classification task. We evaluate performance using both F1 score
and AUC, computed for individual classes and aggregated across the four main classes. Task 2.3 will
be evaluated by both standard automatic measures and human evaluation, following Task 1 on Text
Simplification in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. We also conduct a more detailed overgeneration analysis for Task 2.3.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Participant’s Approaches</title>
      <p>A total of 9 teams submitted 66 runs in total. In the detailed results, we only include runs without errors,
which got a non-zero score.</p>
      <p>
        AIIRLab Largey et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] submitted 10 runs in total for Task 2. They submitted five runs for Task 2.1,
ifve runs for Task 2.2, and none for Task 2.3. They use a combination of four diferent methods for
detecting spurious sentences: an abstract meaning representation, an encoder classifier, majority voting
over three models (QWEN, Mistral, LLaMA), and extensive textual features. Furthermore, they use a
trained RoBERTa classifier for multi-label prediction, and an ensemble of three models (LLaMA, Mistral,
and Openchat) to classify each type of information distortion. Their Task 2.3 were submitted to Task 1.1,
were they deployed their Task 1 approach with special precautions against noise and unwanted output.
DSGT Marturi and Elwazzan [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] submitted 15 runs in total for Task 2. They submitted six runs for
Task 2.1, six runs for Task 2.2, and three runs for Task 2.3. The paper uses an advanced set of approaches,
including classifiers, semantic similarity, entailment, and LLM as a Judge, for Task 2.1. They use
DeBERTa and LLaMA classifiers for Task 2.2. Finally, for Task 2.3, they repurposed the Task 1 two-stage
approach and added a third stage in which they use LLaMA to check and revise the output for content
not in the source document. Details about the Task 1 approach are in a separate paper [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
DUTH Arampatzis and Arampatzis [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] submitted four runs in total for Task 2. They submitted two
runs for Task 2.1, two runs for Task 2.2, and none for Task 2.3. For Task 2.1, the paper uses a set of
classifiers trained on lexical features to detect spurious sentences and obtains high performance. For
Task 2.2, a multi-class classifier is trained on the embeddings of the sentence pairs to classify the pairs
into the given labels.
      </p>
      <p>Mtest (no paper) submitted two runs in total for Task 2. They submitted one run for Task 2.1, one run
for Task 2.2, and none for Task 2.3.</p>
      <p>
        RECAIDS Eugin et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] submitted two runs in total for Task 2. They submitted one run for Task
2.1, one run for Task 2.2, and none for Task 2.3. They explore a T5 model for Tasks 2.1 and 2.2, using a
straightforward T5 completion prompt, with a model fine-tuned on each task.
      </p>
      <p>
        SINAI Collado-Montañez et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] submitted 30 runs in total for Task 2. They submitted 15 runs
for Task 2.1, 15 runs for Task 2.2, and none for Task 2.3. They use a rule-based approach to Task 2.1,
exploiting some features or artifacts of the data and task setup, followed by an LLaMA model for final
classification. Their results show high degrees of efectiveness under both conditions of having access
to the source.
      </p>
      <p>UBO Vendeville et al. [16] submitted two runs in total for Task 2. They submitted one run for Task 2.1,
one for Task 2.2, and none for Task 2.3. The submissions were mostly test submissions, but the paper
documents an interesting LLM approach to directly apply the annotation scheme for information
distortion as used in the human evaluation of Task 2.2.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>This section details the task results for the overgeneration detection subtask, information distortion
detection, classification subtask, and the grounded text simplification subtask.</p>
      <sec id="sec-4-1">
        <title>4.1. Task 2.1: Identify Creative Generation</title>
        <p>Task 2.1 aims to identify overly creative generation in scientific text simplification. This is a new task
that focuses on detecting overgeneration and other information distortion issues in the predictions
of current models. The task raises awareness of remaining information distortion issues in modern
generative models for scientific text simplification, and focuses on post-hoc detection without or with
access to the source text.
count</p>
        <p>Table 4 shows the results of detecting spurious sentences in the generated simplifications of
participants in the track in earlier years. The main task is post-hoc detection without access to the source
texts, which would generalize to generic text generation tasks.</p>
        <p>We make several observations. First, the scores are generally high, with many systems performing
over 90% accuracy, F1, and AUC-PR. The test collection contains a variety of information distortion
issues (see Task 2.2 and Task 2.3 for more details), including some clear “errors” such as leaving in
prompts, or systematic errors in extracting the simplified content from the output of models. However,
it also contains complex cases to detect (like the example in Table 3). Hence, the performance is
encouraging. Second, it is interesting that trained classifiers such as encoders seem to outcompete
larger and modern models as decoders for this task. This may result from the specific task setting,
where efective training will pay of. Third, while the task was intended to present entire abstracts or
documents, a sentence label task was more practical to run in this first year. This may have efectively
reduced the task to a sentence-level task, which may have been easier than a long document-level task.</p>
        <p>Table 5 also shows the results of detecting spurious sentences in the generated simplifications of
participants in the track in earlier years, while having access to pairs of source-prediction content. This
setting exploits the text simplification setting, in which information generation must faithfully reflect
the source content.</p>
        <p>We make several observations. First, access to the sources would intuitively make the task far easier:
human assessors generally rely on this to make their judgments. We see a notable increase in the
performance of models, even in AUC-RO, which was lagging in Table 4 above. Second, similar to above,
we see that trained or fine-tuned encoders are very efective, generally outcompeting larger decoder
models with prompting and few-shot, in-context learning. Third, in the context of source-prediction
pairs of sentences, the task is more straightforward than observing a long source document paired to a
lengthy list of prediction sentences. Still, the near-perfect performance of the best submissions is very
encouraging.</p>
        <p>This completes the discussion of the Task 2.1 experiments. For the source-prediction pairs, we
expected that the better systems would be able to perform close to perfection. These results indicate that
it is possible to detect information distortion errors, such as overgeneration, in the output of current
systems. Current evaluation measures based on the overlap with references are insensitive to such
additions or redundant content freely generated by the models. Efective detection models can help
identify and quantify these issues in the output of models, which is of great importance in further
advancing scientific text simplification models.
count</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Task 2.2: Detect and Classify Information Distortion Errors</title>
        <p>Task 2.2 is a new task that asks not only to detect information distortion in the output of text
simplification models but also to classify the type of error. This task mimics the human manual evaluation we
performed in the track in earlier years.</p>
        <p>We evaluate this task using a corpus of 2,659 manually annotated sentence–simplification pairs. Each
simplified sentence may contain multiple error types, making this a multi-label classification problem.
The error taxonomy is organized hierarchically into four categories (A–D), each comprising several
ifne-grained error types. For evaluation, predicted and gold error labels are aggregated at the group
level: if any fine-grained error from a group is present, the group is considered active. Performance is
then measured per group using both micro and macro F1 scores. We also consider the "No Error" class,
indicating no errors were detected.</p>
        <p>Results are presented in Table 6. The table includes only valid submissions, excluding 39 duplicates
where teams submitted the same method multiple times, where we retain only the run with the highest
F1 score on the No Error class. The results displayed here are limited to the best five runs per team, and
are sorted by F1 score on No Error class.</p>
        <p>From this, we make several observations. First, while some models were able to perform well on No
Error, achieving over 0.65 F1 scores, performance quickly drops, and over half of them do not achieve
0.50 F1 scores. Second, results are quite low for all other groups. For Fluency issues (group A), the five
best systems achieve an F1 score between 0.255 and 0.283. For Alignment (group B), only 55% of the
systems achieved over 0.10 F1 scores, with 20% over 0.25 and up to 0.47. For Information issues (group
C), only two systems achieved over 0.25 F1 scores (with 0.30 and 2.69), with the next 60% achieving
between 0.10 and 0.17. Finally, for Simplification issues (group D), only the same two models were able
to achieve F1 scores above 0.25 (with 0.37 and 0.30) while the next 60% achieved between 0.12 and 0.24.
Third, more generally, the results suggest that detecting specific error categories remains a challenging
task, especially under realistic conditions with a multi-label setting and imbalanced data. The relatively
strong performance on the No Error class demonstrates that distinguishing error-free simplifications
is a realistic and tractable subtask. The gap between detecting no errors and identifying fine-grained
error types remains an open research challenge, and the results of the track highlight the complexity of
accurately modeling semantic information distortions in the output of current models.</p>
        <p>This completes the discussion of the Task 2.2 experiments. The results are mixed. On the one hand,
consistent with the results of Task 2.1, we saw that detecting that a prediction has information distortion
issues is a viable task for current systems. On the other hand, fine-grained annotation of the types of
information distortion remains challenging. This indicates that manual evaluation remains of great
value for scientific text simplification and the automatic evaluation measures. Yet the efort and cost of
manually annotating all output remains very high, and such human evaluation is not reusable and has
to be repeated for every new prediction. One realistic option is to use a hybrid approach. The ability to
automatically filter out the cases with no error and judge samples of the remaining predictions to assess
the error types and distribution can be a pragmatic and more cost-efective way to scale up human
evaluation.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Task 2.3: Avoid Creative Generation</title>
        <p>Task 3.2 aims to avoid overly creative generation in scientific text simplification and showcase systems
that perform grounded generation by design. This is a new task that asks for a pair of submissions, one
of which must make a special efort to avoid overgeneration or other information distortion issues.</p>
        <p>Two teams submitted runs for Task 2.3, indicated by "_grounded" in the run names. Some of these
runs were specifically submitted to Task 2.3, and others were regular submissions to Tasks 1.1 and 1.2.</p>
        <p>
          Table 7 shows the standard evaluation of text simplification output against text overlap with the
reference plain language summaries. We evaluate against the Cochrane-auto aligned cases (top) and
the larger set of original plain language summaries (bottom). We tried to locate the matching baseline
runs, indicated with ⋆ in the tables from the earlier results as displayed for Task 1 in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          We make several observations. First, the performance is generally competitive, and several runs
are among the best-performing runs. This is reassuring, as any attempt to ground the predictions
more closely to the source texts should not lead to a dramatic decrease in performance. Second, the
baseline runs without any precautions observe the highest number of additions, indicating that the
grounded runs are generally more conservative. Third, although we refer to the participants’ papers of
AIIRLab [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and DSGT [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] for specific details, some of the grounding seems to involve more careful
output processing, such as ensuring in the prompts that no extra information other than the text
simplification is output by the model.
        </p>
        <p>More generally, while the primary goal of prediction grounding is not a performance improvement,
it is also the case that other runs with presumably redundant information are not performing less
well. The standard measures based on textual overlap with the references are relatively insensitive to
additional content in the predictions. This invites further analysis to investigate how well the source
information grounds the predictions, and when they are not.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Findings</title>
        <p>This concludes the results for the CLEF 2025 SimpleText Task 2: Controlled Creativity on identify
and avoid hallucination. Our main findings are the following: First, for Task 2.1 on detecting creative
generation, we observed very high performance for identifying overgeneration and other information
distortion. This was hoped and expected for pairs of source-prediction content, but unexpected for
post hoc detection on only the system’s predictions. Second, for Task 2.2 on classifying the type of
information distortion, we observed mixed results. Also here we saw solid performance for the "no
error" cases, yet identifying the precise type of information distortion similar to human evaluation
remains a challenging tasks for current models. Third, for Task 2.3 on avoiding creative generation and
performing grounded generation by design, we observed that text simplification measures are immune
to detecting overgeneration, and that this remains a serious issue in the predictions. More sensitive text
simplification evaluation measures are needed to highlight these aspects and ensure that the research
community further develops grounded generation approaches.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Analysis</title>
      <sec id="sec-5-1">
        <title>5.1. Task 2.1 Analysis</title>
        <p>Task 2.1 focused on identifying spurious sentences i.e. those that introduce content not grounded in
the source text. Participants tackled this task in two settings: post-hoc, where only the prediction
was available, and sourced, where both the source and the generated sentence were provided. This
distinction simulates real-world scenarios where access to the source may or may not be available, and
allows us to explore the limits of detection in both conditions.</p>
        <p>Results were encouraging in both settings, but showed clear diferences. In the post-hoc setting,
several systems still reached high scores over 90% accuracy and F1, suggesting that many spurious
sentences are detectable based on surface cues alone. Obvious cases like prompt leaks or formulaic
overgeneration patterns were often caught even without source access. However, this setting is
inherently more challenging, and performance varied more across teams.</p>
        <p>In the sourced setting, access to the input significantly improved model performance. Top submissions
achieved near-perfect results, with F1 scores up to 0.99 and very high precision and recall. Having
the source allowed systems to make more reliable decisions about whether a sentence was actually
grounded, especially in borderline or more nuanced cases.</p>
        <p>Interestingly, in both settings, trained encoders and task-specific classifiers generally outperformed
larger language models relying on in-context learning. This suggests that for this type of targeted
detection task, fine-tuning on aligned examples still ofers a strong advantage over general LLM
prompting.</p>
        <p>Another important aspect is task framing. While the original goal was to evaluate grounding at the
document level, we focused on sentence-level labels for this first edition. This likely made the task
more approachable, especially for systems that don’t model discourse-level context.</p>
        <p>These results suggest that while source access helps, it’s still possible to detect many hallucinations
post-hoc. Efective detection tools are valuable for both evaluating model outputs and flagging risky
generations in practice.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Task 2.2 Analysis</title>
        <p>Task 2.2 pushed systems beyond simple error detection by asking them to identify what kind of
information distortion occurred, based on a 14-class taxonomy. This proved to be much harder than
determining whether a sentence was error-free.</p>
        <p>Most systems performed reasonably well on the No Error class where several reached F1 scores above
0.70. But performance dropped sharply for the error categories. For example, only a few models scored
above 0.30 F1 on Fluency or Simplification issues, and many had near-zero scores for rarer types like
Overspecification or Topic shift.</p>
        <p>Several factors likely contributed to this. First, the fine-grained labels often overlap and can be subtle,
even for human annotators. Second, systems were trained on synthetic data but evaluated on real,
human-annotated outputs, which may have led to generalization issues. Third, many error types were
underrepresented, and few approaches explicitly addressed this imbalance.</p>
        <p>Interestingly, the systems that performed best on the No Error class also tended to score highest on
Fluency errors (Group A), but this pattern didn’t hold across other categories. For groups B–D, the
correlation with No Error performance was much weaker.</p>
        <p>The strongest results came from ensemble systems like DSGT’s, which combined DeBERTa with
LLaMA-based models, and AIIRLab’s voting ensemble over several LLMs. In contrast, smaller models or
rule-based classifiers struggled more, especially with semantically complex errors.</p>
        <p>These results suggest that while identifying clean outputs is a realistic goal, explaining what went
wrong remains an open challenge. A promising next step could be a hybrid setup: automatic filtering of
error-free outputs, followed by targeted human review of potential issues. This could make evaluation
more scalable without sacrificing quality.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Task 2.3 Analysis</title>
        <p>We analyzed the entire test data set, comprising 666 documents (Task 1.2) and 9,160 sentences (Task 1.1).
This analysis assumes that there is always word overlap between a pair of complex-simple sentences or
abstracts. Moreover, we look specifically for overgenerating output at the sentence’s or abstract’s end.
This is typical of sequence-to-sequence models, which are asked to complete the input with a simplified
version in standard text completion mode.</p>
        <p>Assume we feed the model one long sentence extracted from an abstract, without further context.
Now, due to sentence splitting, the output could contain multiple sentences. However, after the input
sentence is fully simplified, the model wants to complete the text. Without access to the rest of the
source abstract, the model may generate the most likely subsequent sentences. Such sentences are
completely unfounded by the source, and it isn’t easy to spot these cases in the generated text, as they
are indeed coherent and possible continuations. This may occur after every sentence in sentence-level
text simplification.</p>
        <p>In document-level text simplification, this is more likely at the end of the abstract, so we still look at
the end of the source input. We observe, indeed, overgeneration/text completion issues at the end of the
sources/predictions. There are also cases in which there are systematic errors in extracting the output,
with additional content. Increasingly, there is additional LLM commentary other than the requested
output. Accurately removing such additional content can be more challenging for the document-level
submissions than for the sentence-level submissions, as some abstracts are very long.</p>
        <p>Table 8 shows an overgeneration analysis of the Task 2.3 runs. This is done by aligning the source
input to the prediction output regarding their token sequences. If all the source sentence(s) have been
aligned to some prediction sentence(s), we assume the prediction covers all the content of the sources. If
there is still an additional sentence in the prediction, we regard this as spurious content for that specific
input. This is an imperfect proxy, and aligning lengthy documents can be non-trivial. It serves as a
good indicator of spurious content in the predictions and of overgeneration issues in the runs.</p>
        <p>
          We make several observations. First, despite competitive performance in terms of text overlap
with the references, we see widely varying numbers of cases of overgeneration, ranging from a few
percentage points to large fractions of the output. Second, this diference in additional content is not at
all reflected in the evaluation scores, as some of the top-performing runs still exhibit larger fractions of
“extra” content. Some of these may be easily spotted as “noise,” such as systematically left-in prompts.
Other cases may be challenging to detect in the output by users of text simplification systems. Third, in
the context of the task, we see some interesting examples, for example, AIIRLab [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] detected “noise”
and changed the prompts to ensure only the simplified text, and nothing else, was in the model output.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and Conclusions</title>
      <p>This paper describes the setup of the CLEF 2025 SimpleText track, which contains the following three
tasks. Task 1 on Text Simplification : simplify scientific text . Task 2 on Controlled Creativity: identify and
avoid hallucination. Task 3 on SimpleText 2024 Revisited: selected tasks by popular request. This Task
overview focuses on the CLEF 2025 SimpleText Track’s Task 2 on identifying and avoiding information
distortion (or “hallucination”). The main aim of our track, and the CLEF evaluation forum as a whole, is
i) to construct corpora and evaluation resources to stimulate research on scientific text summarization
and simplification, and ii) to foster a community of IR, NLP, and AI researchers working together on
the important task of making science more accessible for everyone.</p>
      <p>Within the CLEF 2025 SimpleText Task 2, we have constructed extensive corpora and references for
evaluation data. First, we exploited the text simplification setup with aligned sources, references, and the
output of generative models to detect, quantify, and avoid spurious information introduced gratuitously
by the generative model. This is what is informally referred to as “hallucinations.” Addressing the
remaining limitations of large generative models is crucial for the scientific use case, as current evaluation
measures are “blind” and don’t punish the unwarranted generation of additional content. Second, we
observed very high accuracy in detecting overgeneration and other types of information distortion
in the output of text simplification systems. This task was based on the real output of CLEF 2024
submissions, and the best systems could detect sentence-level information distortion in the predictions
with near-perfect accuracy in the presence of the sources. Unexpectedly, the accuracy without access
to the source was also very high, even though this may be partly due to the class imbalance in the data.
This is a positive result, as the automatic detection of noise and overgeneration in the output of AI
models appears to be a viable strategy. Third, detailed classification of information distortion, as is
done in small-scale human evaluation, remains challenging to mimic. This may be partly attributed to
human inter-annotator variation and the need for detailed, qualitative judgments that require extensive
expertise and training. This highlights the remaining value of detailed human analysis in addition to
automatic evaluation measures. At the same time, this presents a significant research challenge for
future research to address.</p>
      <p>These reusable corpora and evaluation resources are available to participants and other researchers
who want to work on the important problem of making scientific information open and easily accessible
for everyone. In terms of building a community for researching scientific text summarization and
simplification, the track saw a record attendance in 2025, with significant changes in tasks and the
move to Codabench. More runs were submitted, and the largest number of participating teams ever
was achieved.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We are incredibly thankful to the master’s students in translation and technical writing from the
University of Brest for participating in data annotation. We also thank each of the individual track
participants for their efort in submitting a record number of submissions to Codabench and documenting
these in their papers.</p>
      <p>We thank the CLEF 2025 chairs for hosting us, and the CLEF 2025 Labs and Proceedings chairs for
their excellent assistance and flexibility. It is heartwarming to be part of such a great CLEF family.
We thank Codabench [24] for hosting the competition. Post-competition experiments are ongoing at
https://www.codabench.org/competitions/8400/ (Task 1.1, Task 1.2, and Task 2.3) and https://www.
codabench.org/competitions/8327/ (Task 2.1 and Task 2.2). We hope and expect that these “living test
collections” remain in active use until the next iteration of the track.</p>
      <p>Benjamin Vendeville and Liana Ermakova are partly funded by the French National Research Agency
(ANR-22-CE23-0019-01, Automatic Simplification of Scientific Texts ). Liana Ermakova is further supported
by the CNRS research group MaDICS (https://www.madics.fr/ateliers/simpletext/).</p>
      <p>Jan Bakker and Jaap Kamps are partly funded by the Netherlands Organization for Scientific Research
(NWO NWA # 1518.22.105). Jaap Kamps is further supported by (NWO CI # CISC.CC.016), the University
of Amsterdam (AI4FinTech program), and ICAI (AI for Open Government Lab). Views expressed in this
paper are not necessarily shared or endorsed by those funding the research.</p>
    </sec>
    <sec id="sec-8">
      <title>Disclosure of Interests</title>
      <p>The authors have no competing interests to declare that are relevant to the content of this article.</p>
    </sec>
    <sec id="sec-9">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT and Grammarly in order to: Grammar
and spelling check and Paraphrase and reword. After using these tools/services, the authors
reviewed and edited the content as needed and take full responsibility for the publication’s content.
[16] B. Vendeville, L. Ermakova, P. D. Loor, J. Kamps, UBONLP Report on the SimpleText lab, in: [25],
2025.
[17] P. Kocbek, G. Stiglic, UM-FHS at the CLEF 2025 SimpleText Track: Comparing No-Context and
Fine-Tune Approaches for GPT-4.1 Models in Sentence and Document-Level Text Simplification,
in: [25], 2025.
[18] T. Papandreou, J. Bakker, J. Kamps, University of Amsterdam at the CLEF 2025 SimpleText Track,
in: [25], 2025.
[19] L. Ermakova, V. Laimé, H. McCombie, J. Kamps, Overview of the CLEF 2024 simpletext task 3:
Simplify scientific text, in: G. Faggioli, N. Ferro, P. Galuscáková, A. G. S. de Herrera (Eds.), Working
Notes of the Conference and Labs of the Evaluation Forum (CLEF 2024), Grenoble, France, 9-12
September, 2024, volume 3740 of CEUR Workshop Proceedings, CEUR-WS.org, 2024, pp. 3147–3162.</p>
      <p>URL: https://ceur-ws.org/Vol-3740/paper-307.pdf.
[20] L. Ermakova, I. Ovchinnikova, J. Kamps, D. Nurbakova, S. Araújo, R. Hannachi, Overview of
the CLEF 2022 simpletext task 3: Query biased simplification of scientific texts, in: G. Faggioli,
N. Ferro, A. Hanbury, M. Potthast (Eds.), Proceedings of the Working Notes of CLEF 2022
Conference and Labs of the Evaluation Forum, Bologna, Italy, September 5th - to - 8th, 2022,
volume 3180 of CEUR Workshop Proceedings, CEUR-WS.org, 2022, pp. 2792–2804. URL: https:
//ceur-ws.org/Vol-3180/paper-237.pdf.
[21] L. Ermakova, S. Bertin, H. McCombie, J. Kamps, Overview of the CLEF 2023 simpletext task
3: Simplification of scientific texts, in: M. Aliannejadi, G. Faggioli, N. Ferro, M. Vlachos (Eds.),
Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2023), Thessaloniki,
Greece, September 18th to 21st, 2023, volume 3497 of CEUR Workshop Proceedings, CEUR-WS.org,
2023, pp. 2855–2875. URL: https://ceur-ws.org/Vol-3497/paper-240.pdf.
[22] A. Devaraj, W. Shefield, B. Wallace, J. J. Li, Evaluating factuality in text simplification, in:
S. Muresan, P. Nakov, A. Villavicencio (Eds.), Proceedings of the 60th Annual Meeting of the
Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational
Linguistics, Dublin, Ireland, 2022, pp. 7331–7345. URL: https://aclanthology.org/2022.acl-long.506/.
doi:10.18653/v1/2022.acl-long.506.
[23] B. Vendeville, L. Ermakova, P. D. Loor, Resource for Error Analysis in Text Simplification: New</p>
      <p>Taxonomy and Test Collection, 2025. doi:10.1145/3726302.3730304. arXiv:2505.16392.
[24] Z. Xu, S. Escalera, A. Pavão, M. Richard, W. Tu, Q. Yao, H. Zhao, I. Guyon, Codabench: Flexible,
easy-to-use, and reproducible meta-benchmark platform, Patterns 3 (2022) 100543. URL: https:
//doi.org/10.1016/j.patter.2022.100543. doi:10.1016/J.PATTER.2022.100543.
[25] G. Faggioli, N. Ferro, P. Rosso, D. Spina (Eds.), Working Notes of CLEF 2025: Conference and Labs
of the Evaluation Forum, CEUR Workshop Proceedings, CEUR-WS.org, 2025.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bakker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <article-title>Cochrane-auto: An aligned dataset for the simplification of biomedical abstracts</article-title>
          , in: M.
          <string-name>
            <surname>Shardlow</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Saggion</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Alva-Manchego</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Zampieri</surname>
          </string-name>
          , K. North, S. Štajner, R. Stodden (Eds.),
          <source>Proceedings of the Third Workshop on Text Simplification, Accessibility and Readability (TSAR</source>
          <year>2024</year>
          ),
          <article-title>Association for Computational Linguistics</article-title>
          , Miami, Florida, USA,
          <year>2024</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>51</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .tsar-
          <volume>1</volume>
          .5/. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .tsar-
          <volume>1</volume>
          .5.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Azarbonyad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bakker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vendeville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <article-title>Overview of the CLEF 2025 SimpleText track: Simplify scientific texts (and nothing more)</article-title>
          , in: J. Carrillo de Albornoz,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mothe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Sixteenth International Conference of the CLEF Association (CLEF</source>
          <year>2025</year>
          ), Lecture Notes in Computer Science, Springer,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bakker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vendeville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <article-title>Overview of the CLEF 2025 SimpleText Task 1: Simplify Scientific Text</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>N.</given-names>
            <surname>Largey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Mansouri,</surname>
          </string-name>
          <article-title>AIIRLab Systems for CLEF 2025 SimpleText: Cross-Encoders to Avoid Spurious Generation</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Djoudi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nouali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aabid</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Badache</surname>
          </string-name>
          , A.-G. Chifu, P. Bellot, LIS at the SimpleText 2025:
          <article-title>Enhancing Scientific Text Accessibility with LLMs and Retrieval-Augmented Generation</article-title>
          , in: [ 25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K. C.</given-names>
            <surname>Marturi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. H.</given-names>
            <surname>Elwazzan</surname>
          </string-name>
          ,
          <article-title>Hallucination Detection and Mitigation in Scientific Text Simplification using Ensemble Approaches: DS@GT at CLEF 2025 SimpleText</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K. C.</given-names>
            <surname>Marturi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. H.</given-names>
            <surname>Elwazzan</surname>
          </string-name>
          ,
          <article-title>LLM-Guided Planning and Summary-Based Scientific Text Simplification: DS@GT at CLEF 2025 SimpleText</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Arampatzis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Arampatzis</surname>
          </string-name>
          , DUTH at CLEF 2025 SimpleText Track:
          <article-title>Tackling Scientific Text Simplification and Hallucination Detection</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. M.</given-names>
            <surname>Agüero-Torales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rodríguez-Abellán</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A. C.</given-names>
            <surname>Moraga</surname>
          </string-name>
          ,
          <article-title>Sentence-level Scientific Text Simplification With Just a Pinch of Data</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gallina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jiménez</surname>
          </string-name>
          , S. Huet, University of Avignon at SimpleText 2025:
          <article-title>Guided Medical Abstract Simplification</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chaudhari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hotha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sonawane</surname>
          </string-name>
          , S-3
          <string-name>
            <surname>Pipeline by</surname>
            <given-names>PICT</given-names>
          </string-name>
          /
          <article-title>Pune for Biomedical Text Simplification</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Eugin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ms.Beula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sathvikha</surname>
          </string-name>
          , V. Sangamithra, SimpleText: Simplify Scientific Text, in: [ 25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Dongre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaadiraaju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Madasamy</surname>
          </string-name>
          ,
          <article-title>NITK SCaLAR Lab at the CLEF 2025 SimpleText Track: Transformer-Based Models for Biomedical Sentence Simplification (Task 1.1)</article-title>
          , in: [ 25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Collado-Montañez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Ortiz-Zambrano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Espin-Riofrio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Montejo-Ráez</surname>
          </string-name>
          ,
          <source>SINAI in SimpleText CLEF</source>
          <year>2025</year>
          :
          <article-title>Simplifying Biomedical Scientific Texts and Identifying Hallucinations Using GPT-4.1</article-title>
          and
          <string-name>
            <given-names>Pattern</given-names>
            <surname>Detection</surname>
          </string-name>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dauenhauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. O.</given-names>
            <surname>Dietzler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. D.</given-names>
            <surname>Idahor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. K.</given-names>
            <surname>Kreutz</surname>
          </string-name>
          , THM@
          <article-title>SimpleText 2025 Task 1.1: Revisiting Text Simplification based on Complex Terms for Non-Experts</article-title>
          , in: [25],
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>