<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Systematic Literature Review on Hallucination Detection Methods in LLMs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elisa Bestetti</string-name>
          <email>elisa.bestetti@sanoma.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zheying Zhang</string-name>
          <email>zheying.zhang@tuni.fi</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kostas Stefanidis</string-name>
          <email>konstantinos.stefanidis@tuni.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Science Research Centre, Tampere University</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>Large Language Models (LLMs) are increasingly used across diverse applications, from healthcare and education to legal services and journalism. As their use expands into domains where factual accuracy is critical, the challenge of detecting and mitigating hallucinations has become essential. This research conducts a systematic literature review on hallucination detection methods in LLMs, restricting the scope to detection methods that do not rely on task-specific grounding documents provided at the time of generation (in RAG, summarization, or translation). Instead, the review focuses on open-domain detection, which includes methods that may use public external knowledge (for example Wikipedia or Wikidata) for post-hoc verification of the model's inherent outputs. The review synthesizes 50 peer-reviewed studies published between 2023 and 2025, classifying detection strategies according to hallucination type, i.e. factuality or faithfulness, and technical approach of white-box or black-box. The findings reveal a predominance of factuality-focused methods and black-box techniques, reflecting practical constraints in accessing proprietary model internals. Prominent approaches include LLM-as-a-judge, knowledge graph techniques and fact-checking with external knowledge. This study contributes a systematic synthesis of open-domain LLM hallucination detection methods that do not rely on task-specific grounding documents, providing a structured taxonomy across hallucination type and technical access model, and distilling dominant approaches and evaluation gaps to guide future research.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large language models</kwd>
        <kwd>hallucination detection methods</kwd>
        <kwd>systematic literature review</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Large Language Models are capable of generating sophisticated, coherent, and contextually relevant
text across a wide range of applications. However, a significant weakness of these models is
hallucination, a phenomenon in which the model generates factually incorrect, unfaithful, or nonsensical
information that is not grounded in the provided source content or established world knowledge [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
These fabrications can range from minor inaccuracies to completely invented facts, posing a substantial
risk to user trust and model reliability, especially in domains where factual reliability is fundamental
(such as medicine, law and journalism).
      </p>
      <p>As the adoption of LLMs grows, the need to ensure their reliability and trustworthiness has become
a concern for both researchers and practitioners. This research addresses this challenge by conducting
a Systematic Literature Review (SLR) of methods designed to detect hallucinations in LLMs. The scope
is specifically focused on textual hallucinations, and concentrates exclusively on detection methods.
By mapping the existing landscape of detection techniques, this research aims to provide a clear and
structured overview of the current state of research.</p>
      <p>To guide this systematic review and ensure a comprehensive analysis, the following research questions
has been formulated: What hallucination detection methods have been proposed for Large
Language Models in the existing literature?</p>
      <p>
        The taxonomy proposed by Huang et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] guides the analysis of this research. In the study,
hallucinations are categorized into two main types: factuality hallucination and faithfulness hallucination.
Factuality hallucination occurs when the model generates content that either contradicts established
real-world knowledge or cannot be confirmed as accurate. Conversely, faithfulness hallucination refers
to the extent to which an LLM’s output remains aligned with the user’s instructions, the given context,
and maintains internal coherence.
      </p>
      <p>
        Prior research has examined hallucinations in large language models, though with diferent scopes.
Luo et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] review detection and mitigation methods, including token- and sentence-level approaches.
Their work is a general survey rather than a formal SLR and includes tasks beyond open-domain settings,
such as summarization and translation. It also predates most studies analyzed here. Huang et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
propose a taxonomy distinguishing factuality and faithfulness and discuss causes across data, training,
and inference stages. While comprehensive, their survey spans mitigation and benchmarks, whereas
this research focuses solely on detection.
      </p>
      <p>The paper is divided as follows: Section 2 describes the research method, Section 3 presents the key
ifndings of the research, Section 4 discusses the findings and Section 5 concludes the work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Research Method</title>
      <sec id="sec-2-1">
        <title>2.1. Search Strategy</title>
        <p>
          This research employs the systematic literature review (SLR) methodology proposed by Kitchenham
and Charters [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], which define three main phases: planning, conducting, and reporting the review.
The search was conducted with a metadata-focused search (title, abstract, keywords; or abstract where
needed) across four subscription-accessible digital libraries in computer science and information
technology: ACM Digital Library, Computer Science Database (ProQuest), IEEE Xplore (IEL), and Scopus
(Elsevier). Searches were executed on October 7, 2025. The search string used was:
("language model*" OR LLM* OR "generative AI" OR "foundation model*" OR "
transformer model*" OR "generative model*") AND ("hallucination* detection"
OR "hallucination* identification" OR "hallucination* evaluation" OR
"fact-checking" OR "fact checking")
        </p>
        <p>Inclusion and exclusion criteria were defined broad enough to capture open-domain detection research
while excluding domain-bound or grounding-dependent approaches. The complete list is in Table 1.</p>
        <p>
          Following the protocol, screening proceeded in three stages: (i) title screening removed clearly
irrelevant entries; (ii) abstract screening excluded works outside the open-domain detection scope; (iii)
quality assessment excluded studies failing the predefined threshold. Four otherwise relevant papers
were inaccessible via institutional subscriptions and could not be obtained from authors. In total, 748
studies were retrieved, of which 50 studies were included in the final corpus. The selected studies are
listed in the references as items [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]-[55] and are explicitly marked with paper IDs in the form L* to
enable traceability to the raw data recorded in the spreadsheet.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Data Extraction</title>
        <p>To support transparency and consistent synthesis, a predefined data-extraction form was used. For each
included study, the following information was recorded: bibliographic identifiers, hallucination type
addressed, detection granularity (token/sentence/passage), access to model internals, dependence on
external references, method classification and main technique, a brief method summary, detector output,
evaluation setup and evaluation metrics. Data were extracted in chronological order and consolidated
into a common spreadsheet for analysis, which is publicly available on GitHub1.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>This section answers the research question: what hallucination detection methods have been proposed
for Large Language Models in the existing literature?</p>
      <p>
        Following the taxonomy proposed by Huang et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], hallucination detection strategies are first
divided into two fundamental categories: Factuality hallucination detection, which identifies factual
errors in outputs, and faithfulness hallucination detection, which identifies the faithfulness of outputs
to the provided context.
      </p>
      <p>
        The majority of the studies, 37 in total, proposed methods for factuality hallucination detection,
while 12 focused on faithfulness hallucination detection. In some cases, the distinction between the two
categories was not clear-cut, especially with methods that had application in diferent domains and
diferent phases. One study [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] clearly addressed both strategies.
      </p>
      <p>37
12
1</p>
      <p>Faithfulness hallucination detection
Factuality hallucination detection</p>
      <p>Both</p>
      <p>Despite this primary classification, another important technical distinction between methods was
found: 16 methods needs to access LLM internals or token probabilities (white-box), while 34 methods
34
16</p>
      <p>Method accesses LLM internals</p>
      <p>Method doesn’t access LLM internals
relies solely on the LLM’s outputs (black-box).</p>
      <sec id="sec-3-1">
        <title>3.1. White-box hallucination detection methods</title>
        <p>
          White-box approaches, represented by 16 studies, leverage internal signals such as hidden activations,
attention patterns, gradients, or token-level probabilities , including logits and entropy. These signals
support finer-grained detection at the token or claim level and often enable lightweight inference
without external calls. Some methods combine multiple internal cues , such as embeddings and entropy,
to improve robustness. Detection granularity ranges from token-level ([
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]) to sentence/passage-level
([
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]), with some methods supporting both ([
          <xref ref-type="bibr" rid="ref20">20</xref>
          ],
[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]). All surveyed white-box approaches are reference-free and do not rely on external retrieval, making
them applicable to tasks centered on internal model confidence and uncertainty.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Black-box hallucination detection methods</title>
        <p>Black-box methods (34 studies) do not rely on internals and instead analyze generated text and,
optionally, external evidence. Two broad families emerge: (1) Knowledge-based methods that retrieve
external information, such as web, Wikipedia/Wikidata, and Resource Description Framework (RDF)
knowledge graphs to fact-check outputs. RDF is a standard data model for representing information as
structured ’triples’ (Subject-Predicate-Object), which allows for precise automated verification; and
(2) Reference-free methods that judge reliability using the model’s own behavior, auxiliary LLMs,
structured representations, or consistency signals without retrieval.</p>
        <sec id="sec-3-2-1">
          <title>3.2.1. Knowledge-based methods</title>
          <p>
            The 11 knowledge-based methods target factuality hallucinations and generally follow a fact-checking
paradigm with external retrieval. These methods are in studies [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ], [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ], [
            <xref ref-type="bibr" rid="ref23">23</xref>
            ], [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ], [
            <xref ref-type="bibr" rid="ref25">25</xref>
            ], [
            <xref ref-type="bibr" rid="ref26">26</xref>
            ], [
            <xref ref-type="bibr" rid="ref27">27</xref>
            ], [
            <xref ref-type="bibr" rid="ref28">28</xref>
            ],
[
            <xref ref-type="bibr" rid="ref29">29</xref>
            ], [
            <xref ref-type="bibr" rid="ref30">30</xref>
            ] and [
            <xref ref-type="bibr" rid="ref31">31</xref>
            ]. Knowledge is sourced primarily from Wikipedia/Wikidata or web search, while
some approaches leverage RDF-based knowledge graphs, such as DBpedia, LODsyndesis, OpenDialKG.
          </p>
          <p>Two main phases dominate these methods: fact extraction and verification. Extraction typically
decomposes LLM outputs into atomic claims or constructs knowledge graphs, enabling finer-grained
validation. Verification then matches these units against retrieved evidence using semantic entailment
or graph-based techniques. NLI classifiers and web search are common for textual claims, whereas
KG-based approaches employ structural matching and embedding-based similarity.</p>
          <p>Although most methods adhere to this pipeline, some integrate alternative architectures combining
multi-source retrieval, fusion, and decision-making without explicit decomposition. These systems often
incorporate specialized modules for evidence aggregation and verdict generation, aiming to improve
robustness in detecting and mitigating hallucinations.
1https://github.com/Eelsie/master_thesis_data/blob/main/extraction_forms.csv</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2.2. Knowledge graph: knowledge-based and reference-free methods</title>
          <p>
            As shown in Table 3, seven studies propose methods based on knowledge graphs as the primary
technique: [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ], [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ], [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ], [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ], [
            <xref ref-type="bibr" rid="ref33">33</xref>
            ], [
            <xref ref-type="bibr" rid="ref34">34</xref>
            ], and [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ]. Among these, [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ], [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ] and [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ] rely on external
knowledge in the form of knowledge bases that utilize the RDF data model, such as DBpedia,
LODsyndesis, and OpenDialKG. These methods verify model-generated facts against established knowledge
bases. They transform LLM outputs into triplets (typically subject–predicate–object) statements and
cross-check these against curated or dynamically retrieved graphs.
          </p>
          <p>
            ConFcheKG [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ] links entities to multiple knowledge graphs and scores reliability by contrasting
intersecting versus conflicting subgraphs: low coherence signals hallucination. GPT LODS [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ] prompts
the LLM to produce RDF triples for a question, then checks those triples in real time against DBpedia
or LODsyndesis. [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ] difers because it proposes a method based on Graph Neural Networks: given
a dialogue history and its generated response, it extracts entities or relations from the response and
retrieves KB facts about those entities to form a reference. It then encodes both graphs with RGAT,
pools graph features, and performs binary classification.
          </p>
          <p>
            In contrast, the methods reported in [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ], [
            <xref ref-type="bibr" rid="ref33">33</xref>
            ], [
            <xref ref-type="bibr" rid="ref34">34</xref>
            ] and [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ] are knowledge-free, relying only on the
provided context or the model’s internal knowledge. FactAlign [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ] constructs graphs from the source
text and the output, aligns their triples, and flags low-alignment facts. GraphEval [
            <xref ref-type="bibr" rid="ref33">33</xref>
            ] also turns outputs
into triples but tests each against the given context, improving base NLI detectors while returning
the ofending triples. GCA [
            <xref ref-type="bibr" rid="ref34">34</xref>
            ] builds a graph for each response, models dependencies between facts
with an RGCN, and combines multi-sample consistency with reverse verification to score hallucination.
Finally a semantic-graph uncertainty approach [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ] uses AMR graphs to propagate uncertainty across
entities and sentences and calibrates it with NLI-based contradiction signals.
          </p>
        </sec>
        <sec id="sec-3-2-3">
          <title>3.2.3. LLM-as-a-judge: knowledge-based and reference-free methods</title>
          <p>
            Seventeen methods use LLM-as-a-judge as their main technique or an important part of their method.
LLM-as-a-judge is the practice of using a large language model as an automatic evaluator of other
models’ outputs, typically by prompting it to compare or grade responses on open-ended tasks [
            <xref ref-type="bibr" rid="ref35">35</xref>
            ].
The detection granularity of the methods is at the sentence or passage level. In some methods, claims
are first extracted and verified in a second phase.
          </p>
          <p>Among these studies, six studies use external references to judge the veracity of the outputs. These
methods decompose text into claims, retrieve supporting evidence from the web or knowledge bases,
and use an LLM to verify the consistency between each claim and the retrieved evidence.</p>
          <p>The remaining eleven studies rely exclusively on the model’s internal knowledge and adopt four
distinct strategies:
• Self-consistency methods treat hallucinations as instability in the model’s own behavior: if the
model doesn’t give consistent answers when asked the same question multiple times under
slightly diferent conditions, it’s likely hallucinated.
• Metamorphic-testing methods also treat hallucinations as instability in the model’s own behavior
by applying systematic transformations (metamorphic relations) to the input and checking
whether the output changes in a predictable way.
• LLM-judge classifiers and ensembles methods, where the LLM is prompted as a classifier with no
external retrieval: reliability comes from prompt design, few-shot examples and ensemble voting.
• LLM to train data methods don’t want to call a LLM at inference. Instead, they use a strong
LLM as a teacher to label data (or generate reliability signals) and then train a smaller model or
ensemble of models.</p>
        </sec>
        <sec id="sec-3-2-4">
          <title>3.2.4. Other techniques for knowledge-free black-box methods</title>
          <p>Beyond LLM-as-a-judge and knowledge graphs, the selected studies present several additional
methodologies for hallucination detection that operate without access to model internals or external knowledge
sources. These approaches, nine in total, can be categorized into four groups.</p>
          <p>• Reverse validation method in [47] and [48] assesses reliability by reconstructing the original
query from the generated answer; hallucinations are indicated by reconstruction mismatches.
For example, InterrogateLLM [47] draws on human interrogation techniques, using consistency
across repeated questioning as an indicator of truthfulness.
• Uncertainty estimation method detects hallucinations by quantifying uncertainty in generated
text. For example, [49] paraphrases query into multiple scenarios and applies factor analysis
to separate shared semantic uncertainty from scenario-specific variation, while [ 50] computes
token-level uncertainty using negative log-likelihood and entropy over informative keywords.
• Natural Language Inference (NLI) method casts detection as entailment or contradiction between
model outputs and task-specific textual anchors.[ 51] performs zero-shot detection by checking
entailment between source, hypothesis, and target using pretrained NLI models, while [52]
evaluates internal consistency by decomposing responses into claims and measuring support or
contradiction across multiple sampled answers.
• Synthetic data method is to generate faithful/hallucinated pairs [45] or weak labels [53] with
LLMs, then fine-tune discriminative detectors such as DeBERTa/RoBERTa ([ 53], [54]) or LoRA
adapters [54].</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
      <p>This research systematically reviewed hallucination detection methods for LLMs in open-domain
settings, where no grounding documents are available. The review analyzed 50 peer-reviewed studies
published between 2023 and 2025. The research landscape is recent and rapidly evolving: no studies
were found prior to 2022, and 2024 emerged as the most prolific year, reflecting the surge in interest
following widespread LLM adoption.</p>
      <p>
        Majority of detection methods addressing factuality. Following the taxonomy of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the research
revealed a prominence of factuality hallucination detection (37 studies) over faithfulness detection (12
studies). This imbalance suggests that the primary concern in open-domain generation is the fabrication
of world knowledge (entity errors, factual fabrications) rather than internal logical consistency. This
focus is also intuitive given the open-domain nature of the selected studies: without grounding documents
(as used in RAG, translation or summarization), the truth must be derived from the model’s internal
parametric knowledge or from external verification. However, the distinction is becoming blurred in
knowledge graph -based approaches. These methods transform text into knowledge graphs or semantic
graphs, treating factuality and faithfulness as the same mathematical problem: graph alignment or
consistency.
      </p>
      <p>Predominance of black-box methods. A significant finding of this review, is the majority of
blackbox methods (34 studies) over white-box methods (16 studies). This trend is probably a consequence of
the commercial reality of the LLM landscape: as the most capable proprietary models (GPT, Claude,
Gemini) are accessible primarily via API with no access to internal weights, logits, attention maps,
gradients or hidden activations, researchers have been forced to innovate outside the model architecture.
White-box methods generally ofer finer granularity (token-level detection), but their detection is limited
to open-source models (for example, LLaMA and Mistral). By accessing the model’s uncertainty directly,
these methods also avoid the latency and computational expense of generating multiple external outputs.
More capable proprietary models require also more computationally expensive (black-box) detection
methods.</p>
      <p>LLM-as-a-judge is the most used detection technique. Seventeen methods use LLM-as-a-judge
approaches, where one language model evaluates the outputs of another. This evaluation strategy reflects
both the capabilities and limitations of current AI systems. On one hand, LLMs possess the linguistic
sophistication and reasoning ability to make nuanced judgments about factuality and consistency; on
the other hand, LLM-as-a-judge methods inherit the very vulnerabilities they aim to detect. Evaluator
models may themselves hallucinate, exhibit biases, or demonstrate inconsistent judgment, introducing
the critical risk of recursive hallucinations. This phenomenon occurs when the evaluator model, in
the process of assessing another model’s output, generates its own hallucinations. If the judge model
lacks the necessary world knowledge or shares the same inductive biases as the target model, it may
incorrectly validate a hallucinated claim (a false negative) or flag a correct statement as an error (a false
positive). The literature attempts to mitigate this in diferent ways: grounding the judge in external
evidence, querying multiple judges and using a majority vote, using self-consistency (ask the judge
the same question under paraphrases or perturbations), combining LLM judge with other non LLM
judgments or using LLM judges only to create supervision for separate detectors.</p>
      <sec id="sec-4-1">
        <title>External knowledge detection in black-box methods. Among black-box methods, 11 studies</title>
        <p>rely on external references. External-knowledge methods combines fact extraction into atomic claims
or triples and verification against web search, Wikipedia/Wikidata, or RDF knowledge graphs. These
methods are particularly suited to domains where the relevant information is well covered by public
knowledge bases. However, they are vulnerable to the same knowledge boundaries that afect the
underlying LLMs: long-tail, up-to-date, or copyright-restricted knowledge remains dificult to verify.
They also introduce engineering complexity and runtime cost due to retrieval and reasoning over
external sources.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Reference-free detection in black-box methods. Zero-knowledge detection is attractive because</title>
        <p>it avoids dependencies on external infrastructure and can be applied in settings where retrieval is
unavailable, unreliable, or undesirable. LLM-as-a-judge in this category is the most used technique (11
methods), along with knowledge graphs (4 methods). Moreover, other interesting minority approaches
emerged: reverse validation, uncertainty estimation, Natural Language Inference and generation of
synthetic data.</p>
        <p>
          Broad diversity in evaluation benchmarks. Studies rely on a variety of benchmarks, many of which
difer substantially in task formulation, annotation protocol, size, and granularity. Some detectors are
evaluated primarily on question answering datasets (TruthfulQA [40][
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], FreshQA [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ][40], TriviaQA
[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ][50] and SQuAD [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ][
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]), others on summarisation corpora (xSum [49][
          <xref ref-type="bibr" rid="ref32">32</xref>
          ], SummEval [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]), and
others on dedicated detection datasets like HaluEval [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ][55][
          <xref ref-type="bibr" rid="ref26">26</xref>
          ][
          <xref ref-type="bibr" rid="ref11">11</xref>
          ][
          <xref ref-type="bibr" rid="ref27">27</xref>
          ][
          <xref ref-type="bibr" rid="ref29">29</xref>
          ][
          <xref ref-type="bibr" rid="ref30">30</xref>
          ][
          <xref ref-type="bibr" rid="ref31">31</xref>
          ], SelfCheckGPT
[49][
          <xref ref-type="bibr" rid="ref23">23</xref>
          ][
          <xref ref-type="bibr" rid="ref8">8</xref>
          ][55][
          <xref ref-type="bibr" rid="ref12">12</xref>
          ][40][
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] or SHROOM [51][53][41][46][44]. This diversity complicates systematic
comparison across methods and reported performance is often tightly coupled to the characteristics of
a particular benchmark. Further research could benefit from standardized benchmarks and evaluation
protocols.
        </p>
        <p>Threats to validity. Several limitations should be considered when interpreting these findings. The
review was conducted by the first author, with guidance from the other two authors. Although the
literature search, study selection, and data extraction were discussed among the authors, the risk of
subjectivity remains, particularly in study selection, quality assessment, and data extraction. Involving
multiple reviewers would likely have improved reliability. The exclusion of gray literature ensures a
focus on peer-reviewed quality but may omit relevant preprints. Additionally, four potentially suitable
papers were inaccessible despite attempts to contact the authors. Determining whether studies addressed
open-domain hallucination detection was occasionally ambiguous, potentially introducing selection
bias. Finally, categorizing methods as targeting factuality or faithfulness required interpretive judgment,
particularly for multi-technique and multi-granularity approaches such as knowledge-graph–based
methods. While predefined forms and established taxonomies helped mitigate misclassification risks,
some ambiguity remains.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>This study provides a comprehensive overview of hallucination detection methods for LLMs in
opendomain settings. The findings reveal a rapidly growing and diverse research field, with a strong emphasis
on factuality detection, a predominance of black-box approaches, and widespread reliance on
LLM-as-ajudge techniques. The findings reveal both progress and persistent challenges. Detection methods are
diverse, intending to resolve the balance between scalability, accuracy, and independence from external
resources. White-box approaches ofer more precision but lack applicability to closed models. Black-box
methods are broadly deployable but often computationally expensive or reliant on imperfect judges.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author(s) used OpenAI ChatGPT (GPT-4 and GPT-5 models),
Anthropic Claude (Sonnet 4.5) in order to: Improve the use of spelling and grammar throughout the
text, synthesize or paraphrase complex concepts for comparison with own understanding.
(2024). doi:10.18653/v1/2024.semeval-1.52, L26.
[39] W. Wu, Y. Cao, N. Yi, R. Ou, Z. Zheng, Detecting and reducing the factual hallucinations of large
language models with metamorphic testing, Proc. ACM Softw. Eng. (2025). doi:10.1145/3715784,
L51.
[40] B. Yang, M. A. Al Mamun, J. M. Zhang, G. Uddin, Hallucination detection in large language models
with metamorphic relations, Proc. ACM Softw. Eng. (2025). doi:10.1145/3715735, L56.
[41] R. Sanayei, A. Singh, M. Rezaei, S. Bethard, Maria at semeval 2024 task-6: Hallucination detection
through llms, mnli, and cosine similarity (2024). doi:10.18653/v1/2024.semeval-1.225, L36.
[42] A. Bui, S. Brech, N. Hußfeldt, T. Jennert, M. Ullrich, T. Breuer, N. Khasmakhi, P. Schaer, The two
sides of the coin: Hallucination generation and detection with llms as evaluators for llms (2024).</p>
      <p>L46.
[43] S. Das, R. Śrihari, Compos mentis at semeval2024 task6: A multi-faceted role-based large language
model ensemble to detect hallucination (2024). doi:10.18653/v1/2024.semeval-1.208, L10.
[44] B. Allen, F. Polat, P. Groth, Shroom-indelab at semeval-2024 task 6: Zero- and few-shot llm-based
classification for hallucination detection (2024). doi: 10.18653/v1/2024.semeval-1.120, L44.
[45] Y. Chen, Q. Fu, Y. Yuan, Z. Wen, G. Fan, D. Liu, D. Zhang, Z. Li, Y. Xiao, Hallucination detection:
Robustly discerning reliable answers in large language models (2023). doi:10.1145/3583780.
3614905, L7.
[46] C. Wei, Z. Chen, S. Fang, J. He, M. Gao, Opdai at semeval-2024 task 6: Small llms can accelerate
hallucination detection with weakly supervised data (2024). doi:10.18653/v1/2024.semeval-1.
104, L40.
[47] Y. Yehuda, I. Malkiel, O. Barkan, J. Weill, R. Ronen, N. Koenigstein, Interrogatellm: Zero-resource
hallucination detection in llm-generated answers (2024). doi:10.18653/v1/2024.acl-long.
506, L30.
[48] S. Yang, R. Sun, X. Wan, A new benchmark and reverse validation method for passage-level
hallucination detection (2023). doi:10.18653/v1/2023.findings-emnlp.256, L2.
[49] T. Zhang, L. Qiu, Q. Guo, C. Deng, Y. Zhang, Z. Zhang, C. Zhou, X. Wang, L. Fu, Enhancing
uncertainty-based hallucination detection with stronger focus (2023). doi:10.18653/v1/2023.
emnlp-main.58, L5.
[50] Z. Wen, Z. Liu, Z. Tian, S. Pan, Z. Huang, D. Li, M. Huang, Scenario-independent uncertainty
estimation for llm-based question answering via factor analysis (2025). doi:10.1145/3696410.
3714880, L64.
[51] P. Bhamidipati, A. Malladi, M. Shrivastava, R. Mamidi, Maha bhaashya at semeval-2024 task 6:
Zero-shot multi-task hallucination detection (2024). doi:10.18653/v1/2024.semeval-1.241,
L34.
[52] F. Cheng, V. Zouhar, S. Arora, M. Sachan, H. Strobelt, M. El-Assady, Relic: Investigating large
language model responses using self-consistency (2024). doi:10.1145/3613904.3641904, L42.
[53] F. Borra, C. Savelli, G. Rosso, A. Koudounas, F. Giobergia, Malto at semeval-2024 task 6: Leveraging
synthetic data for llm hallucination detection (2024). doi:10.18653/v1/2024.semeval-1.240,
L35.
[54] J. Lu, S. Li, Roberta with low-rank adaptation and hierarchical attention for hallucination detection
in llms, 2024 International Conference on Image Processing, Computer Vision and Machine
Learning (ICICML) (2024). doi:10.1109/ICICML63543.2024.10957858, L43.
[55] D. Zhang, V. Gangal, B. Lattimer, Y. Yang, Enhancing hallucination detection through
perturbation-based synthetic data generation in system responses (2024). doi:10.18653/v1/
2024.findings-acl.789, L17.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Frieske</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Tiezheng Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ishii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. S.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Madotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fung</surname>
          </string-name>
          ,
          <article-title>Survey of Hallucination in Natural Language Generation</article-title>
          , ACM Computing Surveys,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yu</surname>
          </string-name>
          , W. Ma,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Qin</surname>
          </string-name>
          , T. Liu,
          <article-title>A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions</article-title>
          ,
          <source>ACM Trans. Inf. Syst</source>
          . (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jenkin</surname>
          </string-name>
          , S. Liu, G. Dudek,
          <article-title>Hallucination detection and hallucination mitigation: An investigation</article-title>
          ,
          <source>arXiv preprint arXiv:2401.08358</source>
          (
          <year>2024</year>
          ). URL: https://arxiv.org/abs/ 2401.08358.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kitchenham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Charters</surname>
          </string-name>
          ,
          <article-title>Guidelines for performing systematic literature reviews in software engineering</article-title>
          , Keele University and Durham University Joint Report (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <article-title>Enhancing uncertainty modeling with semantic graph for hallucination detection (</article-title>
          <year>2025</year>
          ). doi:
          <volume>10</volume>
          .1609/aaai.v39i22.
          <volume>34528</volume>
          ,
          <issue>L52</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Suresh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aljundi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Nkisi-Orji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wiratunga</surname>
          </string-name>
          ,
          <article-title>Towards improving open-box hallucination detection in large language models (llms) (</article-title>
          <year>2024</year>
          ).
          <year>L47</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Hademif:
          <article-title>Hallucination detection and mitigation in large language models (</article-title>
          <year>2025</year>
          ).
          <year>L55</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C. Wu, G. Chen,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Embedding and gradient say wrong: A white-box method for hallucination detection (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .emnlp-main.
          <volume>116</volume>
          ,
          <issue>L16</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Haloscope: Harnessing unlabeled llm generations for hallucination detection (</article-title>
          <year>2024</year>
          ).
          <year>L25</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <article-title>Inside: Llms' internal states retain the power of hallucination detection (</article-title>
          <year>2024</year>
          ).
          <year>L28</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Beigi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jin</surname>
          </string-name>
          , C.-
          <string-name>
            <surname>T.</surname>
            Chang-Tien,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
          </string-name>
          , Internalinspector i2:
          <article-title>Robust confidence estimation in llms through internal states (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          . 18653/v1/
          <year>2024</year>
          .findings-emnlp.
          <volume>751</volume>
          ,
          <issue>L29</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sriramanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bharti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sadasivan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kattakinda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Feizi</surname>
          </string-name>
          , Llm-check:
          <article-title>Investigating detection of hallucinations in large language models (</article-title>
          <year>2024</year>
          ).
          <year>L32</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Y.-S.</given-names>
            <surname>Chuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Qiu</surname>
          </string-name>
          , C.-Y. Hsieh,
          <string-name>
            <given-names>R.</given-names>
            <surname>Krishna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Glass</surname>
          </string-name>
          ,
          <article-title>Lookback lens: Detecting and mitigating contextual hallucinations in large language models using only attention maps (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .emnlp-main.
          <volume>84</volume>
          ,
          <issue>L33</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>W.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , Y. Liu,
          <article-title>Unsupervised real-time hallucination detection based on the internal states of large language models (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          . findings-acl.
          <volume>854</volume>
          ,
          <issue>L49</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Joo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-J.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <article-title>Entropy-based sentence-level hallucination score in large language models (</article-title>
          <year>2025</year>
          ).
          <source>doi:10.1109/BigComp64353</source>
          .
          <year>2025</year>
          .
          <volume>00022</volume>
          ,
          <issue>L53</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>R. B. Beyene</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Faghih</surname>
            ,
            <given-names>T. A.</given-names>
          </string-name>
          ,
          <article-title>Hallucination detection in llms via beam search sampling and semantic consistency analysis</article-title>
          ,
          <source>2025 55th Annual IEEE/IFIP International Conference on Dependable Systems</source>
          and Networks
          <string-name>
            <surname>Workshops (DSN-W)</surname>
          </string-name>
          (
          <year>2025</year>
          ). doi:
          <volume>10</volume>
          .1109/DSN-W65791.
          <year>2025</year>
          .
          <volume>00076</volume>
          ,
          <issue>L57</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.</given-names>
            <surname>Arteaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schön</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pielawski</surname>
          </string-name>
          ,
          <article-title>Hallucination detection in llms: Fast and memory-eficient ifne-tuned models (</article-title>
          <year>2025</year>
          ).
          <year>L58</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ciosek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Felicioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghiassian</surname>
          </string-name>
          ,
          <article-title>Hallucination detection on a budget: Eficient bayesian estimation of semantic entropy</article-title>
          ,
          <source>Transactions on Machine Learning Research</source>
          (
          <year>2025</year>
          ).
          <year>L59</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Huo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Mixhd: A method for detecting hallucinations based on the internal state and output probability of large language models</article-title>
          ,
          <source>ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          (
          <year>2025</year>
          ).
          <source>doi:10.1109/ICASSP49660</source>
          .
          <year>2025</year>
          .
          <volume>10889328</volume>
          ,
          <issue>L61</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Fadeeva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rubashevskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petrakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mubarak</surname>
          </string-name>
          , E. Tsymbalov,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kuzmin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          , T. Baldwin,
          <article-title>Fact-checking the output of large language models via token-level uncertainty quantification (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .findings-acl.
          <volume>558</volume>
          ,
          <issue>L18</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sohn</surname>
          </string-name>
          ,
          <article-title>Context-based fact-checking using knowledge graph (</article-title>
          <year>2023</year>
          ).
          <source>doi:10.1109/ BigData59044</source>
          .
          <year>2023</year>
          .
          <volume>10386121</volume>
          ,
          <issue>L3</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Furumai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shinohara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ikeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kato</surname>
          </string-name>
          ,
          <article-title>Detecting dialogue hallucination using graph neural networks</article-title>
          ,
          <source>2023 International Conference on Machine Learning and Applications (ICMLA)</source>
          (
          <year>2023</year>
          ).
          <source>doi:10.1109/ICMLA58977</source>
          .
          <year>2023</year>
          .
          <volume>00128</volume>
          ,
          <issue>L4</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Hallucination detection for generative large language models by bayesian sequential estimation (</article-title>
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .emnlp-main.
          <volume>949</volume>
          ,
          <issue>L6</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mountantonakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tzitzikas</surname>
          </string-name>
          ,
          <article-title>Real-time validation of chatgpt facts using rdf knowledge graphs (</article-title>
          <year>2023</year>
          ).
          <year>L11</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , H. Chen, Factchd:
          <article-title>Benchmarking fact-conflicting hallucination detection (</article-title>
          <year>2024</year>
          ).
          <year>L20</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Reddy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Mujahid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rubashevskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Borenstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pillai</surname>
          </string-name>
          , Factcheck-bench:
          <article-title>Fine-grained evaluation benchmark for automatic fact-checkers (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .findings-emnlp.
          <volume>830</volume>
          ,
          <issue>L21</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Medico:
          <article-title>Towards hallucination detection and correction with multi-source evidence fusion (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          . emnlp-demo.
          <volume>4</volume>
          ,
          <issue>L37</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>H.</given-names>
            <surname>Iqbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Georgiev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geng</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <article-title>Openfactcheck: A unified framework for factuality evaluation of llms (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .emnlp-demo.
          <volume>23</volume>
          ,
          <issue>L41</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>X.</given-names>
            <surname>Cheng</surname>
          </string-name>
          , J.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Gai</surname>
            ,
            <given-names>J.-R.</given-names>
          </string-name>
          <string-name>
            <surname>Wen</surname>
          </string-name>
          ,
          <article-title>Small agent can also rock! empowering small language models as hallucination detector (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .emnlp-main.
          <volume>809</volume>
          ,
          <issue>L45</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>S.</given-names>
            <surname>Heo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Son</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Park</surname>
          </string-name>
          , Halucheck:
          <article-title>Explainable and verifiable automation for detecting hallucinations in llm responses</article-title>
          ,
          <source>Expert Systems with Applications</source>
          (
          <year>2025</year>
          ). doi:
          <volume>10</volume>
          .1016/j.eswa.
          <year>2025</year>
          .
          <volume>126712</volume>
          ,
          <issue>L60</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Towards detecting llms hallucination via markov chainbased multi-agent debate framework</article-title>
          ,
          <source>ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          (
          <year>2025</year>
          ).
          <source>doi:10.1109/ICASSP49660</source>
          .
          <year>2025</year>
          .
          <volume>10889448</volume>
          ,
          <issue>L63</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rashad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zahran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Amin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdelaal</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>AlTantawy, Factalign: Fact-level hallucination detection and classification through knowledge graph alignment (</article-title>
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          . trustnlp-
          <volume>1</volume>
          .8,
          <issue>L19</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>H.</given-names>
            <surname>Sansford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Richardson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Maretić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Saada</surname>
          </string-name>
          ,
          <article-title>Grapheval: A knowledge-graph based llm hallucination evaluation framework (</article-title>
          <year>2024</year>
          ).
          <year>L22</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>X.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Zero-resource hallucination detection for text generation via graph-based contextual knowledge triples modeling (</article-title>
          <year>2025</year>
          ). doi:
          <volume>10</volume>
          .1609/aaai.v39i22.
          <volume>34559</volume>
          ,
          <issue>L50</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , W.-L. Chiang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Stoica</surname>
          </string-name>
          ,
          <article-title>Judging LLM-as-a-judge with MT-bench and chatbot arena</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          <volume>36</volume>
          (NeurIPS
          <year>2023</year>
          ), Datasets and
          <string-name>
            <given-names>Benchmarks</given-names>
            <surname>Track</surname>
          </string-name>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. Das</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Malin</surname>
          </string-name>
          , S. Sricharan,
          <article-title>Sac3: Reliable hallucination detection in black-box language models via semantic-aware cross-check consistency (</article-title>
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          . findings-emnlp.
          <volume>1032</volume>
          ,
          <issue>L12</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>P.</given-names>
            <surname>Manakul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Liusie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gales</surname>
          </string-name>
          , Shroomshroom:
          <article-title>Zero-resource black-box hallucination detection for generative large language models (</article-title>
          <year>2023</year>
          ).
          <year>L13</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hoblitzell</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. O'Keefe</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Jang</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Varma</surname>
          </string-name>
          ,
          <article-title>Halu-nlp at semeval-2024 task 6: Metacheckgpt - a multi-task hallucination detection using llm uncertainty and meta-models</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>