<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Nara, Japan
$ carlo.bono@polimi.it (C. A. Bono)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Eficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Carlo Alberto Bono</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federico Belotti</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Palmonari</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano, DEIB</institution>
          ,
          <addr-line>Via Ponzio 34/5, Milano, 20133</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Università degli Studi di Milano Bicocca, DISCo</institution>
          ,
          <addr-line>Viale Sarca, 336, 20126, Milano</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Linking textual values in tabular data to their corresponding entities in a Knowledge Base is a core task across a variety of data integration and enrichment applications. Although Large Language Models (LLMs) have shown State-of-The-Art performance in Entity Linking (EL) tasks, their deployment in real-world scenarios requires not only accurate predictions but also reliable uncertainty estimates, which require resource-demanding multi-shot inference, posing serious limits to their actual applicability. As a more eficient alternative, we investigate a self-supervised approach for estimating uncertainty from single-shot LLM outputs using token-level features, reducing the need for multiple generations. Evaluation is performed on an EL task on tabular data across multiple LLMs, showing that the resulting uncertainty estimates are highly efective in detecting low-accuracy outputs. This is achieved at a fraction of the computational cost, ultimately supporting a cost-efective integration of uncertainty measures into LLM-based EL workflows. The method ofers a practical way to incorporate uncertainty estimation into EL workflows with limited computational overhead.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Entity Linking</kwd>
        <kwd>Uncertainty</kwd>
        <kwd>LLMs</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>of confidence scores enables the identification of potential errors by prioritizing links that humans could
review to improve the quality of the results. Understanding and quantifying the uncertainty of the LLM
output is essential to ensure the robustness and trustworthiness of the model output, to highlight where
the model is fragile or ungrounded, and to direct human intervention in order to maximize its eficacy.</p>
      <p>An established way of measuring uncertainty in LLMs is to determine the likelihood of an answer being
consistent over multiple independent generations [13, 14, 15]. In the context of LLMs, this introduces
significant computational overhead, as the generation time scales linearly with the number of tokens in both the
prompt and the output, even with caching techniques [16]. Additionally, if we consider that EL may be
applied to large datasets, this approach may be unsustainable. To mitigate the resource demands of large-scale
EL scenarios, we propose an eficient approach to estimate uncertainty in LLMs without relying on multiple
generations during inference. More precisely, we propose an eficient self-supervised method that learns
to estimate the uncertainty observed over multiple generations of an LLM using observables from a single
generation. The method leverages token-level features, derived from the probability distribution over
the output vocabulary, to train a lightweight regression model targeting the “true” observed uncertainty.</p>
      <p>Although in this paper we focus on EL on tabular data, we believe that our method could be applied
to diferent closed-form tasks that can benefit from uncertainty quantification.</p>
      <p>This paper makes the following contributions:
• Formalizes uncertainty-aware EL on tabular data and proposes a self-supervised regressor that
learns to approximate multi-shot uncertainty from single-shot token-level features;
• Introduces a lightweight, model-agnostic feature set from output-layer probabilities and optional
intermediate-layer signals, requiring no task labels;
• Evaluates the approach across several instruction-tuned LLMs, showing how
uncertaintyawareness can efectively lead to the detection of low-accuracy outputs, which can then be
corrected to improve accuracy under a constrained review budget.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Uncertainty in EL for Tabular Data. Associating confidence scores in matching tasks, including EL for
tabular data, is a well-established practice [17, 18, 4], where scoring a list of retrieved candidates is a typical
intermediate step. This is also true in approaches combining a pretrained language model with classifiers,
such as TURL [10] and UNICORN [19]. While scores computed in pre-LLMs methods somehow support
confidence estimation, less attention has been dedicated to systematic analyses of uncertainty estimation
in this task. One approach based on a deep neural network proposes a supervised method that considers
the matching score of the best candidate and its distance from the second-best to quantify confidence [ 20].
The authors show that computed confidence scores help prioritize links to revise with a progressive budget.
To the best of our knowledge, recent approaches based on LLMs [7, 8, 9, 3, 5] have not investigated how to
exploit uncertainty measures in generative approaches. Uncertainty measures may support the decision
whether to link or not a top candidate (e.g., based on a threshold), including the detection of NIL entities
(i.e., values associated with entities not in the KB) [4]. Additionally, little attention has been dedicated to
the impact of the variability of the links predicted by LLMs under diferent generations within this task.
Uncertainty Estimation for LLMs. Confidence Score methods leverage single-shot proxies from
output probabilities (e.g., entropy, log-probabilities, perplexity) but can be overconfident when
wrong [21, 13, 22, 23, 24]; Semantic Consistency approaches, involving the generation of multiple outputs
for the same prompt and measuring their consistency or semantic similarity, are efective but
computationally expensive [22, 14]; Supervised methods learn calibrated uncertainty from features of generated text
or hidden states and often outperform unsupervised heuristics, but require an annotated dataset [25, 26];
Ensemble- and Bayesian-Inspired approaches estimate approximate Bayesian uncertainty (e.g., deep
ensembles, Monte-Carlo dropout) but are generally impractical at LLM scale [27, 28, 29]; Verbalized and
Self-Reported Uncertainty methods improve interpretability by prompting models to report confidence,
with mixed reliability across tasks [30, 31, 32]. Finally, [33, 34] employ Uncertainty-Aware In-Context
Learning to filter/refine or to guide iterative prompting, improving reliability on open-ended tasks.
This work.</p>
      <p>Our approach intersects multiple-generations and supervised paradigms by leveraging
a self-supervised uncertainty regressor that learns from multiple-generations outputs how to estimate
uncertainty from a single generation. Although this work shares some conceptual similarities with [26],
our method does not require supervised labels and focuses on tabular EL.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <sec id="sec-3-1">
        <title>3.1. Problem Formulation</title>
        <p>Let a tabular EL instance be defined by an input tuple (, , KG), where  is a table,  is a mention
at coordinates (, ), and KG = {1, . . . ,  } is a retriever-provided set of candidate entities, dependent
on the particular Knowledge Graph KG. An LLM  conditions on a prompt x built from (, , KG)
and generates a textual answer  from which a deterministic post-processor extracts a selected candidate
^() ∈ KG. For non-deterministic decoding, repeated generations {() 
}=1 induce an empirical
distribution ^ans(·| x) over answers3, from which an uncertainty (· ) can be estimated. Here, (· ) is the
uncertainty score derived from the empirical distribution of generations4. Our goal is to output, for each
x, both a selected candidate ^() and an uncertainty score ^() that correlates with (· ) computed a
posteriori from multiple generations. We learn this score via an uncertainty regressor ℎ from single-shot
token-level features (cf. Section 3.4). At deployment, we use ^() to flag instances for manual review.
Algorithm 1 outlines the complete training (warm-up) and inference workflow.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Measures of Uncertainty</title>
        <p>
          (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
To quantify the uncertainty (· ) of LLM outputs, we adopt two widely recognized measures: Predictive
Entropy (PE), which captures the uncertainty inherent to a model’s answer distribution, and Semantic
Entropy (SE), which considers semantic equivalence classes rather than raw answers for measuring
uncertainty [14]. PE is defined as the entropy of the observed output sequences , realizations of a
random variable  , conditional on an input sequence x, such that:
 (x) = (^ans(·| x)) = −
∑︁ ^ans(|x) log ^ans(|x)
where each class refers to a shared meaning:
Lower PE indicates that the output distribution is concentrated around a single answer, while higher
PE reflects uncertainty over multiple possible answers. SE is instead calculated over the distribution
of answer meanings ^sem(·| x)5, that is, over the semantic equivalence classes  observed on the output,
(x) = (^sem(·| x)) = −
∑︁ ^sem(|x) log ^sem(|x)
where (x) is the same as in [14]. To enable a comparison between entropy measures computed over
discrete distributions with varying support sizes, both measures are normalized by the logarithm of
the number of unique outputs. We also report the sequence perplexity (PP) as a confidence baseline.
For an answer  = (1, . . . ,  ) of  tokens,


  () = exp −  =1
︃(
1 ∑︁ log (︀ |&lt;, x)︀
)︃
where (|&lt;, x) is the model’s next-token probability.
4In Section 3.2 we instantiate it using the entropy of ^ans and, via semantic grouping [14], of ^sem.
3In particular, ^ans(|x) = 1 ∑︀
        </p>
        <p>=1 I[() = ], where I[] = 1 if the predicate  holds, and 0 otherwise.
5Let  :  →  map an answer  to its semantic class  = () [14]. The empirical distribution over semantic classes is
^sem(|x) = ∑︀
∈(x) ^ans(|x)I[() = ] = 1 ∑︀</p>
        <p>=1 I[(()) = ], where (x) is the set of observed distinct answers.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Uncertainty Estimation</title>
        <p>Our practical objective is to estimate uncertainty and use it to flag uncertain answers for manual review,
without relying on an explicit ground truth. The underlying intuition is that higher uncertainty is
correlated with lower accuracy, while low uncertainty alone does not imply higher accuracy. Accuracy
is calculated as the average answer correctness across  generated outputs for each input prompt.
Uncertainty–Accuracy Relationship. Figure 1 highlights the efect of setting a threshold on
uncertainty to flag items for manual review. Interpreting the flag decision as a binary classifier, the green
area contains correctly flagged low-accuracy cases (true positives). Reviewing these comes at the cost
of also reviewing cases in the yellow area (false positives), where accuracy would be satisfactory but
uncertainty is high. Cases in the red area are consistently wrong but with low variability; these are
false negatives and are not recoverable by uncertainty-based thresholding. Intuitively, lowering the
uncertainty threshold expands the region of flagged cases. On the other hand, raising the threshold
shrinks both green and yellow zones, reducing workload at the risk of missing additional bad cases.
Notably, the subset of red points with zero uncertainty, which are consistently wrong across generations,
remains invariant to the threshold and unrecoverable by uncertainty thresholding alone. Sweeping the
threshold traces the ROC analysis discussed later.</p>
        <p>Eficient Uncertainty Estimation via Self -Supervised Regression. A large class of uncertainty
estimation methods, including widely used metrics like PE and SE, depend on repeated sampling for
their computation, which comes with significant resource overhead. We therefore propose an eficient
self-supervised method to approximate such uncertainty measures using information from a single
generation, making them more suitable for downstream use. The target variable – either PE or SE,
measured on  independent runs of an LLM – is regressed from observable features using a Random
Forest model implemented via XGBRFRegressor, with 100 estimators [35, 36]. We optimize the Mean
Squared Error (MSE) between the predicted score ^() and the normalized multi-shot target (^ans)
(Section 3.2). We utilize 10-fold cross-validation grouped by prompt to avoid data leakage and to obtain
reliable performance estimates on unseen cases. The regressor is trained in a warm-up phase, during
which  generations per prompt are collected and the derived target variable is learned. At runtime, the
regressor estimates the uncertainty from a single generation in response to the same prompt. We denote
the learned uncertainty regressor by ℎ, which maps features  (, x) extracted from a target model
 ’s single generation  =  (x) to an estimate ^() = ℎ( (, x)). The method is self-supervised
as it does not require externally labeled data. Moreover, the regression incurs negligible computational
overhead. Further considerations regarding learning convergence are provided in Section 4.5.</p>
        <p>
          We collect the features, as defined in Section 3.4, from  independent runs of an LLM. We benchmark
the proposed method considering: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) the first 10 generated tokens, and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) all the tokens from the
ifxed-width tail of the prompt 6. This allows us to investigate the contribution of diferent prompt
segments to the quality of the uncertainty estimates. We also evaluate diferent feature sets as input,
as detailed in Section 4.2, to assess the relative importance of individual feature groups.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Features for Uncertainty Estimation</title>
        <p>We rely on features  (, x) that are observable or easily computable during the inference process of
an LLM  . We compute token-wise features for specific portions of the prompt and generated tokens.
Instead of aggregating the features at the sequence level [13], we maintain the features at the token level.
Let  be a vocabulary of tokens and let x be a sequence of tokens from  . Given the input prompt x,
an LLM computes a logit vector z =  (x) ∈ R| | for each token. It then applies a softmax function
to derive the stochastic vector p, i.e., the probability of  being the next token for every  ∈  :
exp(/ )
p ∈ R| | s.t.  =  (z) = ∑︀∈[| |] exp(/ )
where  is the temperature parameter, which controls the randomness of the output distribution. We
calculate for each token7:
• Max probability. The maximum probability observed across the output vocabulary, measured
at the output layer  of the model.
• Entropy. The entropy of the token probability distribution</p>
        <p>M(p) = m∈[|ax|] 
H(p) = −
∑︁  log 
∈[| |]
over all vocabulary tokens at the output layer . Lower entropy indicates higher confidence, while
higher entropy reflects more uncertainty.
• LogitLens [37]. The Kullback-Leibler (KL) divergence between the probability distribution at each
intermediate layer  and the probability distribution at the output layer  [37].</p>
        <p>
          KL(p ‖ p) =
∑︁  log
∈[| |]


where p is measured for each layer  ∈ [− 1] and p is measured on , resulting in − 1 divergence
measures quantifying the deviation between the distribution at the intermediate layers and the
output layer. At each layer , p =  (z) =  (︀ LayerNorm[h] )︀ is obtained by applying the last
LayerNorm to the hidden representation h and multiplying it by the unembedding matrix  .
6The token indices collected in (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) are fixed across models, whereas those in (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) are model-dependent.
7We use the notation [] to indicate {1, ..., }.
        </p>
        <sec id="sec-3-4-1">
          <title>Workflow overview for self-supervised uncertainty estimation Warm-up (ofline)</title>
          <p>1. For each prompt x, collect  independent generations from the LLM  to obtain {()}=1.
2. Compute the a posteriori uncertainty target (^ans/sem) either as normalized PE or SE
(Equations 1 and 2, respectively).
3. For each generation  ∈ [ ], extract per-generation features  ()(, x) using token-level
observables (Section 3.4) over selected segments (e.g., Postilla, first  Generated tokens).
4. Form training pairs (︀  ()(, x), (^ans/sem)︀) for  = 1, . . . ,  and for all prompts x; train
the regressor ℎ on the training pairs.</p>
        </sec>
        <sec id="sec-3-4-2">
          <title>Inference (online)</title>
          <p>1. For each new prompt x, run a single generation to produce an answer  =  (x) and extract
 (, x).</p>
          <p>2. Return  and predict ^() = ℎ( (, x)) and use it to decide if the item should be reviewed.</p>
        </sec>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Runtime Considerations</title>
        <p>At inference time, the proposed approach reduces the number of LLM generations used to estimate
uncertainty from  to 1 and adds only a lightweight regression pass on features extracted during
generation. In typical settings, the regression overhead is negligible compared to a single forward pass
of the LLM; thus, single-shot estimated PE/SE achieves most of the benefit of multi-shot uncertainty
at substantially lower cost. For a theoretical analysis of Transformer time complexity with and without
KV-cache during generation, see Appendix F.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Evaluation</title>
      <p>In this work, we address the following research questions:</p>
      <p>Q1 Do our single-shot uncertainty estimates identify low-accuracy answers?
Q2 How much does uncertainty-guided manual correction improve accuracy under a budget ?
Q3 How much warm-up data is needed to learn the uncertainty regressor ℎ efectively?
Q4 How does temperature  afect the uncertainty/accuracy trade-of?
Reproducible code and plotting scripts are available at: https://github.com/carloalbertobono/llm-u
ncertainty.</p>
      <sec id="sec-4-1">
        <title>4.1. Dataset</title>
        <p>We address the Entity Linking (EL) task on tabular data, formalized in Section 3.1, where each table
mention must be linked to its corresponding entity in a Knowledge Base. The dataset used,
TableInstructEL-2K, is adapted from the TableInstruct EL test set [5], which is formatted for LLMs and contains 2,000
mentions annotated with Wikidata entities, with exactly one correct entity per mention. The original
dataset included ∼ 600 mentions with only one candidate (i.e., the correct one), which limits our ability to
evaluate answer variability. To address this issue, the candidates were enriched with those retrieved using
LamAPI [38], a full-fledged retriever that returns richer candidate sets. In the resulting dataset, 1,650
mentions (∼ 91%) include at least 45 candidate entities each. Prompts follow the “Entity Linking” template
from [5] and include the following segments: an Instruction that provides context and task guidelines; an
Input table, Markdown-serialized; a Question asking which of the provided referent candidates corresponds
to a specific table mention; a fixed-width Postilla that clarifies the expected answer format. The Generated
segment contains the model’s predicted answer. To support understanding, a minimal, self-contained
illustrative example of the prompt, candidate list, and expected answer format is provided in Appendix A.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Experimental Settings</title>
        <p>Never correct
(w/o Unc.)</p>
        <p>Never correct
(w/ Unc.)</p>
        <p>Sometimes
correct</p>
        <p>Recoverable</p>
        <p>total
Experiments are performed using the following models: Gemma-2-2B-Instruct, Gemma-2-9B-Instruct,
Llama-3.1-8B-Instruct, TableLlama, Qwen2.5-7B-Instruct. These models were chosen as a representative
set of instruction-tuned, open-source language models. Additionally, TableLlama is considered
state-of-the-art in the Entity Linking task on tabular data [3]. The number of generations for each prompt
is set to  = 10, while the temperature  is set to 1.0. Features (observables) are extracted over diferent
portions of tokens, specifically from the Postilla and from the first 10 tokens of the Generated segment8,
considering features from the output layer alone (M(p) and H(p)), from the intermediate layers alone
(LogitLens), and combined. These combinations are motivated by the hypothesis that uncertainty may
manifest diferently across prompt segments and feature groups, necessitating an empirical assessment
of their impact. We compare the entropy values derived with our method against the following baselines:
PE, SE, PP, and an oracle with access to the true answers.</p>
        <p>In the remainder of the paper, we utilize two diferent concepts of accuracy. Answer-level
accuracy refers to the accuracy of the answers to a given prompt, computed over multiple generations.
Answer-level accuracy can then be aggregated at the level of a set of prompts; we refer to dataset-level
accuracy to indicate the average answer-level accuracy computed over the whole dataset.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Task Accuracy and Recoverable Errors</title>
        <p>Regarding the performance of the diferent LLMs on the task, Table 1 summarizes the proportion of
answers with and without uncertainty – Recoverable and Unrecoverable, respectively – considering
 = 10 runs per item. Answers with no associated uncertainty are grouped into Always correct and
Never correct; answers with associated uncertainty are grouped into Never correct and Sometimes correct.
The rightmost column (Recoverable total) adds Never correct (w/ Unc.) and Sometimes correct to give
the recoverable fraction, i.e., the share of items that uncertainty thresholding can, in principle, surface
for manual correction. It should be noted that Never correct items are partitioned into zero-variance
items (w/o Unc.) and items with observed uncertainty (w/ Unc.).</p>
        <p>Gemma-2-2B-Instruct, the smallest model in terms of parameters, and Llama-3.1-8B-Instruct show
the lowest Always correct proportion, suggesting limited reliability. In contrast, TableLlama and
Gemma-2-9B-Instruct show the highest Always correct proportion, with Qwen2.5-7B-Instruct falling in
between. The Recoverable total highlights the best-case pay-of of uncertainty-guided review. In general,
the fraction of unrecoverable-never-correct cases remains modest.
8The performance difers depending on whether features are extracted from the Postilla or Generated tokens. A more detailed
analysis can be found in Appendix C.</p>
        <p>Gemma-2-2B-it</p>
        <p>Gemma-2-9B-it</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Q1 and Q2: Uncertainty Estimates Assessment</title>
        <p>We assess the reliability of our estimates through a series of targeted experiments. First, we test whether
the uncertainty measures can identify mentions with low answer-level accuracy, where “low” means
that the answer-level accuracy is below 0.5 over  = 10 generations. To this end, we perform a ROC
analysis [39] on both the estimated and baseline entropy measures. We report the results in terms of
true positive rate (TPR) and false positive rate (FPR), where TPR corresponds to the actual low-accuracy
cases that the uncertainty-based method correctly flags. At the same time, FPR is the fraction of
high-accuracy cases (accuracy ≥ 0.5) that the method wrongly flags as low-accuracy. Here, the positive
class corresponds to the decision “flag for manual review.” In the ROC plot, each curve traces the trade-of
between TPR and FPR as a threshold on the uncertainty score—used to decide “flag or not”—is varied.
The diagonal line represents random chance; curves that bow upward (high TPR at low FPR) indicate
efective uncertainty signals for identifying low-accuracy items.</p>
        <p>Moreover, we assess the impact of the uncertainty measures on budget-dependent manual correction.
To this end, we investigate the average accuracy obtained by selecting the most uncertain items up to
a budget , according to a given uncertainty measure. We assess the aggregate accuracy at the dataset
level after correcting a fraction  of the items.</p>
        <p>Detecting Low-Accuracy Answers We evaluate how well the uncertainty estimate can identify
low accuracy items (answer-level accuracy ≤ 0.5). Figure 2 shows the ROC curves for the considered
LLMs, comparing the diferent uncertainty estimates and baselines. ROC curves are a well-established
metric for assessing classification performance, and illustrate the trade-of between true positive rate
and false positive rate across diferent thresholds, in our case, on uncertainty estimates. PE and SE based
on  generations (dashed lines) show the highest performance. The gap between these baselines and the
estimated PE/SE (solid lines) captures the performance loss due to estimating the multiple-generations
entropy with the information from a single generation. Notably, when using the best-performing models,
TableLlama and Gemma-2-9B-Instruct, PE and SE computed on the  multiple generations closely
approximate the estimated PE and SE (cf. green and purple solid and dashed lines). However, across
other models, the PE and SE baselines show a better ability to distinguish high- and low-accuracy cases.
Moreover, PP is consistently outperformed by other methods. In general, the performance of our method
falls between PE/SE and PP. We recall that the estimates of our method are obtained at a fraction of the
computational cost, by computing  = 1 generations instead of  = 10.</p>
        <p>Regarding the use of features derived from diferent stages of the generation process, estimates based
on the features observed in the Postilla segment appear to be generally less informative compared to
the Generated segment. However, focusing on the utilization of the LogitLens features, we highlight
that their contribution is substantial when using the Postilla features, while it becomes negligible when
using the Generated features.</p>
        <p>Uncertainty-Guided Correction under Budget  We evaluate how much accuracy improves when
a human annotator—assumed to be always correct—uses a limited budget  to manually revise the most
uncertain cases. Items are ranked by the various uncertainty signals, and the top  fraction is corrected; the
resulting dataset-level accuracy after these ideal corrections is shown in Figure 3. Each curve corresponds
to a diferent ranking method: the proposed PE estimate (e.g., PE(Generated, MaxP+H)), the a posteriori
PE/SE baseline, perplexity, random selection, and an oracle that ranks strictly by true low-accuracy
severity. The gray oracle curve defines the upper bound, while the yellow random curve gives a reference
for uninformed correction. Across the evaluated models, uncertainty-guided correction substantially
outperforms random selection, with the largest marginal gains at small budgets. The single-shot regressors
closely track the multi-shot PE/SE baselines, recovering most of their improvement while reducing LLM
calls from N to 1, whereas the Perplexity baseline produces consistently lower curves. TableLlama and
Gemma-2-9B exhibit steeper initial slopes, indicating a high concentration of correctable errors among
the most-uncertain items, while other models obtain comparable gains only at larger budget fractions .
TableLlama
Gemma2-9b-it
Qwen2.5-7b-it
LLaMA-3.1-8b-instruct</p>
        <p>Gemma2-2b-it
4.5. Q3: Learnability of the uncertainty regressor ℎ
Since our self-supervised regressor ℎ requires a warm-up phase for learning, we conducted two
supplementary experiments to assess its practical applicability. First, we estimated the number of examples
that are necessary to reach the stability of the regressors’ performance. As a complementary analysis, we
assess if a proxy target measure can be derived, trading some accuracy for a reduced computational cost.
Regressor Convergence as a Function of Training Size We train the regressor with an increasing
number of cases and assess its Spearman correlation with the target. We perform a 10-fold cross-validation,
where in each fold we keep the validation set fixed and train the regressor by expanding the training set
one-by-one. Figure 4 reports the average performance over all the folds for each model. We regress the
PE target only, using the M(p) and H(p) features on the first 10 Generated tokens. Stable performance
is reached with a limited number of items, e.g., 10% − 20% of the dataset, depending on the model.
This highlights how, even with a limited number of examples, the approach can successfully learn to
estimate uncertainty. This, in turn, impacts the overall eficiency and sustainability of the method, since
the number of multiple-generations cases necessary to train the model appears to be limited. We recall
that, since the method is self-supervised, it does not rely on external annotations.</p>
        <p>TableLlama
Gemma-2-9B-it Avg. Accuracy</p>
        <p>Gemma-2-2B-it Avg. Accuracy
TableLlama Avg. Accuracy</p>
        <p>Approximated Ground Truth We finally consider whether reducing the resources dedicated to the
multiple generations can retain enough information to approximate the full uncertainty – that is, the
uncertainty observed with the original  generations when considering all the generated tokens. For this
experiment, performed for illustration purposes, we focus on the best-performing model, TableLlama, and
examine the PE target only. We cap the number of generations and the number of generated tokens, both
of which impact the computational burden linearly. Figure 5 shows the Spearman correlation between the
approximated PE and the full PE. In this setup, even under computational constraints, the derived
approximate PE is reasonably correlated with the full PE, i.e., computed on the complete set of generations and
tokens. This highlights how an approximate target signal could be derived by trading accuracy for eficiency.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.6. Q4: Temperature Sensitivity</title>
        <p>Our approach proposes a measure of uncertainty related to output variability. However, the temperature
setting afects the trade-of between output variability – needed to estimate uncertainty – and task
accuracy. In this section, we evaluate the impact of the temperature on the task performance. We
summarize the performance as the area under the curve (AUC) that describes dataset-level accuracy
as a function of correction budget  (cf. Figure 3 for comparison). We chose this metric as it summarizes
the practical downstream usability of the proposed method.</p>
        <p>We systematically sweep the temperature value in the range 0.0 ≤  ≤ 2.0 with steps of 0.1 to assess
its efect on the AUC. Due to computational budget constraints, we utilize a subset of 200 elements from
the original dataset and focus on PE. Figure 6 reports the AUC over the temperature and the average
overall accuracy obtained with selected models. It can be observed that the most convenient performance
is obtained for balanced  values, for which dataset-level accuracy does not change drastically, implying
that temperature variations up to a certain level do not compromise the task performance. Setting
smaller  values degrades the performance, since lower output variability directly reduces observable
uncertainty. Similarly, setting a higher  value leads to higher observed uncertainties, which are related
to noisy outputs and do not reflect the actual answer uncertainty. Based on these observations, we
set  = 1.0 for all the experiments previously discussed in this paper. Nonetheless, this experiment
illustrates how temperature can be adjusted depending on the specific scenario.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Future Work</title>
      <p>In this work, we propose a lightweight, self-supervised approach for obtaining uncertainty estimates. An
approximate entropy measure is regressed by leveraging features that are observable at inference time,
based on final- and intermediate-layer token probability distributions. This approximation is calculated,
after an initial warm-up phase, without relying on multiple, costly generation runs, making it suitable
for practical use in real-world EL workflows.</p>
      <p>We validate the proposed method on the task of Entity Linking for tabular data, showing a strong
correlation with uncertainty observed over multiple generations estimates, with significantly reduced
computational overhead. Empirical evaluations, performed across several instruction-tuned LLMs, demonstrate
that the method is highly efective at identifying low-accuracy outputs, in particular when using features
derived from generated tokens. We also quantitatively assess the number of cases needed for learning
to predict uncertainty, showing that a limited warm-up phase is suficient. Finally, we measure how the
LLM temperature influences the trade-of between uncertainty and the overall accuracy of the method.</p>
      <p>The proposed method is general, and it may benefit other closed-form tasks beyond the considered
context. Future experiments, aimed at testing its generalizability across tasks and domains – particularly
in open-ended EL settings, where diferent uncertainty measures could have diferent behaviors and
eficiency trade-ofs – could extend the results provided in this study. A warm-up phase, involving
multiple LLM generations, is required for training, which introduces a computational cost. However,
this cost is mitigated by the rapid convergence of the learning phase, maintaining the eficiency of the
overall process. Additionally, evaluating the transferability of learned regressors across datasets related
to the same task may provide further insights into the robustness of the method.</p>
      <p>While our method is model-agnostic and applies unchanged to any model exposing token
logprobabilities (and optionally hidden states), we leave the profiling and evaluation of larger-capacity
models to future work. When only output-layer probabilities are available (e.g., closed models), our
output-layer-only variant remains applicable.</p>
      <p>We also show that some of the features related to the intermediate layers during the generation process
are partially able to fill the performance gap observed between prompt-related and generation-related
tokens. The use of internal LLM state features to anticipate properties of the generated output
represents a promising research direction. Overall, the results reinforce the practical applicability
of uncertainty-aware methods in LLM-based EL workflows, particularly in contexts where quality,
eficiency, and scalability are critical. Beyond mention-level scores, a natural extension is to target the
specific decision tokens (e.g., the span that selects a candidate, or yes/no tokens in classification prompts)
to obtain position-aware confidence; for multi-answer outputs, per-span uncertainties could further
support selective review and partial acceptance.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The work has received funding from the European Union’s Horizon Europe research and innovation
programme under grant agreements No. 101189771 (DataPACT) and No. 101070284 (enRichMyData),
and the Italian PRIN project Discount Quality for Responsible Data Science (202248FWFS).</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT (GPT-4-turbo) in order to perform
grammar and spelling checks. After using this tool, the authors reviewed and edited the content as
needed and take full responsibility for the publication’s content.
model, ACM SIGMOD Record 53 (2024) 44–53.
[20] R. Avogadro, M. Ciavotta, F. De Paoli, M. Palmonari, D. Roman, Estimating link confidence
for human-in-the-loop table annotation, in: 2023 IEEE/WIC International Conference on Web
Intelligence and Intelligent Agent Technology (WI-IAT), IEEE, 2023, pp. 142–149.
[21] Y. Huang, J. Song, Z. Wang, S. Zhao, H. Chen, F. Juefei-Xu, L. Ma, Look before you leap: An
exploratory study of uncertainty analysis for large language models, IEEE Transactions on Software
Engineering 51 (2025) 413–429. doi:10.1109/TSE.2024.3519464.
[22] Z. Lin, S. Trivedi, J. Sun, Generating with confidence: Uncertainty quantification for black-box
large language models, Transactions on Machine Learning Research (2023).
[23] H. Ma, J. Chen, G. Wang, C. Zhang, Estimating LLM uncertainty with logits, 2025. URL:
https://arxiv.org/abs/2502.00290. arXiv:2502.00290.
[24] B. Plaut, N. X. Khanh, T. Trinh, Probabilities of chat LLMs are miscalibrated but still predict
correctness on multiple-choice Q&amp;A, 2025. URL: https://arxiv.org/abs/2402.13213. arXiv:2402.13213.
[25] A. Vazhentsev, L. Rvanova, I. Lazichny, A. Panchenko, M. Panov, T. Baldwin, A. Shelmanov,
Token-level density-based uncertainty quantification methods for eliciting truthfulness of large
language models, in: Proceedings of NAACL 2025, 2025, pp. 2246–2262.
[26] L. Liu, Y. Pan, X. Li, G. Chen, Uncertainty estimation and quantification for LLMs: A simple
supervised approach, 2024. URL: https://openreview.net/forum?id=g3aGMMFHW0.
[27] E. Fadeeva, R. Vashurin, A. Tsvigun, A. Vazhentsev, S. Petrakov, K. Fedyanin, D. Vasilev, E.
Goncharova, A. Panchenko, M. Panov, et al., LM-Polygraph: Uncertainty estimation for language
models, arXiv preprint arXiv:2311.07383 (2023).
[28] B. Lakshminarayanan, A. Pritzel, C. Blundell, Simple and scalable predictive uncertainty estimation
using deep ensembles, Advances in neural information processing systems 30 (2017).
[29] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, R. Salakhutdinov, Dropout: a simple way
to prevent neural networks from overfitting, The Journal of Machine Learning Research 15 (2014)
1929–1958.
[30] S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds,
N. DasSarma, E. Tran-Johnson, S. Johnston, S. E. Showk, A. Jones, N. Elhage, T. Hume, A. Chen,
Y. Bai, S. Bowman, S. Fort, D. Ganguli, D. Hernandez, J. Jacobson, J. Kernion, S. Kravec, L. Lovitt,
K. Ndousse, C. Olsson, S. Ringer, D. Amodei, T. Brown, J. Clark, N. Joseph, B. Mann, S. McCandlish,
C. Olah, J. Kaplan, Language models (mostly) know what they know, CoRR (2022).
[31] S. Lin, J. Hilton, O. Evans, Teaching models to express their uncertainty in words, arXiv preprint
arXiv:2205.14334 (2022).
[32] M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He, B. Hooi, Can LLMs express their uncertainty? an empirical
evaluation of confidence elicitation in LLMs, arXiv preprint arXiv:2306.13063 (2023).
[33] Y. Yang, H. Li, Y. Wang, Y. Wang, Improving the reliability of large language models by leveraging
uncertainty-aware in-context learning, arXiv preprint arXiv:2310.04782 (2023). arXiv:2310.04782.
[34] Y. A. Yadkori, I. Kuzborskij, A. György, C. Szepesvári, To believe or not to believe your LLM, arXiv
preprint arXiv:2406.02543 (2024).
[35] L. Breiman, Random forests, Machine Learning 45 (2001) 5–32.
[36] T. Chen, C. Guestrin, XGBoost: A scalable tree boosting system, in: Proceedings of the
22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,
KDD’16, Association for Computing Machinery, New York, NY, USA, 2016, p. 785–794. URL:
https://doi.org/10.1145/2939672.2939785. doi:10.1145/2939672.2939785.
[37] nostalgebraist, Interpreting GPT: The logit lens, https://www.lesswrong.com/posts/AcKRB8wD
pdaN6v6ru/interpreting-gpt-the-logit-lens, 2020. LessWrong, 2020.
[38] R. Avogadro, M. Cremaschi, F. D’Adda, F. De Paoli, M. Palmonari, et al., Lamapi: a comprehensive
tool for string-based entity retrieval with type-base filters., in: OM@ ISWC, 2022, pp. 25–36.
[39] K. Hajian-Tilaki, Receiver operating characteristic (roc) curve analysis for medical diagnostic test
evaluation, Caspian journal of internal medicine 4 (2013) 627.
[40] E. S. Keeping, Introduction to statistical inference, Courier Corporation, 1995.</p>
    </sec>
    <sec id="sec-8">
      <title>A. Illustrative Example</title>
      <p>The example below illustrates the approach to resolve a mention in a table to its corresponding entity,
using an LLM. A high-level instruction, describing the entity linking task, is provided. A question
contains the mention to be linked, together with a list of entity candidates extracted from a retriever,
e.g., Wikidata Lookup Service or LamAPI. The candidates are provided in the following format: &lt;label
[DESC] description [TYPES] type1, type2, ..., typeK&gt;. A snapshot of the input table is provided,
i.e., N rows above and below the row containing the mention, in Markdown format. The LLM selects
a candidate from the predefined list and outputs it verbatim as an answer.</p>
      <sec id="sec-8-1">
        <title>Entity Linking Prompt Example</title>
        <p>Instruction: This is an entity linking task. The goal for this task is to link the selected entity mention in the
table cells to the entity in the knowledge base. You will be given a list of referent entities, with each one composed
of an entity name, its description and its type. Please choose the correct one from the referent entity candidates.
Input: [TLE] List of high schools in South Dakota.
col: |school|type|city|county|mascot|
row 0: |Aberdeen High School|Private|Aberdeen|Brown|Knights|
row 1: |Agar High School|Public|Agar|Sully|
[...]
row 35: |Crazy Horse High School|Public|Wanblee|Jackson|Chiefs|
row 36: |Crow Creek High School|Public|Stephan|Hyde|Chieftains|
row 37: |Custer High School|Public|Custer|Custer|Wildcats|
[...]
Question: The selected entity mention is: ‘Hyde’. The column name for ‘Hyde’ is ‘county’. The referent
candidates are:
&lt;Dr. Jekyll and Mr. Hyde [DESC] fictional characters [TYPE] group of fictional characters&gt;
&lt;Hyde v Hyde [DESC] landmark case of the English Court of Probate and Divorce [TYPE] legal case&gt;
&lt;Hyde County [DESC] county in South Dakota, United States [TYPE] county of South Dakota&gt;
&lt;Douglas Hyde [DESC] first President of Ireland (1860-1949) [TYPE] linguist&gt;
&lt;Hyde County [DESC] county in North Carolina, United States [TYPE] county of North Carolina&gt;
&lt;Hyde [DESC] civil parish in Bedfordshire, UK [TYPE] civil parish&gt;
&lt;Hyde Park [DESC] town in Dutchess County, New York, United States [TYPE] town of New York&gt;
&lt;Strange Case of Dr Jekyll and Mr Hyde [DESC] novel by R. L. Stevenson [TYPE] literary work&gt;
&lt;Hyde Park [DESC] neighborhood in Chicago, Illinois [TYPE] neighborhood&gt;
[...]
What is the correct referent entity for ‘Hyde’?
Postilla: Answer with just a candidate, selected from the provided referent entity candidates list, and nothing
else. The selected candidate must be reported verbatim from the list provided as input. Each candidate in the list
is enclosed between &lt; and &gt; and reports [DESC] and [TYPE] information.</p>
        <p>Answer: &lt;Hyde County [DESC] county in South Dakota, United States [TYPE] county of South
Dakota&gt;</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>B. Representative Qualitative Examples</title>
      <p>We provide representative examples of the strengths and shortcomings of our method, reporting for each
example: input table, mention, candidate list, generated answers with their frequencies, baseline, and
estimated uncertainties. Recoverable error 1 reports the standard case: a model’s output is uncertain,
with borderline accuracy (0.5), and our one-shot prediction successfully highlights such uncertainty.
Recoverable error 2 reports a more complex case, in which the same mention is referred to over
multiple seasons. As a consequence, the model struggles to fix on one single entity. Again, our method
is able to recover this case by predicting an uncertain outcome. We also report an unrecoverable error
case, in which there is no answer variability. Our method correctly predicts low answer variability and,
consequently, the case is not marked for correction. However, the answer is wrong, and the error cannot
be recovered. Notably, the regressor accurately captures uncertainty across all the examples.</p>
      <p>Recoverable error 1: high answer variability → low accuracy</p>
      <p>
        Candidate entities with description and type (right and wrong answers)
Observed answers (count):
&lt;Madrid [DESC] capital city of Spain [TYPE] municipality of Spain&gt;
&lt;Madrid [DESC] Spanish Congress Electoral District
[TYPE] electoral district of the Spanish Congress&gt;
&lt;Madrid [DESC] None [TYPE] electoral district&gt;
Predictive Entropy:
Semantic Entropy:
0.649 Predicted using our method (avg): 0.588
0.478 Predicted using our method (avg): 0.421
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
Name
Serie A
Recoverable error 2: high answer variability → low accuracy
Observed answers (count):
&lt;1986-87 Serie A [DESC] sports season [TYPE] sports season&gt;
&lt;1994-95 Serie A [DESC] sports season [TYPE] sports season&gt;
&lt;Serie A [DESC] top Italian football league [TYPE] annual sporting event&gt;
&lt;1998-99 Serie A [DESC] sports season [TYPE] sports season&gt;
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
Predictive Entropy:
Semantic Entropy:
0.639 Predicted using our method (avg): 0.571
0.638 Predicted using our method (avg): 0.559
      </p>
      <sec id="sec-9-1">
        <title>Unrecoverable error: no answer variability &amp; zero accuracy</title>
        <p>name
Agaw
Amhara
Beja
Bilen
Gurage
Oromo
Saho
family
Cushitic
Semitic
Cushitic
Cushitic
Semitic
Cushitic
Cushitic
language region</p>
        <p>country
Agaw Horn of Africa
Amharic Horn of Africa
Beja Horn of Africa
Bilen Horn of Africa
Gurage Horn of Africa
Afan Horn of Africa
Oromo
Saho Horn of Africa
...</p>
        <p>Ethiopia. Eritrea
Ethiopia
Sudan. Eritrea
Eritrea
Ethiopia
Ethiopia. Somalia.</p>
        <p>Sudan. Kenya
Eritrea. Ethiopia</p>
        <p>
          Candidate entities with description and type (right and wrong answers)
Observed answers (count):
&lt;Bilen people [DESC] ethnic group [TYPE] ethnic group&gt;
(
          <xref ref-type="bibr" rid="ref10">10</xref>
          )
Predictive Entropy:
Semantic Entropy:
        </p>
        <p>Predicted using our method (avg): 0.044</p>
        <p>Predicted using our method (avg): 0.041</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>C. Position-Dependent Feature Contributions</title>
      <p>
        Figure 2 reveals that performance difers depending on whether features are extracted from the Postilla
or Generated tokens. To investigate this further, we designed an experiment that progressively expands a
sliding window over the concatenated Postilla and Generated segments, measuring how information
accumulates as more tokens are included. We evaluate three feature configurations: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) output-layer features,
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) LogitLens features from intermediate layers, and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) their combination. Figure 7 shows the Spearman
correlation between each configuration and the baseline PE as a function of the window size. Correlation
increases gradually while the window traverses the Postilla tokens, with a pronounced jump once the
Generated tokens are reached. The improvement over Postilla is non-uniform: certain positions—especially the
ifnal tokens—contribute disproportionately, indicating that the informative signal is unevenly distributed.
Additionally, LogitLens features provide a benefit within the Postilla segment but not after entering
the Generated portion, suggesting that during generation, the output-layer features already capture
suficient information, whereas intermediate-layer representations are more useful prior to generation.
These observations imply that feature selection should be adapted based on token origin and position
to maximize efectiveness. This also motivates extensions that localize uncertainty at the answer-bearing
token(s), enabling per-token confidences and finer-grained handling of multi-answer outputs.
      </p>
      <p>E. Answer Variability as a Function of Temperature
D. Spearman Correlation Between Estimated and True PE/SE</p>
      <p>
        To evaluate the variability of answers and recoverable cases, we systematically sweep temperature
values in the range 0.0 ≤  ≤ 2.0 with steps of 0.1 to assess the temperature efect on (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) the number
of items that show output variability and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) the average accuracy. For capacity reasons, we utilize a
subset of 200 elements from the original dataset. Figure 9 shows the results of the experiment for two
selected LLMs. The yellow series shows average accuracy, highlighting how higher  values deteriorate
the performance on the task. At the same time, lower  settings yield the highest proportion of always
correct cases (green line). However, again at low  , output variability is also lowest; the uncertainty of
these cases cannot be estimated, making this setting unsuitable for recovery. Operationally, a reasonable
trade-of for  is to minimize the accuracy loss while maximizing the number of cases recoverable
through uncertainty, that is, those above the red line. Another way to view this trade-of is by examining
the diference between the red and green lines: this diference represents the proportion of unrecoverable
cases, that is, cases that are always wrong and have no output variability. Making these cases recoverable
requires “paying” by reducing always correct cases, while trying to preserve average accuracy (yellow)
as much as possible. This experiment illustrates how temperature can be adjusted depending on specific
combinations of model, task, and application constraints.
      </p>
    </sec>
    <sec id="sec-11">
      <title>F. Transformer-Based Time Complexity</title>
      <p>To further assess the practicality and usability of our approach, in this section we derive the time
complexity of a Transformer-based architecture, such as the ones used by the models considered in
this work. We consider a Transformer with  layers, and  hidden size. The context length is set to  ,
while the number of generated tokens is . The time complexity of a single forward pass over a prompt
 ∈ R× can be decomposed into the following components:
• Self-attention: Given the , ,  ∈ R×  matrices, the time complexity of the self-attention
mechanism is ( 2 · ), where  is the context length and  is the hidden size. The attention
scores are computed as  , and the output is computed as  ( ) . The time complexity
of this operation is ( 2 · ).
• Feed-forward: During the feed-forward step, the time complexity is (8 · 2) = ( · 2)
overall, where  is the hidden size. In this we can include also the projection of the input to the
 space and the final projection to the output space.</p>
      <p>In total one has a time complexity of ( 2 ·  +  · 2) for a single Transformer layer, which becomes
([ 2 ·  +  · 2]), where  is the number of layers.</p>
      <p>If we now suppose to generate  tokens without the use of a KV-cache, the time complexity of the
-th generation step is ([( + )2 ·  + ( + ) · 2]), for every  ∈ []. If we then sum over all
the generations, we have:</p>
      <p>⎛  ⎞
 ⎝∑︁  [︀ ( + )2 + ( + )2]︀ ⎠ =</p>
      <p>=1
⎛  ⎞
 ⎝ ∑︁ [︀ ( + )2 + ( + )2]︀ ⎠ =</p>
      <p>=1
⎛  ⎞
 ⎝ ∑︁ [︀ ( 2 + 2  + 2) + ( + )2]︀ ⎠ =</p>
      <p>=1
 (︀  [︀ ( 2 + 2 2 + 3) + ( + 2)2]︀)</p>
      <p>
        =
 (︀  [︀ ( + )2 + ( + )2]︀) =
 (︀  [︀ ( + )2 + ( + )2]︀)
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
(
        <xref ref-type="bibr" rid="ref7">7</xref>
        )
(
        <xref ref-type="bibr" rid="ref8">8</xref>
        )
(
        <xref ref-type="bibr" rid="ref9">9</xref>
        )
which shows that the time complexity of generating  tokens is quadratic in the number of overall
tokens  + , when  ≪  , otherwise it would become cubic in the number of generated ones.
      </p>
      <p>When a KV-cache is used, while the time for processing the prompt is the same, a major computational
saving is obtained during the generation phase. In this case, the time complexity of the -th generation
step can be decomposed into the following components:
• Self-attention: During the self-attention,  reduces to a single vector  ∈ R1× , while , 
becomes ,  ∈ R(+− 1)× . Overall, the time complexity of this operation is (( + ) · ).</p>
      <p>If we then sum over all the generations  ∈ [], we have:</p>
      <p>=1
which shows that the time complexity of generating  tokens is linear in the number of overall tokens
 + , when  ≪  , otherwise it would become quadratic in the number of generated ones.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Bhagavatula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Noraset</surname>
          </string-name>
          , D. Downey,
          <article-title>TabEL: Entity linking in web tables</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2015</year>
          , pp.
          <fpage>425</fpage>
          -
          <lpage>441</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <article-title>Entity linking via joint encoding of types, descriptions, and context</article-title>
          ,
          <source>in: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>2681</fpage>
          -
          <lpage>2690</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Belotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dadda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cremaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Avogadro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <article-title>Evaluating LLMs on entity disambiguation in tables</article-title>
          ,
          <source>arXiv preprint arXiv:2408.06423</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chabot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.-P.</given-names>
            <surname>Huynh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Labbé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Monnin</surname>
          </string-name>
          ,
          <article-title>From tabular data to knowledge graphs: A survey of semantic table interpretation tasks and methods</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>76</volume>
          (
          <year>2023</year>
          )
          <fpage>100761</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>TableLlama: Towards open large generalist models for tables</article-title>
          , in: K. Duh,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , S. Bethard (Eds.),
          <source>Proceedings of the</source>
          <year>2024</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Association for Computational Linguistics</article-title>
          , Mexico City, Mexico,
          <year>2024</year>
          , pp.
          <fpage>6024</fpage>
          -
          <lpage>6044</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .
          <article-title>naacl-long</article-title>
          .
          <volume>335</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .
          <article-title>naacl-long</article-title>
          .
          <volume>335</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yashar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Fainman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Chaudhuri, Table-GPT:
          <article-title>Table-tuned GPT for diverse table tasks</article-title>
          ,
          <source>arXiv preprint arXiv:2310.09263</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cremaschi</surname>
          </string-name>
          ,
          <string-name>
            <surname>F. D'Adda</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Maurino</surname>
          </string-name>
          ,
          <article-title>stEELlm: An LLM for generating semantic annotations of tabular data</article-title>
          ,
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          (
          <year>2025</year>
          ). URL: https://doi.org/10.1145/3719206.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>I. Jayawardene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Avogadro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Soylu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Tablinkllm:</surname>
          </string-name>
          <article-title>An llm-based approach for entity linking in tabular data</article-title>
          ,
          <source>in: 2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT)</source>
          , IEEE,
          <year>2024</year>
          , pp.
          <fpage>206</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Nararatwong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kertkeidkachorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ichise</surname>
          </string-name>
          ,
          <article-title>Evaluating tabular and textual entity linking in ifnancial documents</article-title>
          ,
          <source>in: 2024 IEEE 18th International Conference on Semantic Computing (ICSC)</source>
          , IEEE,
          <year>2024</year>
          , pp.
          <fpage>130</fpage>
          -
          <lpage>133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>X.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Yu, TURL: Table understanding through representation learning</article-title>
          ,
          <source>ACM SIGMOD Record</source>
          <volume>51</volume>
          (
          <year>2022</year>
          )
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>J.-C. Klie</surname>
          </string-name>
          , R. Eckart de Castilho, I. Gurevych, From Zero to Hero:
          <article-title>Human-In-The-Loop Entity Linking in Low Resource Domains</article-title>
          , in: D.
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chai</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Schluter</surname>
          </string-name>
          , J. Tetreault (Eds.),
          <article-title>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>6982</fpage>
          -
          <lpage>6993</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>624</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>624</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Arnold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dziuba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Löser</surname>
          </string-name>
          , Tasty:
          <article-title>Interactive entity linking as-you-type</article-title>
          ,
          <source>in: Proceedings of COLING</source>
          <year>2016</year>
          ,
          <source>the 26th International Conference on Computational Linguistics: System Demonstrations</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>111</fpage>
          -
          <lpage>115</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Manakul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Liusie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Gales</surname>
          </string-name>
          ,
          <article-title>SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models</article-title>
          ,
          <source>in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>9004</fpage>
          -
          <lpage>9017</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Farquhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kossen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gal</surname>
          </string-name>
          ,
          <article-title>Detecting hallucinations in large language models using semantic entropy</article-title>
          ,
          <source>Nature</source>
          <volume>630</volume>
          (
          <year>2024</year>
          )
          <fpage>625</fpage>
          -
          <lpage>630</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>H.-Y. Huang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>A survey of uncertainty estimation in LLMs meets practice</article-title>
          ,
          <source>arXiv preprint arXiv:2410.15326</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Luohe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Keep the cost down: A review on methods to optimize LLM's KV-cache consumption</article-title>
          ,
          <source>in: First Conference on Language Modeling</source>
          ,
          <year>2024</year>
          , p.
          <fpage>19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Batini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Scannapieco</surname>
          </string-name>
          , et al.,
          <source>Data and information quality</source>
          , Cham, Switzerland: Springer International Publishing
          <volume>63</volume>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>G.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Human-in-the-loop data integration</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>10</volume>
          (
          <year>2017</year>
          )
          <fpage>2006</fpage>
          -
          <lpage>2017</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <article-title>Unicorn: a unified multi-tasking matching</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>