<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>X ( B. Sáez);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>MRT at IberLEF-2025 PRESTA Task: Maximizing Recovery from Tables with Multiple Steps</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maximiliano Hormazábal Lagos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Álvaro Bueno Sáez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Héctor Cerezo-Costas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pedro Alonso Doval</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Alcalde Vesteiro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fundación Centro Tecnolóxico de Telecomunicacións de Galicia (GRADIANT)</institution>
          ,
          <addr-line>Vigo</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>This paper presents our approach for the IberLEF 2025 Task PRESTA: Preguntas y Respuestas sobre Tablas en Español (Questions and Answers about Tables in Spanish). Our solution obtains answers to the questions by implementing Python code generation with LLMs that is used to !lter and process the table. This solution evolves from the MRT implementation for the Semeval 2025 related task. The process consists of multiple steps: analyzing and understanding the content of the table, selecting the useful columns, generating instructions in natural language, translating these instructions to code, running it, and handling potential errors or exceptions. These steps use open-source LLMs and !ne-grained optimized prompts for each step. With this approach, we achieved an accuracy score of 85% in the task.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Table Question Answering</kwd>
        <kwd>Large Language Models</kwd>
        <kwd>Code generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Natural Language Processing (NLP) is nowadays constrained by the amount of information that can
be processed by Large Language Models (LLMs) due to the limited capacity of input data that they
can handle. Some applications that have this limitation are response generation using RAG systems in
which the LLM generates answers with a small subset of retrieved documents[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] as context. In table
questions answering this limitation is aggravated as tables and databases can have millions of records
and columns [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which by today standards will not !t completely in the LLM context.
      </p>
      <p>
        In this paper, we improve the algorithm of Maximizing Recovery from Tables with Multiple Steps
(MRT)[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a multi-step process that implements LLMs and Python code generation to answer questions
as objectively as possible. Our system implements a sequential divide-and-conquer approach in which
LLMs or heuristic algorithms are executed at each step with very speci!c tasks. These steps range from
describing the tables and generating instructions in natural language to producing the source code to
implement the previous instructions, executing it, and parsing the output to obtain the !nal answer.
In comparison with end-to-end solutions, our approach is more explainable as it is easy to debug and
trace which was the cause of a good or bad response.
      </p>
      <p>
        This paper addresses the IberLEF 2025 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] Task PRESTA: Preguntas y Respuestas sobre Tablas en
Español (Questions and Answers about Tables in Spanish) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The code that generated these results is
publicly available1.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>
        As the name implies, question answering (QA) consists of answering questions that normally have
objectively correct answers. Tabular QA requires the system to retrieve responses from knowledge
bases represented as datasets in tables. Recent methods for QA in tabular data, such as TAPAS[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
TAPEX [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] or Omnitab[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], integrate transformers with architectures speci!cally adapted to extract
answers directly from tables as they integrate said tables as context. Parallel to this, LLMs have also
been employed in zero-shot and few-shot strategies, and have proven to show high levels of usability
in QA due to their prior knowledge. One of the bene!ts of zero/few shot strategies is the fact that
domain-speci!c !ne-tuning could be avoided, eliminating the usual needs of gathering, cleaning, and
performing human validation for the task. In addition, recent LLMs include reasoning skills by default.
This helps, but still presents di"culties with complex queries, involving multiple columns, large tables,
or ambiguous questions requiring common-sense knowledge.
      </p>
      <p>
        Another approach is to parse natural language questions into formal queries in programming
languages such as SQL. There are systems like Seq2SQL[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] or TableGPT2[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which are designed to
create SQL queries from relational database queries or Python code, respectively. These methods are
theoretically independent of table size (there is no context limitation) and provide greater transparency
by including intermediate steps to supervise generated queries.
      </p>
      <p>
        Several datasets to evaluate TableQA strategies have been released in the past years. WikiSQL [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
based on Wikipedia and TabFact [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] provided structured evaluation of tabular data. However, these
datasets do not convey the heterogeneity of real-world tabular data, which are usually more complex,
less homogeneous, and unstructured. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. To address this problem, DataBench[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] has been developed,
which brings together 65 real-world datasets with more than 1, 300 manually crafted question-answer
pairs in multiple domains.
      </p>
      <p>
        Apart from the methods with LLMs and code generators mentioned above, others use alternative
strategies such as Retrieval Augmented Generation (RAG) or Chain-of-Thoughts (CoT). For example,
TableRag [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] proposes the use of RAG systems for tabular comprehension tasks, such as QA, employing
techniques such as query expansion and a double transformation to query languages. This process
translates, on the one hand, the schema to be interacted with and, on the other hand, the operation
necessary to identify the cells with the answer. A noteworthy proposal is Chain-of-Table [16], which
implements CoT as an iterative reasoning mechanism. Instead of executing the code in one shot,
programming instructions are executed iteratively to add or discard information from the table until
the !nal answer is found.
      </p>
      <p>Although there has been recent progress in this !eld, certain challenges still persist, such as the
enhancement of reasoning across multiple rows and columns, managing multiple domains and languages,
mixing data from multiple tables, or improving explainability.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Background of this task</title>
      <p>Similar tasks to PRESTA challenge have been recently published. This is the case of SemEval 2025 Task
8: Question-Answering over Tabular Data challenge [17], a task with the same purpose as this, mainly
in English, with tables from other contexts and a wider range of topics.</p>
      <p>
        For this task, we developed MRT: Maximizing Recovery from Tables with Multiple Steps [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This
solution treats the tables as Pandas Dataframes and uses LLMs to generate Python code that could
extract an answer to each question. We now present brie#y the MRT system, which is our starting
point in the PRESTA challenge.
      </p>
      <p>The initial step is the column descriptor module, which analyses the data of each column (description
of the column content meaning, type of data, frequent values, max and min values, etc.). Secondly, the
explainer module, using the previous analysis, generates natural language instructions with an LLM.
These instructions contain the steps needed to get the answer. Then, the coder module generates, also
with an LLM, Python code from the text instructions, and afterward, this code is executed by the runner
module. If an exception occurs during the code execution or answer parsing, the system steps back into
the coder in an iterative looping process until it gets a valid answer or a limit of attempts is exceeded.
Finally, the interpreter module and formatter module implement di$erent approaches for obtaining the
answer in the desirable data type in order to match the expected result for the task. In each step, the
same or di$erent LLM could be used. Usually, we employ an LLM !netuned for code generation within
the coder module.</p>
      <p>Some of the limitations of MRT were:
• It struggles to !lter categorical values when the value in the dataset has a di$erent representation
as it appears in the question. For example, when asked for Obama the system may not !nd the
value with a strict match if the representation in the table is Barack Obama instead.
• The generated natural language instructions were more intricate than they need to be to get the
right response. Sometimes !ltering instructions were added that are not needed to obtain the
response. The generated code produced several exceptions with !lters that at !rst glance are
actually easy to execute.
• It does not scale well with a large number of columns. Some modules include in the prompt
all the columns of the table, their descriptions and statistical information and frequent values.</p>
      <p>However, this does not scale properly when the number of columns is large.</p>
    </sec>
    <sec id="sec-4">
      <title>4. System overview</title>
      <p>The system with the di$erent components is presented in Figure 1. New components were added from
the previous version, column selector and others were deeply modi!ed (explainer, coder to make the
system more resilient to the new challenges posed by the PRESTA challenge.</p>
      <sec id="sec-4-1">
        <title>4.1. New challenges addressed in this task</title>
        <p>Compared with the SemEval task, the PRESTA dataset patented new limitations that our former MRT
system had:
• A larger number of columns in each table implied enormous prompts in the explainer module.</p>
        <p>Sometimes, this leads to exceptions running LLMs. Having more columns increases the changes
of selecting wrong columns to answer a question. For comparison, Semeval task tables had an
average of 24.8 columns, while in this IberLEF task they have an average of 174.1 columns.
• Ambiguous names. Some columns have names that without all the context (or even with it) are
not informative about the content of their cells, or that use initials that are not intuitive. In those
cases, their values are di"cult to interpret even by humans. Some clarifying examples are N_R,
¿Usted es? (LEER_SÓLO PARA LOS QUE HAN CONTESTADO QUE "TRABAJA" EN LA P2014) or
other boolean columns that are possible answers to question, which could not be easily inferred
by the name of the column: Analgésico_antiin!amatorio_1, Antidepresivos (!uoxetina, sertralina,
escitalopram)_1...
• Columns with mixed types. For example, columns that are essentially numerical but actually
contain strings for some of the values. This is the case of columns like Del 1 al 10, con qué
probabilidad votarías al partido político BNG? (From 1 to 10, with which probability you would
vote to BNG party?), whose values are 5, 6, 7, etc., but also contain these options 1 - No le votaría
nunca, 10 - Le votaría siempre. Using numerical operations directly in these columns would throw
exceptions.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Development of new features</title>
        <p>To mitigate and correct both the already known limitations inherited from the original MRT work and
face the new challenges found in the training dataset for PRESTA, we developed and improved some
features.</p>
        <sec id="sec-4-2-1">
          <title>4.2.1. Column selector module</title>
          <p>We designed the column selector module to make an initial !lter of columns before feeding the prompt
of the explainer, to avoid huge prompts that can lead to exceptions or errors. Due to its nature in the
current implementation, the explainer has to know in advance all the columns involved in answering
the question and their descriptions in order to obtain correct natural language instructions. Those
questions must be included in the prompt with the question. Henceforth, this new module tries to !lter
the columns to avoid crashes by large prompts whilst leaving the relevant questions untouched.</p>
          <p>This new step uses an LLM to ask which columns are potentially relevant to that question. To
achieve this, it prompts iteratively the LLM in groups of 25 columns each, giving their names and
descriptions along the question. The prompt emphasizes that, in case of doubt, it should return the
column, in order to not miss the relevant information in this step.</p>
          <p>The output of this module is the set of columns that the LLM considered useful for the explainer
module, conditioned by the query.</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.2.2. Rules for removing uninformative columns</title>
          <p>Column names can be ambiguous and mislead the di$erent modules. We identi!ed exceptional columns
that led to signi!cant and recurrent errors, that we consider are not needed in any of the questions seen
and will not make sense to use in questions of this style. This is the case of "N_R_" ("No Recuerda"/"Don’t
remember"), which the column descriptor wrongly described as "Number of respondents". Having bad
descriptions will produce wrong answers if those columns are used to obtain the answer instead of the
good ones.</p>
          <p>Columns consisting of numerations or that share almost the name with others, where it is di"cult to
identify semantic di$erences between them, were discarded. For example, there is one table that has
columns identi!ed as "Ns_Nc_0", "Ns_Nc_1", "Ns_Nc_2", etc.</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>4.2.3. Clarification instructions in natural language explanations</title>
          <p>The original MRT code returned raw text instructions that were then parsed to obtain a list of steps
that should be coded afterwards.</p>
          <p>The evolution over the original MRT implementation corrects mistakes due to the wrong naming of
columns or variables in the natural language instructions.</p>
          <p>In the new version, the raw text format using in the output was changed to a JSON with the following
properties: instructions: the list of natural language instructions that must be applied to the data to get
the answer, columns: the list of columns used in the instructions and "lter_values, values that will be
used to !lter the information of the table.</p>
          <p>After the natural language instructions are calculated a check is performed to correct the column
names that are misspelled by the model. Levenshtein distance is used to obtain the closest column to
the one appearing in the instructions, directly switching the name if they are not equal.</p>
          <p>With the "lter_values a similar procedure was implemented. Instead of substituting the values, a
clari!cation instruction was added to those originally generated by the system including the old value.
Those instructions have the following format "Be careful!. The value →old value↑ appears in the database
with the following format: →new value↑".</p>
          <p>Furthermore to simplify the task for the coder information about the column types and the typical
values of the column (only for the not-numerical columns) are added at the end of the instructions: "The
column →column name↑ is of type →column type↑ and has the following example values: →column values↑".</p>
          <p>Table 1 shows an example without clari!cation instructions and with clari!cation instructions. The
instructions on the right are more complete and easier to understand than their counterparts on the left
that lack certain information and are more prone to fail.</p>
        </sec>
        <sec id="sec-4-2-4">
          <title>W/o clarification Instructions With clarification Instructions</title>
          <p>1) Count the total number of surveys conducted in 1) Count the total number of surveys conducted in
January January
2) Compare the count of surveys conducted in 2) Compare the count of surveys conducted in
January with the total count of surveys to determine January with the total count of surveys to determine
if most surveys were conducted in January. if most surveys were conducted in January.
3) Be careful! The value enero appears in the
database with the following format: ’Enero’
4) The column ’Mes de realización’ is of type ’object’
and has the following example values: Enero,
Febrero, Marzo</p>
        </sec>
        <sec id="sec-4-2-5">
          <title>4.2.4. Custom Functions for Code Generation</title>
          <p>Assigning the full responsibility of both orchestrating and generating code to a single model could
overload its capabilities. Empirical results supported this hypothesis, as evidenced by recurrent instances
of poor coding practices despite explicit instructions to avoid them. Common problematic patterns
included:
1. Misuse of the group_by function, leading to code errors instead of employing suitable alternatives
that are less prone to produce exceptions.
2. Contradictory ordering operations, such as sorting in descending order initially, followed by an
explicit ascending sort, e$ectively negating previous instructions.</p>
          <p>Inspired by the methodology of [18], though without precomputing function-result cubes, we
developed instead pre-coded generic functions that can be used in multiple contexts but that solve many of
these common mistakes. The prompt conditions the model to use these functions as an alternative to the
pandas implementations. This allowed the model to primarily focus on code orchestration, signi!cantly
reducing its code-generation workload.</p>
          <p>The objective behind these functions was to maintain general applicability rather than targeting
overly speci!c scenarios. For instance, functions are generalized to handle common data tasks, such as
counting occurrences or !ltering data based on numerical conditions.</p>
          <p>These functions were created through a semi-automated approach involving both Large Language
Models (LLMs) and human input:
1. An LLM analyzed the training dataset and proposed generic function templates that could broadly
address the questions of the dataset. Additionally, the LLM had the option to reuse existing
templates. The outcome was a set of function templates capable of solving a substantial portion
of the queries. This functions were obtained from the training split.
2. Human developers implemented these function templates, subsequently verifying that the model
actively employed them, therefore improving empirical performance.
3. During the testing and debugging phase, we validated that pre-coded, generalized functions
e$ectively resolved recurring issues, prompting further expansion of the function pool.
4. Additionally, certain functions incorporated fuzzy decision-making capabilities to enhance
performance, a methodology detailed in the subsequent section, Fuzzy Search of Categorical Values.
The !nal set of developed generic functions is in appendix I.</p>
          <p>Additionally, we speci!cally addressed the challenge of columns misidenti!ed as numerical due
to naming conventions through the dedicated internal function extract_numeric, detailed in the
Appendix A.</p>
          <p>Lastly, another way to conceptualize this methodology is by viewing the provision of these functions
as enabling the model to utilize Python tool-calling capabilities, directly resolving common coding
problems encountered by the model [19].</p>
        </sec>
        <sec id="sec-4-2-6">
          <title>4.2.5. Fuzzy Search of Categorical Values</title>
          <p>As previously described, certain functions within the model pipeline employ fuzzy matching techniques
to mitigate typical errors encountered during data processing, particularly those stemming from
inconsistencies or typographical variations in categorical values. These errors frequently arise due to
the complexity and variability of data names, especially after undergoing multiple processing stages.</p>
          <p>To clarify the motivation behind implementing fuzzy matching, we !rst outline the challenges
experienced in earlier iterations of the pipeline and also in the version presented in this paper. Speci!cally,
accessing column and row names in a reliable manner posed substantial di"culties due to discrepancies
that emerged across the !ve stages of large language model (LLM) processing. Such discrepancies
typically result in false positives and false negatives:
1. False positives occur when the model incorrectly forwards values due to minor deviations in
the text, such as pluralization or capitalization inconsistencies. For example, the value "item"
might erroneously be transformed into "items" or "Item."
2. False negatives occur when originally correct but unusually formatted values are improperly
corrected, thereby introducing errors. For instance, the value "iTem" might incorrectly be normalized
to "item," altering the intended representation.</p>
          <p>To address these challenges, we adopt established fuzzy matching methodologies as detailed in prior
research such as [20]. Rather than relying solely on exact matches—which fail to accommodate minor
textual variations—fuzzy matching techniques allow the model to recognize and utilize values that
closely resemble the target values based on a de!ned similarity threshold.</p>
          <p>An illustrative example of such a fuzzy matching approach in our pipeline is presented through the
Python functions shown in appendix B.</p>
          <p>These functions perform sequentially the following steps:
1. Identifying the target column and value.
2. Recognize empirically that exact matches are often not achievable due to textual variations.
3. If an exact match is unavailable, apply fuzzy matching in a secondary step, selecting the closest
matching value based on the prede!ned similarity threshold and subsequently operating on it.</p>
          <p>In practice, this fuzzy matching strategy enhances signi!cantly the robustness and overall accuracy
of the pipeline, e$ectively reducing the incidence of errors due to minor textual discrepancies.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental setup</title>
      <p>The experimental setup was executed in batches, where all the questions in the dataset were run
through one step before advancing to the following one. Thus, the number of times a model has to
be loaded/unloaded was optimized as each of the steps may use di$erent models. Also, results of the
column descriptor were cached between experiments and was executed only once, given that its output
for a table is independent of the questions.</p>
      <p>The tests were executed in a NVIDIA RTX-a6000 that combines 84 second-generation RT cores, 336
third-generation Tensor cores, and 10,752 CUDA cores with 48 GB of graphics memory for performance.</p>
      <sec id="sec-5-1">
        <title>5.1. Dataset splits</title>
        <p>Although no training of any model has been performed, the splits of the dataset are shown below (see
table 2). The tables used in the test consist of the same tables of the train and the dev splits.</p>
        <sec id="sec-5-1-1">
          <title>Split</title>
          <p>train
dev
test</p>
          <p>Train and dev splits have been used for the development of the modules, whereas test split was solely
used for the validation of the system against the o"cial platform used in the benchmark.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Models</title>
        <p>We decided to use Qwen2 models for the di$erent modules of the system. Speci!cally, we executed two
types of Qwen models: Qwen 2.5 14B3 for all the modules, excepting the coder which used Qwen 2.5
coder 14B4.</p>
        <p>We also made some tests using the recently published model Qwen 3 5 in the explainer, maintaining
Qwen 2.5 and Qwen 2.5 coder for the rest of the modules.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <sec id="sec-6-1">
        <title>6.1. Performance in validation test</title>
        <p>2https://huggingface.co/Qwen
3https://huggingface.co/Qwen/Qwen2.5-14B-Instruct
4https://huggingface.co/Qwen/Qwen2.5-Coder-14B-Instruct
5https://huggingface.co/Qwen/Qwen3-14B</p>
        <sec id="sec-6-1-1">
          <title>Total</title>
          <p>0.71
100</p>
        </sec>
        <sec id="sec-6-1-2">
          <title>Boolean</title>
          <p>0.75
20</p>
        </sec>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Performance in test set</title>
        <p>We submitted 3 results to the task. One is with the results of our system using Qwen 2.5 14B for all the
steps that involve LLMs, except the coder which used Qwen Coder 14B. The second one selects the
output of the interpreter module in that same execution. Finally, a third one changes only the model of
the explainer to Qwen 3 14B. They achieve scores of 85%, 85%, and 83% respectively. We show in
the table 4 our best submission broken down by type of expected answer. All the experiments were
repeated 8 times, taking the most repeated answer (a simple majority voting strategy) as the !nal result.</p>
        <p>Score
Size</p>
        <sec id="sec-6-2-1">
          <title>Total</title>
          <p>0.85
100</p>
        </sec>
        <sec id="sec-6-2-2">
          <title>Boolean</title>
          <p>0.95
20</p>
        </sec>
        <sec id="sec-6-2-3">
          <title>Number</title>
          <p>0.9
20</p>
        </sec>
        <sec id="sec-6-2-4">
          <title>Category</title>
          <p>0.8
20</p>
        </sec>
        <sec id="sec-6-2-5">
          <title>List[Category]</title>
          <p>0.8
20</p>
        </sec>
        <sec id="sec-6-2-6">
          <title>List[Number]</title>
          <p>0.75
20</p>
          <p>In comparison with the score of validation set, in this case is 15% higher. It matches our intuition
that is that the test set questions are in average simpler than the validation set questions.</p>
        </sec>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Manual Error Analysis in the validation set</title>
        <p>We performed a manual error analysis of the answers for the validation set #agged as an error by the
evaluator. The results are summarized in Table 5. The main source of errors is the wrong generation of
instructions in the explainer module. Some of the errors involve not selecting the correct column to
use, although sometimes the ambiguity of the column names makes this choice di"cult. For example,
columns such us ’Edad’ vs ’Edad_recodi!cada’ in which one is just a higher level of abstraction from
the other. Other cases involved just wrong natural language instructions. Adding unnecessary extra
!lters (removing nulls, zeros, empty lists, etc.) was very frequent source of errors.</p>
        <p>The other main source of errors is the removal of relevant columns in the column selector. Sometimes
it struggles with columns that have long names and even involve complex semantics such as conditionals,
like survey questions present in this dataset.</p>
        <p>We identi!ed a few errors in the interpreter for cases where the runner actually obtained the correct
response. This case usually involves incorrectly removing symbols or clari!cations. For instance, there
are two related to age intervals: "+65" and "18-24" which were changed to "65" and "1824". Other less
frequent errors were due to the transformations of the output that make the metric implemented to
consider it as an error. The removal of the parenthesis and the information within in ’PP’ in the expected
answer ’PP (Partido Popular)’. This last example is #agged as incorrect according to the benchmark
metrics but that could be perfectly be deemed as acceptable by common sense with human feedback.</p>
        <p>Compared to manual error analysis in our previous work for the analogous SemEval task, we can
prove the bene!ts of some of our new features taking into account that wrong cell !ltering is not an
issue anymore as it was before and code errors and code exceptions have been notably reduced.</p>
      </sec>
      <sec id="sec-6-4">
        <title>6.4. Ablation study</title>
        <p>We have performed an ablation study to evaluate the impact of some of the features on the performance
of our system. To do this, we deactivate one by one the new modules and execute the benchmark with
the validation set. Every con!guration is repeated 8 times, taking a majority voting ensemble as the</p>
        <sec id="sec-6-4-1">
          <title>Description</title>
          <p>Wrong Instructions</p>
          <p>Wrong column filtering</p>
          <p>Formatting (transformations)
Code Generation (incl. exceptions)</p>
          <p>Others
!nal result, discarding thrown exceptions or bad results (e.g. responses such as No matching records
were found).</p>
          <p>Table 6 shows the scores of each of the tries broken down by the type of expected answer. As can be
seen, the overall metric is always in the range of 0.69 and 0.74. The number of question-answer pairs is
very low (100) and the diversity of tables used in the validation set (4) is not enough to con!rm whether
the new modules improved or not the performance of the system in the test. For example, the column
selector module was implemented to avoid throwing exceptions when the number of columns was very
large. Nevertheless, none of the 4 datasets used in the validation throws this exception. Hence !ltering
the columns could have a detrimental e$ect if a good column was removed. Filtering out columns has a
positive impact on time consumption, as the execution of the overall system is much faster (e.g. about
three times in our experiments within the PRESTA dataset).</p>
        </sec>
        <sec id="sec-6-4-2">
          <title>Score in dev</title>
        </sec>
        <sec id="sec-6-4-3">
          <title>Boolean</title>
        </sec>
        <sec id="sec-6-4-4">
          <title>Number</title>
        </sec>
        <sec id="sec-6-4-5">
          <title>Category</title>
        </sec>
        <sec id="sec-6-4-6">
          <title>Scenario</title>
        </sec>
        <sec id="sec-6-4-7">
          <title>Scenario</title>
          <p>All (formatter)</p>
          <p>All (Interpreter)
w/o column selector
w/o custom functions
w/o explainer corrector
w/o retries in coder
w/o fuzzy subs.</p>
          <p>By seeing these results, one question that might arise is whether all the experiments are failing in the
same questions. In other words, if the errors in the validation set are not addressed by the new modules.
Figure 2 shows the error repetition frequency of errors in all the experiments. We can see that more
than half of the errors are coincident in all the experiments.</p>
          <p>Combining the information of the most repeated errors (6 and 7 times) with the information of the
manual analysis, we !nd that most errors came from the bad selection of columns (9) or the explainer
not being able to generate good instructions in natural language (7). Some errors were formatting issues
(2) and errors due to the validation metric but they were essentially correct (3).</p>
          <p>Finally, we can observe the e$ect of using an ensemble with majority voting in Figure 3. An ensemble
of 5 experiments should be enough, although we have used 8 repetitions in our con!guration.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions</title>
      <p>In this work, we presented our system MRT for answering queries over tables. Our system generates
Python code involving multiple steps: describing and !ltering columns, generating natural language
instructions, code generation, and formatting the answer. This strategy builds upon our previous work
in which we addressed some of the common sources of problems made by the previous version: adding
auxiliary functions to avoid recurring exceptions, fuzzy match of cell and column names to improve
the understanding of the question and instructions, and the selection of relevant columns to avoid
LLM exceptions due to context size and improving the speed of the system at the same time. We used
middle-size pretrained LMs with 14B parameters achieving a third place in the task with a 0.85% of
accuracy. One clear bene!t of our approach is its explainability as the user can very easily understand
what is the source of the errors by seeing the natural language instructions and the generated code.</p>
      <p>However, evaluating the bene!ts of the theoretical improvements is di"cult as the dataset lacks the
size and diversity in order to be statistically relevant. The di$erences between the con!gurations are
very small (between 1 and 5 question/answer pairs). In the future, we plan to test the system against
larger datasets in order to gain more insights into the relevance of each block in the !nal answer.</p>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used Grammarly in order to: Grammar and spelling
check and GPT-4 in order to: suggestions for academic writing style. After using these tools, the authors
reviewed and edited the content as needed and take full responsibility for the publication’s content.
[16] Z. Wang, H. Zhang, C.-L. Li, J. M. Eisenschlos, V. Perot, Z. Wang, L. Miculicich, Y. Fujii, J. Shang,
C.Y. Lee, T. P!ster, Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding,
2024. URL: https://arxiv.org/abs/2401.04398. arXiv:2401.04398.
[17] J. Osés Grijalba, L. A. Ureña-López, E. Martínez Cámara, J. Camacho-Collados, SemEval-2025 Task
8: Question Answering over Tabular Data, in: Proceedings of the 19th International Workshop on
Semantic Evaluation (SemEval-2025), Association for Computational Linguistics, Vienna, Austria,
2025.
[18] F. Zhou, M. Hu, H. Dong, Z. Cheng, S. Han, D. Zhang, Tacube: Pre-computing data cubes for
answering numerical-reasoning questions over tabular data, 2022. URL: https://arxiv.org/abs/2205.
12682. arXiv:2205.12682.
[19] S. He, Achieving tool calling functionality in llms using only prompt engineering without
!netuning, 2024. URL: https://arxiv.org/abs/2407.04997. arXiv:2407.04997.
[20] S. Sheu, A. Chang, W. Huang, Fast similarity search in string databases, in: 19th International
Conference on Advanced Information Networking and Applications (AINA’05) Volume 1 (AINA
papers), volume 1, 2005, pp. 617–622 vol.1. doi:10.1109/AINA.2005.185.</p>
    </sec>
    <sec id="sec-9">
      <title>A. Appendix I: Custom functions for coder</title>
      <p>This section contains the de!nitions of the functions that have been implemented to help the coder to
perform some common operations.
def flatten_column_values_from_df(df: pd.DataFrame, column: str) -&gt; pd.DataFrame:
def get_top_n_records_with_non_nan_column_value(
df: pd.DataFrame, column: str, number: int
) -&gt; pd.DataFrame:
def get_tail_n_records_with_non_nan_column_value(
df: pd.DataFrame, column: str, number: int
) -&gt; pd.DataFrame:
def delete_rows_by_column_value(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; pd.DataFrame:
def sort_dataframe_column_alphabetical_order(df: pd.DataFrame, column_name: str):
def filter_rows_by_column_equals_or_less_than_numeric_value(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; pd.DataFrame:
def filter_rows_by_column_strictly_less_than_numeric_value(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; pd.DataFrame:
def filter_rows_by_column_equals_or_higher_than_numeric_value(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; pd.DataFrame:
def filter_rows_by_column_strictly_higher_than_numeric_value(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; pd.DataFrame:
def filter_rows_that_contain_column_value(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; pd.DataFrame:
def filter_rows_that_do_not_contain_column_value(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; pd.DataFrame:
def exists_value_in_column(df: pd.DataFrame, column: str, value) -&gt; bool:
def count_elements_equal_to_value_in_column(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; int:
def count_elements_containing_value_in_column(
df: pd.DataFrame, column: str, value: Any = None
) -&gt; int:
def find_n_most_frequent_elements_in_column_subset(
df: pd.DataFrame, target_column: str, subset_column: str, filter_value, n: int
) -&gt; list:
def find_most_frequent_element_in_column_subset(
df: pd.DataFrame, target_column: str, subset_column: str, filter_condition
):
def find_most_frequent_element_in_column(df: pd.DataFrame, column: str = None):
def find_n_most_frequent_elements_in_column(
df: pd.DataFrame, column: str, n: int
) -&gt; list:</p>
    </sec>
    <sec id="sec-10">
      <title>B. Appendix II: Fuzzy filters</title>
      <p>This section contains the code of the functions that implement the fuzzy match !ltering in order to
explain the steps that this !ltering follows.
pd.api.types.is_string_dtype(df[column])
and isinstance(value, str)
and value != ""
best_match = _best_fuzzy_match(df[column], value, threshold)
if best_match is not None:</p>
      <p>fuzzy = df[df[column] == best_match]
if _round_was_useful(original_len, len(fuzzy)):</p>
      <p>return fuzzy</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Rethinking Tabular Data Understanding with Large Language Models</article-title>
          , in: K. Duh,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , S. Bethard (Eds.),
          <source>Proceedings of the</source>
          <year>2024</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Association for Computational Linguistics</article-title>
          , Mexico City, Mexico,
          <year>2024</year>
          , pp.
          <fpage>450</fpage>
          -
          <lpage>482</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .
          <article-title>naacl-long</article-title>
          .
          <volume>26</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .
          <article-title>naacl-long</article-title>
          .
          <volume>26</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ruan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lan</surname>
          </string-name>
          , J. Ma,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <article-title>Language modeling on tabular data: A survey of foundations, techniques</article-title>
          and evolution,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2408.10548. arXiv:
          <volume>2408</volume>
          .
          <fpage>10548</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hormazabal-Lagos</surname>
          </string-name>
          , Álvaro Bueno Saez,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cerezo-Costas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Doval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Vesteiro</surname>
          </string-name>
          , MRT at SemEval
          <article-title>-2025 Task 8: Maximizing Recovery from Tables with Multiple Steps</article-title>
          ,
          <source>in: Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Vienna, Austria,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>González-Barba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chiruzzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <article-title>Overview of IberLEF 2025: Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the 41st Conference of the Spanish Society for Natural Language Processing (SEPLN 2025), CEUR-WS</article-title>
          . org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Osés-Grijalba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Ureña-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Cámara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Camacho-Collados</surname>
          </string-name>
          , Overview of PRESTA at IberLEF 2025:
          <article-title>Question Answering Over Tabular Data In Spanish, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the 41st Conference of the Spanish Society for Natural Language Processing (SEPLN 2025), CEUR-WS</article-title>
          . org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Herzig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Nowak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piccinno</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Eisenschlos,</surname>
          </string-name>
          <article-title>TaPas: Weakly Supervised Table Parsing via Pre-training, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</article-title>
          , Association for Computational Linguistics,
          <year>2020</year>
          . URL: http://dx.doi. org/10.18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>398</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>398</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ziyadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , J.-G. Lou,
          <article-title>Tapex: Table Pre-Training via Learning a Neural SQL Executor</article-title>
          ,
          <source>arXiv preprint arXiv:2107.07653</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Neubig</surname>
          </string-name>
          , W. Chen,
          <article-title>OmniTab: Pretraining with Natural and Synthetic Data for Few-Shot Table-based Question Answering</article-title>
          ,
          <source>arXiv preprint arXiv:2207.03637</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Fan,
          <string-name>
            <given-names>G.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <article-title>A Survey of NL2SQL with Large Language Models: Where are we, and Where are we Going?</article-title>
          ,
          <source>arXiv preprint arXiv:2408.05109</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , G. Zhang, G. Chen, G. Zhu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          , et al.,
          <article-title>TableGPT2: A Large Multimodal Model with Tabular Data Integration</article-title>
          ,
          <source>arXiv preprint arXiv:2411</source>
          .
          <year>02059</year>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <article-title>Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning</article-title>
          ,
          <year>2017</year>
          . URL: https://arxiv.org/abs/1709.00103. arXiv:
          <volume>1709</volume>
          .
          <fpage>00103</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , W. Y. Wang,
          <article-title>TabFact: A Large-scale Dataset for Table-based Fact Veri!cation, 2020</article-title>
          . URL: https://arxiv.org/abs/
          <year>1909</year>
          .02164. arXiv:
          <year>1909</year>
          .02164.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>W.</given-names>
            <surname>Hwang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Seo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Comprehensive</given-names>
            <surname>Exploration on WikiSQL with Table-Aware Word</surname>
          </string-name>
          <string-name>
            <surname>Contextualization</surname>
          </string-name>
          ,
          <year>2019</year>
          . URL: https://arxiv.org/abs/
          <year>1902</year>
          .01069. arXiv:
          <year>1902</year>
          .01069.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J. Osés</given-names>
            <surname>Grijalba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Ureña-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Martínez</given-names>
            <surname>Cámara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Camacho-Collados</surname>
          </string-name>
          ,
          <article-title>Question Answering over Tabular Data with DataBench: A Large-Scale Empirical Evaluation of LLMs</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            , M.-
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Hoste</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lenci</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sakti</surname>
          </string-name>
          , N. Xue (Eds.),
          <source>Proceedings of the 2024 Joint International Conference on Computational Linguistics</source>
          ,
          <article-title>Language Resources and Evaluation (LREC-COLING 2024), ELRA</article-title>
          and
          <string-name>
            <given-names>ICCL</given-names>
            ,
            <surname>Torino</surname>
          </string-name>
          , Italia,
          <year>2024</year>
          , pp.
          <fpage>13471</fpage>
          -
          <lpage>13488</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          . lrec-main.
          <volume>1179</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.-A.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Miculicich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Eisenschlos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fujii</surname>
          </string-name>
          , H.-T. Lin,
          <string-name>
            <given-names>C.-Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          , T. P!ster,
          <source>TableRAG: Million-Token Table Understanding with Language Models</source>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2410.04739. arXiv:
          <volume>2410</volume>
          .
          <fpage>04739</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>