<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Navigating the Fermi Multiverse: Assessing LLMs for Complex Multi-hop Queries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mostafa Rahgouy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hamed Babaei Giglou</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dongji Feng</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Taher Rahgooy</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gerry Dozier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cheryl D. Seals</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Software Engineering, Auburn University</institution>
          ,
          <addr-line>Auburn, AL</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mathematics</institution>
          ,
          <addr-line>Computer Science</addr-line>
          ,
          <institution>and Statistics, Gustavus Adolphus College</institution>
          ,
          <addr-line>MN</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Meta</institution>
          ,
          <addr-line>Menlo Park, CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>TIB Leibniz Information Centre for Science and Technology</institution>
          ,
          <addr-line>Hannover</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recently, large language models (LLMs) have gained significant attention in the field of Natural Language Processing (NLP) and have shown promise across various tasks, even when given only a few examples to learn from. However, their ability to understand and reason with natural language remains uncertain. While there have been attempts to evaluate these models using reasoning tests, these evaluations have mostly focused on models' final answers, often overlooking the step-by-step reasoning processes behind their performance. Additionally, these analyses have typically concentrated on just one or a few aspects of reasoning, especially for tasks that do not require much complex thinking to find the answer. This limits our understanding of LLMs' potential and limitations when it comes to more complex and realistic questions. To address this issue, we conduct a comprehensive analysis of LLMs using the existing Fermi reasoning challenge, a task that combines diferent aspects of reasoning into a single question-answering format, requiring deeper levels of reasoning. In this paper, we examine various advanced LLMs in this reasoning challenge and explore how their performance is afected by their size (i.e., the number of parameters). We also investigate how these models behave with diferent levels of supervision, ranging from having all the information to no evidence at all. Furthermore, we compare the two primary methods of teaching these LLMs, fine-tuning, and few-shot learning, using the Chain-of-Thought approach. We provide a detailed case study highlighting the most common limitations of these models. While our results imply that these models may have a long journey ahead to reach human-level reasoning, our work can be considered a robust baseline for the community to strive toward achieving this ambitious goal. Our code is available on GitHub https://github.com/MostafaRahgouy/LLMs_for_FPs for the community.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;NLP</kwd>
        <kwd>Natural Language Reasoning</kwd>
        <kwd>LLMs</kwd>
        <kwd>QA</kwd>
        <kwd>Fermi Problems</kwd>
        <kwd>Few-shot Learning</kwd>
        <kwd>Fine-tuning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Throughout history, humans have grappled with the concept of reasoning, seeking to define
and understand it. This pursuit dates back to the early Greek philosophers, who posed profound
questions such as “What can be known?” and “What does that mean someone knows something?”
in their quest to illuminate the nature of reasoning [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A more recent definition involves
stepby-step or systematic thinking that guides humans toward correct answers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. With the advent
of Artificial Intelligence (AI), the development of systems facilitating reasoning has emerged
as a paramount objective for researchers in this domain ([
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]). This aspiration has drawn
nearer to realization with the introduction of LLMs. These models undergo a two-phase training
process, commonly referred to as pre-training and fine-tuning. In the pre-training phase, they
are exposed to vast volumes of data, equipping them with a foundational understanding of
language. Subsequently, in the fine-tuning phase, these models refine their capabilities by
learning to excel at specific downstream tasks. Prominent exemplars of such LLMs include
BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], BART [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and T5 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which have consistently outperformed their predecessors
across a spectrum of NLP tasks, including reasoning challenges. Researchers have pushed the
boundaries further by introducing super-large language models capable of addressing various
questions with minimal or even zero examples, denoted as few-shot and zero-shot learning,
respectively [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Models such as GPT-4 and LLaMA [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], have demonstrated remarkable prowess
in these areas. Nevertheless, in numerous NLP tasks, LLMs often approach or even surpass
human-level performance. However, their ability to reason falls significantly short of human
capabilities, warranting further in-depth investigation to enhance these models. For example,
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] has pointed out a significant issue where LLMs like GPT-3 and BLOOM struggle with
simple common-sense planning tasks, which humans find easy. Moreover, [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] conducted
experiments and found that existing LLMs are still unable to pass the Theory-of-Mind tests,
where Theory-of-Mind tests aim to assess LLMs abilities to understand and infer the intentions,
emotions, and mental states of others. The root cause of the uncertainty regarding the ability of
LLMs in reasoning can be traced back to the initial works that published tasks and benchmark
datasets that are overly simplistic for these large models. Such simplicity inadvertently provides
opportunities for models to employ suboptimal techniques, thus potentially skewing their
performance [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. To address this issue, we have recently witnessed significant eforts
within the community aimed at devising more intricate tasks that LLMs to not only comprehend
but also engage in reasoning when responding to questions [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ]. These endeavors have
resulted in the emergence of more intricate reasoning tasks, typically taking the form of QA
formats, which require multi-hop reasoning capabilities. Nevertheless, while these studies have
been invaluable and have contributed intriguing insights into LLMs, they have inadvertently
overlooked certain crucial factors. Firstly, many of these tasks involve a limited number of
hops, often restricted to two or three steps. Secondly, these tasks are typically structured as
true/false or multiple-choice questions, which may not capture the nuanced behavior of LLMs
in responding. Additionally, these investigations failed to consider the correlation between
the level of supervision and LLMs performance. In essence, it is essential to explore how
LLMs perform when provided with various levels of information, ranging from complete
information as seen in mathematical word problems to partial or even no information, to gain a
comprehensive understanding of their behavior and decision-making processes. To this end, we
selected the existing Fermi Reasoning Challenge introduced by [17]. This selection addressed
the aforementioned issues by presenting multi-level tasks that demand a more profound level
of reasoning. Furthermore, Fermi Problems (FP) inherently require an approximation in their
responses, as precise answers are often impossible or impractical to attain. Our contribution
can be summarized as:
1. We present a comprehensive assessment of LLMs applied to Fermi Problems across
diferent levels of supervision. Our study focuses on leading LLMs, including T5, Flan-T5,
and models from the GPT family, marking the first of its kind in applying these models to
the FP reasoning task. Therefore, this paper can serve as a reference point for approaching
this challenging task.
2. We investigate various approaches for training and inferring from LLMs. Specifically, we
assess the impact of fine-tuning, both with and without prompting. We also explore the
utility of Chain-of-Thought prompting [18] and evaluate the efectiveness of few-shot
learning while varying the number of provided examples. Furthermore, we delve into the
application of zero-shot learning for implicit reasoning in FP.
3. We ofer a more in-depth analysis of the behavior exhibited by LLMs, particularly
focusing on the most common errors they tend to make. Additionally, we investigate the
relationship between the size of LLMs and their performance for FP. Furthermore, we also
unearth intriguing latent insights that ofer valuable guidance for potential enhancements
in these models.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Fermi Problem</title>
      <p>The Fermi challenge [17] inspired by Enrico Fermi, the Nobel winner in physics known for
his remarkable skill in making accurate estimates of complex numerical problems, is often
referred to as “Fermi problems”. These problems typically involve making assumptions and
approximations to arrive at a rough estimate, rather than a precise calculation. Owing to the
intrinsic complexity of the reasoning questions involved, FPs have been appropriated for use in
science Olympiads and interviews.
2.1. Tasks
FP encompasses three distinct tasks: perfect-context, distractor-context, and full, which
are designated as Task 1, Task 2, and Task 3, respectively. These tasks can be seen as diferent
levels of supervision or evidence provided to a model. Figure 1 illustrates an example of the FP
for these three tasks.</p>
      <sec id="sec-2-1">
        <title>2.1.1. Task 1: Perfect-Context</title>
        <p>At this level, alongside the given question, all essential knowledge (defined as a set of facts)
required to answer the question is integrated into the input. This task bears resemblance to math
word problems, which often feature a concise narrative outlining a scenario and presenting
a question related to an unknown quantity [19]. In both of these scenarios, information is
explicitly provided, eliminating the need for retrieving knowledge, and instead emphasizing
how such information is interconnected and can be used as guidelines to arrive at the final
answer.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.1.2. Task 2: Distractor-Context</title>
        <p>In realistic scenarios, the input context often comprises non-relevant information, and models
must efectively discern which pieces are pertinent to the given question. In line with this
If every human that's ever lived was
brought back to life. What would be the
population density of people per square</p>
        <p>kilometer?
F1: The total surface area of Earth is
510e+6 km square
F2: The total population on earth is
7.2e+9
F3: Around 55e+6 people die each year
F4: It has been 120000 years since the
start of human-life on earth
rFe5q:uTirheed mtoinriempuompunlautmebtheer eonftpireeohpuleman
race is 500
s
tac F6: There are 192 countries in the world
F
ittrrsacoD 5FF1720:01:T0Th0eh0et0ot0toatslaqllainncdoamr...eeagoenneeraarttehdislast
year in the world is 80e+12
real-world complexity, the FP includes this task that involves the deliberate inclusion of some
distractor facts alongside the question and relevant information.
2.1.3. Task 3: Full
The ultimate goal is to enable models to answer complex questions without any provided
information. In essence, the model should learn to retrieve supporting facts and demonstrate
the reasoning behind the answer based on the retrieved supporting facts. Achieving this level
of abstraction is akin to human-style problem-solving.
2.1.4. Outputs
Regardless of the chosen task, two acceptable outputs can be provided, which may serve as
alternatives to each other: Ans  and Program P. Ans  signifies the direct answer to a given
question. Furthermore, each FP question is complemented by an explanation in the form of an
executable program. This program delineates the facts, values, and mathematical computations
essential for deriving the answer. Program P has the potential to demonstrate how models
engage in reasoning across various facets, including question decomposition, fact matching,
and the relationship between questions in terms of mathematical operations. Given the paper’s
primary focus on LLM behavior and Program P’s role in facilitating the interpretability of</p>
        <sec id="sec-2-2-1">
          <title>REALFP</title>
        </sec>
        <sec id="sec-2-2-2">
          <title>SYNTHFP 8000</title>
          <p>Train
185
Validation
125*
1000
Test
558
1000</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>Statistics for both RealFP and SynthFP datasets. However, [17] mentioned that the validation set (* ) of RealFP contains 185 samples, but we found only 125 in their released dataset.</title>
          <p>models, we will exclusively present experiments based on this output, excluding Ans . This
decision allows us to delve deeper into the model’s reasoning processes and provides valuable
insights into its decision-making capabilities.
2.2. Datasets
[17] presented two distinct datasets for analysis: RealFP, comprising of real-world FP collected
from various internet pages, quizzes, and Fermi problem Olympiads; and SynthFP, a larger
synthetic dataset created manually from 12 diferent templates. Table 1 shows statistics of the
datasets. As the test set of RealFP better represents the Fermi challenges in the real world,
throughout this paper, we will use this set to report the performance of the models. Additionally,
the RealFP validation set is employed for early stopping mechanisms for the models in our
experiments, and we chose not to use the validation and test sets of Synthetic to maintain
consistency and fairness across all models.
2.3. Metrics</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Answer Evaluation:</title>
        <p>To evaluate the performance of a predicted direct answer ′ over the
true answer , the  _ with the following definition will be employed:
 _ =  0, 1 −
︂{
1
3 ⃒
⃒
⃒⃒ 10  ⃒⃒
′ ⃒⃒ }︂
(1)
This metric considers the imprecision and uncertainty of answers by assigning a full score to a
prediction that produces an answer within the same order of magnitude as the reference gold
answer. Conversely, for each order of magnitude that the prediction diverges from the reference
answer, the score is reduced by 1/3 points.</p>
        <p>Program Evaluation: Explanations (programs) are evaluated along three criteria:
receives a score of 0.</p>
        <p>_.
• Validity (valid?): This assesses whether the program is syntactically valid by evaluating it
to determine if it results in a numerical output. A Python program executor is employed
for this purpose. A score of 1 is assigned if the execution is successful, otherwise, it
• Answer Accuracy Evaluation ( ): After the program successfully generates a numeric
answer (passed the valid? evaluation with score 1), this metric assesses the accuracy of
the generated numeric answer ′ with ground truth  based on the previously defined
• Fact Identification ( Facts): Determines whether the program includes all and only the
specified gold facts F, using an F1 measure for assessment.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Design</title>
      <p>Our experiments are structured around diferent LLMs training approaches, including
finetuning, few-shot, and zero-shot learning, with and without Chain-of-Thought (CoT) prompting.
3.1. Fine-tuning Setting
Fine-tuning without Chain-of-Thought: The introduction of Transformers, along with
the attention mechanism [20], marked a new era in transfer learning. This innovation enabled
the training of large models in an unsupervised manner on extensive datasets, allowing them to
acquire semantic knowledge of the language, thus forming a robust foundation for adapting to
downstream NLP tasks. This approach has significantly improved performance across a wide
range of NLP tasks with limited training samples, a practice commonly referred to as fine-tuning.
Notable examples of these models include BERT, BART, and T5. Among the aforementioned
models, T5 stands out as particularly adept at handling reasoning tasks due to its inherent
architecture of sequence-to-sequence (seq2seq), which aligns well with the demands of such
tasks. Consequently, we have chosen this model as the foundation for our fine-tuning setting.
To achieve this, we employ a straightforward approach: we concatenate the available supportive
facts with the question (for tasks 1 and 2) to construct the input, which the model processes,
ultimately yielding Program P as the output. Furthermore, we assess the performance of T5
using various versions, including T5-small and T5-base. It is worth noting that we also explored
a larger variant of T5, namely T5-large, which boasts 770 million parameters. However, our
ifndings indicated a degradation in results, which could potentially be attributed to the limited
number of samples available within the datasets.</p>
      <p>Fine-tuning with Chain-of-Thought: In the realm of instruction-based fine-tuning, [ 21]
conducted an investigation into the impact of various factors, including scaling the number
of tasks and model size, and the incorporation of CoT data during the fine-tuning process.
Specifically, in their paper, they outlined their objective as follows:
“The goal of Flan finetuning is to produce an improved checkpoint across a range of
evaluations, which includes multi-step reasoning ability in addition to traditional NLP
tasks”.</p>
      <p>To achieve this goal, they integrated nine CoT datasets into the fine-tuning phase and
demonstrated the positive impact of this approach on unseen reasoning tasks. Additionally, they made
Flan-T5 checkpoints publicly available, maintaining consistency with prior versions of publicly
released T5 checkpoints. As a result of these considerations, we selected the Flan-T5 model as
our experimental model, which incorporates the CoT capability. However, prompt engineering
can be beneficial in tailoring prompts to suit specific tasks. Nonetheless, we opted to maintain
consistency by using the original CoT prompt (“Answer the following question step by step”),
as utilized in the original paper, to ensure fairness across diferent tasks and model sizes in our
experiments.
3.2. Few-Shot Setting
LLMs have popularized the notion of few-shot learning, where these models can acquire new
tasks with just a small number of examples [22, 23]. Few-shot learning ofers significant benefits
as it reduces the necessity for extensive data collection, which can be costly. Investigating
few-shot reasoning for FPs can determine whether such models can address intricate questions
in an interpretable manner. To this aim, we explore the capabilities of the GPT-based family
with varying numbers of provided examples, as typically encountered in few-shot learning
scenarios.
3.3. Zero-Shot Setting
Embracing the use of super-large LMs has opened the door to unlocking zero-shot reasoning
capabilities. Prominent models in this category include GPT-4, LLama, Flan-PaLM, Flan-T5,
and Bloom. Furthermore, the fusion of these models with the CoT prompting has yielded
significant improvements in various tasks [ 24]. However, CoT imposes a requirement for
these models to generate answers through step-by-step reasoning processes. Nonetheless,
these answers may not align with FP’s Program P output. This discrepancy arises because
the specified program P does not conform to the standard output generated by such models,
and they encounter challenges in producing responses without any prior exposure to relevant
samples. Consequently, we explored and evaluated zero-shot reasoning by focusing solely on
direct answers (Ans ). While this approach may limit interpretability to some extent, assessing
their performance under these conditions can ofer valuable insights.
3.4. Experiments Setup
In our fine-tuning process, we used a constant random seed value throughout all experiments.
We maintained a batch size of 8 for the entire duration. Additionally, we set the learning
rate to 1e-3 and utilized the Adam optimizer [25]. To monitor model performance and ensure
reproducibility, we employed the validation set from the real dataset. The best model checkpoints
were saved based on this validation set loss. For consistency and fairness in reporting results, we
conducted fine-tuning for 50 epochs on real data and 5 epochs each on synthetic and combined
(both) data. Moreover, our GPU of choice was the NVIDIA A100 SXM4 40 GB. In few-shot
and zero-shot settings, we adjusted the temperature parameter to &lt;= 0.1 to enhance model
determinism and minimize variation.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Performance Analysis</title>
      <p>4.1. Quantitative Results
Program-Based results: Table 2 provides an overview of our findings obtained through
ifne-tuning and few-shot learning experiments. As observed in the table, increasing the model</p>
      <sec id="sec-4-1">
        <title>Task 1: Perfect-Context</title>
      </sec>
      <sec id="sec-4-2">
        <title>Program P</title>
        <p>Valid? Facts</p>
      </sec>
      <sec id="sec-4-3">
        <title>Task 2: Distractor-Context</title>
      </sec>
      <sec id="sec-4-4">
        <title>Program P</title>
        <p>Valid? Facts</p>
      </sec>
      <sec id="sec-4-5">
        <title>Task 3: Full</title>
      </sec>
      <sec id="sec-4-6">
        <title>Program P</title>
        <p>Valid?
size from small to base resulted in performance improvements for both the T5 and FLAN-T5
models. Notably, FLAN-T5 achieved the highest performance in task-1 with a precision score of
0.49, while in task-2, T5-base obtained the highest score (0.28) among all fine-tuning settings.
However, it is important to highlight that task 3 yielded comparatively lower results across all
ifne-tuning settings. Of particular interest is the performance in the 5-shot learning scenario,
where GPT-3.5-turbo surpassed all fine-tuned models and achieved a new state-of-the-art
benchmark on all three tasks. Another noteworthy observation relates to the "Valid?" score in
Task 3 when compared to Task 1 and Task 2 in the Fine-tuning setting. In most instances, the
model tends to generate more valid programs in Task 3. This can be attributed to the model’s
greater freedom in generating its own approach, which results in more preferred ways of
generating solutions, often leading to shorter answers and reducing the likelihood of producing
invalid responses.</p>
        <p>GPT-4
GPT-3.5-Turbo
FLAN-T5-XL</p>
      </sec>
      <sec id="sec-4-7">
        <title>Task 1: Perfect-Context</title>
        <p>Ans 
0.66
0.72
0.34</p>
      </sec>
      <sec id="sec-4-8">
        <title>Task 2: Distractor-Context Ans</title>
        <p>0.55
0.46
0.12</p>
        <p>Direct-Answer-based results As previously mentioned, attempting to generate the
executable Program P without providing any examples to LLMs is not practically achievable in a
zero-shot context. Therefore, in this section, we evaluate the results based on Ans . Table 3
showcases significant findings in this scenario, with scores of 0.72 and 0.55 for tasks 1 and 2,
respectively. These scores suggest that LLMs can ofer reasonably accurate answers through
implicit reasoning when provided with useful information(either with only pertinent
information or in conjunction with distractors.). However, they struggle to clearly explain how they
arrived at their answers (see the superior results in table 3 compared to table 2). Regardless
of the output format, it becomes evident that the presence of distractor information can pose
challenges for models and adversely afect their performance. Importantly, the results indicate
that the provision of relevant information can enable models to generate estimations, whereas
the absence of such knowledge reduces LLMs to a trivial baseline level. In essence, [17] reported
a constant model that predicts a random value (a logarithmic sweep between 1010 and 10− 10),
which can achieve a result of 0.22 of Ans .
4.2. Qualitative Results
Figure 2 illustrates various examples generated by fine-tuned and few-shot models for tasks 1
and 3. In the leftmost example in task 1, we observe that the few-shot model correctly produced
a program, while the fine-tuned model made an error concerning the selection of the appropriate
mathematical operation, mistakenly choosing multiplication instead of division to derive the
answer. Conversely, on the right side of the figure, the opposite situation occurred, where
the few-shot model erroneously included the original question in the decomposition question
and incorporated it into the calculation (Program: div(Q1, Q2, Q3)). More interesting, in task
3 where models have the freedom to forge their unique path of reasoning, leveraging their
supportive facts and calculations, the left example showcases the performance of the few-shot
model. In this task, the few-shot model interprets the question and deduces a suitable estimated
number based on its own knowledge. Notably, in this particular example, the few-shot learner
employs “liters” as a unit of measurement to provide the answer. Conversely, the fine-tuned
model, while correctly identifying the appropriate mathematical operation (division) for the two
facts, falls slightly short in its generation of numbers and decomposition questions. Finally, in
the rightmost example, where the question can be considered more challenging than the other
examples, the fine-tuned model encountered a significant issue. It duplicated its generation,
resulting in an answer that lacked reasonability, essentially attempting to mimic the structure it
learned during the fine-tuning phase. Conversely, the few-shot model demonstrated a more
Task 1
Task 3</p>
        <p>INPUT:
F1: The volume of the swimming pool is</p>
        <p>2500 cubic meters
F2: The volume of a single water drop is</p>
        <p>18e-10 cubic meter
Q: How many drops of water are needed to
fil an Olympic swimming pool?</p>
        <p>PROGRAM:
Q1: What is the volume of an Olympic</p>
        <p>swimming pool?
Q2: What is the volume of a single drop of</p>
        <p>water?
A1: 2500 m**3 A2: 18e-10 m**3
Q2 -&gt; A2 | F2 Q1 -&gt; A1 | F1</p>
        <p>P: Div(Q1, Q2)
Q1: What is the volume of the swimming pool?
Q2: What is the volume of a single water drop?</p>
        <p>A1: 2500 m**3 A2: 18e-10 m**3
Q2 -&gt; A2 | F2 Q1 -&gt; A1 | F1</p>
        <p>P: Mul (Q1, Q2)
Q1: What is the volume of the swimming pool?
Q2: What is the volume of a single water drop?</p>
        <p>A1: 2500 m**3 A2: 18e-10 m**3
Q2 -&gt; A2 | F2 Q1 -&gt; A1 | F1</p>
        <p>P: Div (Q1, Q2)
x</p>
        <p>INPUT:
F1: The total heat generated by a heater</p>
        <p>is 9600 W
F2: The average heat given off by a single</p>
        <p>person is 80 W
Q: How many children would a couple
have to have in order to negate
the need of a heating system
in a modern average-sized</p>
        <p>home?</p>
        <p>PROGRAM:
Q1: What is the total heat needed to generate
to compensate for the need of a heating</p>
        <p>system?
Q2: What is the average heat produced by the</p>
        <p>presence of a single child?
A1: 9600 W A2: 80 W
Q2 -&gt; A2 | F2 Q1 -&gt; A1 | F1</p>
        <p>P: Div (Q1, Q2)
Q1: What is the total heat generated by a heater?
Q2: What is the average heat given off by a single</p>
        <p>person?
A1: 9600 W
Q2 -&gt; A2 | F2</p>
        <p>A2: 80 W</p>
        <p>Q1 -&gt; A1 | F1</p>
        <p>P: Div (Q1, Q2)
Q1: What is the total heat generated</p>
        <p>by a heater?
Q2: What is the average heat given off by a</p>
        <p>single person?
Q3: How many children does a couple
have to have in order to
negate the need of a heating
system in a modern
average-sized home?
A1: 9600 W A2: 80 W
Q2 -&gt; A2 | F2 Q1 -&gt; A1 | F1</p>
        <p>P: Div (Q1, Q2, Q3)
x</p>
        <p>INPUT:
Q: How many jel y beans fil a bucket?</p>
        <p>PROGRAM:
F1: The average volume of a bucket is 0.67</p>
        <p>cubic feet
F2: The average volume of a jel y bean is</p>
        <p>0.00012 cubic feet
Q1: What is the average volume of a bucket?
Q2: What is the average volume of a jel y</p>
        <p>bean?
A1: 0.67 ft**3
Q2 -&gt; A2 | F2</p>
        <p>A2: 0.00012 ft**3</p>
        <p>Q1 -&gt; A1 | F1</p>
        <p>P: Div(Q1, Q2)
F1: The volume of a bucket of jel y beans</p>
        <p>is 62000 cc
F2: The volume of a single jel y bean is</p>
        <p>62000 cm**3
Q1: What is the volume of a single bucket of</p>
        <p>jel y beans?
Q2: What is the volume of a single jel y bean?
A1: 62000 cm**3
Q2 -&gt; A2 | F2</p>
        <p>A2: 62000 cm**3</p>
        <p>Q1 -&gt; A1 | F1
P: Div(Q1, Q2)</p>
        <p>x
F1: The volume of a bucket is 5 liters
F2: The volume of a jel y bean is 0.001 liters
Q1: What is the volume of a bucket?
Q2: What is the volume of a single jel y bean?</p>
        <p>A1: 5 liters A2: 0.001 liters
Q2 -&gt; A2 | F2 Q1 -&gt; A1 | F1</p>
        <p>P: Div(Q1, Q2)
blue components represent answers obtained from a 5-shot GPT-3.5-turbo, and the orange components
show the answers generated by the fine-tuned FLAN-T5-base model. The ground truth for each example
is depicted in gray.
meaningful expansion of its decomposition, even though it failed to identify a correct path to
the original question, yielding an incorrect result. Based on these examples, it becomes evident
that models, especially larger ones like GPT-based models, have the capacity to retrieve relevant
information, but the relationship between such information requires further investigation.
4.3. Fine-tuning vs Few-shot Learning
compared to the few-shot model across all tasks (denoted as Fp’s tasks). In Task 1, depicted
as the leftmost image, where all relevant facts are provided, the fine-tuned model generated
programs with hop counts that closely matched the two most frequent counts (specifically, 2-hop
and 4-hop). This suggests that fine-tuning may be influenced by an unbalanced distribution of
examples and may tend to produce the most common strategy learned during the fine-tuning
phase. Furthermore, this tendency can explain the high value of the “valid?” score (0.86) for
this model. In contrast, the few-shot model appears to mimic the ground-truth behavior more
closely. For instance, in Task 1, the few-shot model generates answers with hop counts of 1 and
5, which the fine-tuned model did not produce. Interestingly, in Task 2, where the fine-tuned
model displays a similar behavior, the few-shot learning approach generates answers with
higher hop counts (e.g., 7, 8, 9) due to the increased number of facts in this task. Furthermore,
in Task 3, where models are given the freedom to explore their own reasoning methods, both
models exhibit a tendency to generate programs with a lower number of hops. This behavior is
noteworthy and may provide insight into why these models face challenges with this task. The
inclination towards simpler answers in response to complex questions, such as in the case of
Task 3, can potentially lead to suboptimal results.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. LLMs’ Limitations &amp; Possible Solutions</title>
      <p>Facts’ Units Identifiers: One common issue found in LLMs is the mishandling of FP
identiifers. This often results in incorrect responses, where LLMs choose mathematical operations for
facts that lack compatibility or fail to yield a correct answer. For example, in Figure 2 leftmost
hand side, a fine-tuned model selected multiplication despite the query explicitly requesting a
unitless numerical answer. A potential solution involves integrating fact’s unit checking into
the generation process, using a multi-task approach or a non-sequential procedure, as shown in
[19] for addressing equation-based questions with a tree-structured decoder.
Reasoning Deadlock (LLMs’ Predicament): A significant challenge in LLMs is their
performance in Task 3 especially in the Few-shot setting, where they must provide autonomous
reasoning. In this task, LLMs often struggle with their reasoning processes, leading to invalid or
incorrect answers. Common issues, including but not limited to, raising sub-questions without
answers, assigning numeric values to undecomposed questions, failing to link supportive facts
to questions, and condensing multiple mathematical relations into a single equation, such as
“Div(Q1, Mul(Q2, Q3), Q4)” which increases the likelihood of encountering impossible equations.
An efective approach to address this issue involves the deployment of LLMs within iterative
loops, as opposed to requiring them to perform reasoning in a single comprehensive round.
Utilizing LLMs in multiple rounds ensures the consistency and reliability of the generated answers.
This perspective can yield significant benefits, as the model becomes progressively informed
about the components it has expanded upon and those that remain unexplored, resulting in a
more coherent and accurate reasoning process.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Related Work</title>
      <p>Recently, LLMs have emerged as a promising avenue for improving reasoning tasks, particularly
in the context of retrieving relevant and supportive information. For instance, [26] demonstrates
that LLMs possess the ability to discern implicit relations. They achieve this by decoupling the
process of inferring reasoning steps from their execution. This partially aligns with our findings
for task 3 (implicit reasoning) of FPs, where LLMs appear to be more successful at retrieving
information than conducting reasoning over the retrieved information. Additionally, there
have been attempts to enhance existing reasoning and QA tasks by generating intermediate
knowledge to facilitate multi-hop reasoning. In their work, [27, 28] employ LLMs to create
benchmarks, and they find that this technique can improve both the performance of models and
their interpretability. In line with these eforts, other studies [ 29, 30] have demonstrated that
incorporating intermediate supervisory information into the input can enhance the performance
of such models. This aligns with our findings, wherein the fine-tuning of task 1, incorporating
supportive facts concatenated with the input yielded superior results than the absence of such
knowledge. Furthermore, [31] introduced an agent communication mechanism for addressing
complex reasoning questions. In this approach, a model engages in a series of QA interactions
with agents, such as TextQA and TableQA, to arrive at the final answer. However, this approach
has the potential to harness the capabilities of LLMs as agents for solving FPs. Nevertheless,
accomplishing this without auxiliary supervision remains a challenging endeavor, necessitating
significant dataset modifications to adapt it for FPs. In addition, exploring other directions in
the context of complexity is also noteworthy. In terms of complexity, some studies, such as
[32], propose that introducing a progressive task complexity framework can yield advantages
for LLMs. [33] proposed TELeR, a general guideline for LLMs in prompt designing to perform
complex tasks. However, Fermi Problems regard this complexity as comprising three distinct
and isolated tasks. The possibility of merging these complex paradigms represents a promising
avenue for future research, which we defer to further exploration. Finally, similar to our work,
[34] also assessed the capabilities of LLMs in the domain of logical reasoning tasks. They carried
out their experiments in a few-shot learning scenario, utilizing methodologies such as prompt
engineering and CoT. However, it’s important to note that their primary focus was solely on
logical reasoning, whereas our study diverges by primarily delving into mathematical reasoning,
particularly within the context of Fermi Problems.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>This paper explored the performance of LLMs in various FP scenarios, including fine-tuning,
few-shot, and zero-shot settings. Our findings indicate that despite the advancements in LLMs,
there is still a need for further enhancements to enable these models to exhibit creativity at a
level comparable to human capabilities
[17] A. Kalyan, A. Kumar, A. Chandrasekaran, A. Sabharwal, P. Clark, How much cofee was
consumed during emnlp 2019? fermi problems: A new reasoning challenge for ai, arXiv
preprint arXiv:2110.14207 (2021).
[18] J. Wei, X. Wang, D. Schuurmans, M. Bosma, brian ichter, F. Xia, E. H. Chi, Q. V. Le, D. Zhou,
Chain of thought prompting elicits reasoning in large language models, in: A. H. Oh,
A. Agarwal, D. Belgrave, K. Cho (Eds.), Advances in Neural Information Processing Systems,
2022.
[19] Z. Xie, S. Sun, A goal-driven tree-structured neural model for math word problems., in:</p>
      <p>Ijcai, 2019, pp. 5299–5305.
[20] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I.
Polosukhin, Attention is all you need, Advances in neural information processing systems 30
(2017).
[21] H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M.
Dehghani, S. Brahma, et al., Scaling instruction-finetuned language models, arXiv preprint
arXiv:2210.11416 (2022).
[22] T. Schick, H. Schütze, Exploiting cloze questions for few shot text classification and natural
language inference, arXiv preprint arXiv:2001.07676 (2020).
[23] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are
unsupervised multitask learners, OpenAI blog 1 (2019) 9.
[24] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, Y. Iwasawa, Large language models are zero-shot
reasoners, Advances in neural information processing systems 35 (2022) 22199–22213.
[25] D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, arXiv preprint
arXiv:1412.6980 (2014).
[26] U. Katz, M. Geva, J. Berant, Inferring implicit relations with language models, arXiv
preprint arXiv:2204.13778 (2022).
[27] E. Zelikman, J. Mu, N. D. Goodman, Y. T. Wu, Star: Self-taught reasoner bootstrapping
reasoning with reasoning (2022).
[28] J. Welbl, P. Stenetorp, S. Riedel, Constructing datasets for multi-hop reading comprehension
across documents, Transactions of the Association for Computational Linguistics 6 (2018)
287–302.
[29] N. Wies, Y. Levine, A. Shashua, Sub-task decomposition enables learning in sequence to
sequence tasks, arXiv preprint arXiv:2204.02892 (2022).
[30] G. Recchia, Teaching autoregressive language models complex tasks by demonstration,
arXiv preprint arXiv:2109.02102 (2021).
[31] T. Khot, K. Richardson, D. Khashabi, A. Sabharwal, Hey ai, can you solve complex tasks
by talking to agents?, arXiv preprint arXiv:2110.08542 (2021).
[32] M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan,
A. Lewkowycz, M. Bosma, D. Luan, et al., Show your work: Scratchpads for intermediate
computation with language models, november 2021, URL http://arxiv. org/abs/2112.00114
(2021).
[33] S. K. K. Santu, D. Feng, Teler: A general taxonomy of llm prompts for benchmarking
complex tasks, arXiv preprint arXiv:2305.11430 (2023).
[34] A. Creswell, M. Shanahan, I. Higgins, Selection-inference: Exploiting large language
models for interpretable logical reasoning, arXiv preprint arXiv:2205.09712 (2022).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bassignana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramponi</surname>
          </string-name>
          , Preface to the
          <source>Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI)</source>
          ,
          <source>in: Proceedings of the Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI</source>
          <year>2023</year>
          )
          <article-title>co-located with 22th International Conference of the Italian Association for Artificial Intelligence (AI* IA</article-title>
          <year>2023</year>
          ),
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Fagin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Halpern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Moses</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vardi</surname>
          </string-name>
          ,
          <article-title>Reasoning about knowledge</article-title>
          , MIT press,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. C.-C. Chang</surname>
          </string-name>
          ,
          <article-title>Towards reasoning in large language models: A survey</article-title>
          ,
          <source>arXiv preprint arXiv:2212.10403</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Kolodner</surname>
          </string-name>
          ,
          <article-title>An introduction to case-based reasoning</article-title>
          ,
          <source>Artificial intelligence review 6</source>
          (
          <year>1992</year>
          )
          <fpage>3</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Prakken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sartor</surname>
          </string-name>
          ,
          <article-title>A dialectical model of assessing conflicting arguments in legal reasoning, Logical models of legal argumentation (</article-title>
          <year>1997</year>
          )
          <fpage>175</fpage>
          -
          <lpage>211</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ghazvininejad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          , L. Zettlemoyer, Bart:
          <article-title>Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension</article-title>
          , arXiv preprint arXiv:
          <year>1910</year>
          .
          <volume>13461</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>
          ,
          <source>The Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>5485</fpage>
          -
          <lpage>5551</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Subbiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Neelakantan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          , et al.,
          <article-title>Language models are few-shot learners</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Izacard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Martinet</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Rozière</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Hambro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Azhar</surname>
          </string-name>
          , et al.,
          <article-title>Llama: Open and eficient foundation language models</article-title>
          ,
          <source>arXiv preprint arXiv:2302.13971</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Valmeekam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Olmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sreedharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kambhampati</surname>
          </string-name>
          ,
          <article-title>Large language models still can't plan (a benchmark for llms on planning and reasoning about change)</article-title>
          ,
          <source>arXiv preprint arXiv:2206.10498</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ullman</surname>
          </string-name>
          ,
          <article-title>Large language models fail on trivial alterations to theory-of-mind tasks</article-title>
          ,
          <source>arXiv preprint arXiv:2302.08399</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Helwe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Clavel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Suchanek</surname>
          </string-name>
          ,
          <article-title>Reasoning with transformer-based models: Deep learning, but shallow reasoning</article-title>
          ,
          <source>in: International Conference on Automated Knowledge Base Construction (AKBC)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Misra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Scales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Chi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schärli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Large language models can be easily distracted by irrelevant context</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>31210</fpage>
          -
          <lpage>31227</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mihaylov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Khot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sabharwal</surname>
          </string-name>
          ,
          <article-title>Can a suit of armor conduct electricity? a new dataset for open book question answering</article-title>
          , arXiv preprint arXiv:
          <year>1809</year>
          .
          <volume>02789</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rastogi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A. M.</given-names>
            <surname>Shoeb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Santoro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Garriga-Alonso</surname>
          </string-name>
          , et al.,
          <article-title>Beyond the imitation game: Quantifying and extrapolating the capabilities of language models</article-title>
          ,
          <source>arXiv preprint arXiv:2206.04615</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>