<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Models for the Symbol Grounding Task in ABC Repair System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pak Yin Chan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xue Li</string-name>
          <email>xue.shirley.li@ed.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan Bundy</string-name>
          <email>A.Bundy@ed.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Informatics, The University of Edinburgh</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The ABC Theory Repair System (ABC) has demonstrated its success in facilitating users to repair faulty theories utilizing distinct techniques. Yet, comprehending ABC-repaired theories becomes more challenging due to the presence of dummy constants or predicates introduced by ABC. In this paper, we propose a grounding system by incorporating Large Language Models (LLMs) to provide these dummy items with meaningful names. By applying ABC and grounding alternately, the resulting theory is both fault-free and semantically meaningful. Moreover, our study shows that LLMs without fine-tuning still exhibit capabilities of common knowledge, and their grounding performances are enhanced by providing suficient background or asking for more returns. Large language model, Closed-book question answering, Faulty logical theory repair, Automated theorem Logical theory stands as a reasoning tool in the field of artificial intelligence (AI), representing structured and precise representations of relationships among objects [1]. The theory needs to be refined to cope with new observations [ 2]. When users introduce novel information, the original theory may either make incorrect predictions or fail to predict the expected truth. [3].</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>CEUR
Workshop
Proceedings
dummy items. In the previous example, users can deduce that “dummyConst1” and
“dummyConst2” are referring to a birth mother and a stepmother respectively. Yet, there are instances
when users face uncertainty in naming these items. Since the current implementation of ABC
does not encompass the consideration of their semantic implications, numerous generated
repaired theories might be logically consistent but devoid of semantic meaning.</p>
      <p>Example 1 : A Comparison of Original and Repaired Motherhood Theory
Original Theory:
( ,  ) ∧ ( ,  ) ⟹  = 
Repaired Theory:
( ,  ,  1
⟹ ( ,   )
⟹ (,   )
) ∧ ( ,  ,</p>
      <p>
        The challenge of attributing meanings to meaningless symbols is known as the “symbol
grounding problem”, an important problem in Cognitive Science [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. To tackle this problem
automatically, we require tools with access to common-sense knowledge, enabling them to
suggest potential names for users to consider. Although Large Language Models (LLMs) exhibit
inconsistencies in reasoning [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], studies indicate that they store extensive relational knowledge
of the training datasets during pretraining [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10, 11, 12</xref>
        ]. This suggests that we can leverage LLMs
to propose possible meanings for dummy items by presenting propositions involving these
items in natural language.
      </p>
      <p>
        Our study demonstrates an application of LLMs to solve the symbol grounding problem.
The primary objective of this paper is to enrich the semantic content of the repaired theory by
utilizing LLMs to replace the names of dummy items with semantically meaningful content. We
hypothesize that the closed-book question answering (CBQA) task [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] with LLMs helps to
conduct the symbol grounding challenge within ABC, which in the CBQA task, LLMs generate
responses based on their training data solely, without having access to external sources [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ].
To explore how well can LLMs provide meanings of dummy items in the repaired theory by
the ABC, as a way of enhancing the semantics of the repaired theory, we propose a system
of symbol grounding for ABC to determine the meanings of dummy items that involve user
interactivity.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. LLM-grounding system</title>
      <p>
        Figure 1 illustrates the process employed by the grounding system. We first parse the input
theory in Datalog, a subset of First Order Logic (FOL) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and the system sets up records of
constants and predicates (G1). These records facilitate the detection of ungrounded dummy
items. For each dummy item, the system interprets the associated axioms into a natural language
question (G2). After grounding with the LLM that users choose (G3), the system presents all
available choices by that LLM. Users can choose the LLM-suggested answers or suggest new
answers by themselves (G4a). After that, the system also recommends users use the existing
theory items with high similarity to any previously suggested answers (G4b). Once all dummy
items are successfully grounded, the system proceeds to replace these items with the selected
answers (G5) and exports all possible grounded theories in Datalog [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        Before grounding starts, ABC detects and repairs the fault in the input theory  when it
conflicts with the given preferred structure ℙ , which represents users’ observations [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Once
the repair is done, users need to manually copy a repaired theory to start the grounding process.
      </p>
      <p>
        We design a heuristic to formulate a prompt in the ungrounded theory. We suspect that
LLMs perform grounding more accurately if we provide suficient knowledge in the prompt, so
we contain assertions without dummy items and multiple axioms with the same dummy item
in the same prompt. For each proposition with a maximum arity of 3 in axioms, we convert it
into natural language with the interpretation in Table 2 in the Appendix. The interpretation is
similar to [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], except we substitute the dummy name by the item’s type - “property” for the
dummy predicate, and “entity”/“kind” for the dummy constant. We gather the propositions
into rules using conditional sentences and append the specification of the word limit at the end
of the prompt to avoid LLMs returning lengthy answers. Example 2 illustrates the resulting
prompt for grounding  1 using the above heuristics, with setting the word limit as 5.
      </p>
      <p>For each grounding of dummy items, users are presented with one to two rounds of
recommendations. The initial recommendations are from the “LLM-grounding Recommender”,
a phase where the LLM’s suggestions are displayed. Users are provided with the option to
directly select the LLM-suggested answers, retain the dummy name, or propose new names.
The inclusion of the latter choice allows users to refine the grounding names based on the
LLM-generated answers or to tailor them to their preferences. In the subsequent phase
“Existeditems Recommender”, the system explores the presence of existing items within the theory
that exhibit high similarity to the chosen or suggested groundings from the previous phase.
This comparison process involves assessing the resemblance of all prior suggestions against
constants or predicates within the theory, depending on the type of dummy item. This phase
employs the F1 score of BERTScore vanilla (referred to as F1 BERTScore) [15] to gauge the
similarity between items. BERTScore utilizes embeddings from the pre-trained BERT model,
calculating the cosine similarity of embeddings to measure word matches between candidates
and references [15]. Higher scores correspond to more significant similarity. Users also have
the flexibility to set a threshold for the F1 BERTScore, enabling the system to recommend items
that surpass the specified F1 BERTScore.</p>
      <p>Example 2: Prompt Formulating in Repaired Tweety Theory
 ( ,</p>
      <p>aNotice that the names of the penguins Tweety and Polly are not capitalized in the prompt.</p>
      <p>An important consideration is that the grounding process has the potential to reintroduce
faults into repaired theories. In response, we incorporate a safeguard as an extra feature
by aiding users in re-running ABC following the exportation of all grounded theories. This
validation step assesses whether the theories remain free from faults. This iterative approach
involving repair and grounding persists until users attain a satisfactory theory. A practical
example of the interplay between repair and grounding is depicted in the appendix.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Grounding Performance</title>
      <p>
        We experimented with GPT-3.5 Turbo (4K context version) and GPT-4 (8K context version)
[
        <xref ref-type="bibr" rid="ref10">10, 16</xref>
        ] using OpenAI’s ChatComplete API without further fine-tuning. As ABC is a
domainindependent repair system, we intentionally omitted both fine-tuning and few-shot learning to
gauge how these models perform without such adjustments.
      </p>
      <p>As ABC uses Datalog, a subset of FOL, the grounding system cannot be evaluated with major
FOL datasets [17, 18]. We compromised to examine the performance of LLMs. We substituted
some items with dummy names in the theories and studied if the LLMs could ground similar
items as the original ones in our system. We constructed theories automatically from two
knowledge bases, enriched WebNLG dataset [19] and excerpt of DART [20], and replaced
some items with dummy names. These theories serve as simulations of the generated repaired
theories, with both assertions and rules. Details of the construction of the evaluation dataset
are in the project’s GitHub repository1.</p>
      <p>We adopted two semantic-based metrics, F1 BERTScore [15] and SAS [21], to evaluate the
semantic similarity of answers in LLM-grounding Recommender, as they are shown to have
1https://github.com/HistoChan/ABCGrounding
a certain correlation between human judgement [21]. The former one is the same as the one
used in the “Existed-items Recommender”. SAS, however, uses a pre-trained cross-encoder and
applies the model by concatenating two texts with a separator token in between. Diferent
from BERTScore, SAS considers two inputs together [21]. We calculated the scores of the LLM
outcomes with the original items’ names. All the metrics values range from 0 to 1, with values
closer to 1 indicating greater semantic similarity between candidates and references.</p>
      <p>Prompt content</p>
      <p>number of output
basic
w/ background
w/ multi axioms
w/ both
w/ both</p>
      <p>Table 1 is the statistics of the experiment, which shows that an LLM without any fine-tuning
can still have an adequate grounding performance in zero-shot. The increase in the metrics
confirms that the inclusion of additional background information in the prompt can obtain
higher-quality grounding outcomes. The utilization of background information emerges as a
more impactful hint for successful grounding, surpassing the efectiveness of querying multiple
axioms in a single prompt. Moreover, the performances generally enlarge with model size, and
an increase in the number of groundings would correspondingly enhance overall performance.
We also adopted other models such as T5 models by Google [22], OpenLLaMA models from
OpenLM Research [23], and Dolly 2.0 models by Databricks [24]. Their performances also
support the above statements, which the project’s GitHub repository1 contains the statistics of
their performances. Some case studies are in the appendix.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusions</title>
      <p>In this paper, we have proposed a system of symbol grounding for the ABC repair system.
We formulate the grounding challenge into a CBQA task and require LLMs to return possible
answers. The system also embraces user interactivity, in which users have a right to control the
model use, types of formatting prompts and grounding results. This system not only helps to
enhance the semantics in the repaired theory but also determines the rationality of the repair
plan. Yet, we suspect that the grounding performance is limited by the quality of the prompt,
which can be improved in the future. Additionally, we facilitate using LLMs without either
ifne-tuning or few-shot learning for CBQA tasks by providing suficient background information
and enhancing the number of outputs. Nonetheless, we do not deem that few-shot learning can
be replaced by providing background information. It is worth studying if few-shot learning can
achieve similar performance of fine-tuning, and if the performance enhancement with few-shot
learning is limited by the scale of the model.
[15] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, Y. Artzi, BERTScore: Evaluating Text
Generation with BERT, arXiv preprint arXiv:1904.09675 (2019). URL: https://github.com/
Tiiiger/bert.
[16] OpenAI, GPT-4 Technical Report, arXiv preprint arXiv:2303.08774 (2023). URL: http:
//arxiv.org/abs/2303.08774.
[17] S. Han, H. Schoelkopf, Y. Zhao, Z. Qi, M. Riddell, L. Benson, L. Sun, E. Zubova, Y. Qiao,
M. Burtell, D. Peng, J. Fan, Y. Liu, B. Wong, M. Sailor, A. Ni, L. Nan, J. Kasai, T. Yu,
R. Zhang, S. Joty, A. R. Fabbri, W. Kryscinski, X. V. Lin, C. Xiong, D. Radev, FOLIO: Natural
Language Reasoning with First-Order Logic, arXiv preprint arXiv:2209.00840 (2022). URL:
http://arxiv.org/abs/2209.00840.
[18] J. Tian, Y. Li, W. Chen, L. Xiao, H. He, Y. Jin, Diagnosing the First-Order Logical Reasoning
Ability Through LogicNLI, Proceedings of the 2021 Conference on Empirical Methods in
Natural Language Processing (2021) 3738–3747.
[19] T. C. Ferreira, D. Moussallem, S. Wubben, E. Krahmer, Enriching the WebNLG corpus, in:
Proceedings of the 11th International Conference on Natural Language Generation, 2018,
pp. 171–176. URL: http://data.statmt.org/wmt17_systems.
[20] L. Nan, D. Radev, R. Zhang, A. Rau, A. Sivaprasad, C. Hsieh, X. Tang, A. Vyas, N. Verma,
P. Krishna, Y. Liu, N. Irwanto, J. Pan, F. Rahman, A. Zaidi, M. Mutuma, Y. Tarabar, A. Gupta,
T. Yu, Y. C. Tan, X. V. Lin, C. Xiong, R. Socher, N. F. Rajani, DART: Open-Domain Structured
Data Record to Text Generation, in: Proceedings of the 2021 Conference of the North
American Chapter of the Association for Computational Linguistics: Human Language
Technologies, Association for Computational Linguistics, Online, 2021, pp. 432–447. URL:
https://aclanthology.org/2021.naacl-main.37. doi:10.18653/v1/2021.naacl- main.37.
[21] J. Risch, T. Möller, J. Gutsch, M. Pietsch, Semantic Answer Similarity for Evaluating
Question Answering Models, arXiv preprint arXiv:2108.06130 (2021). URL: http://arxiv.
org/abs/2108.06130.
[22] C. Rafel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu,
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Journal
of Machine Learning Research 21 (2020) 1–67. URL: http://jmlr.org/papers/v21/20-074.html.
[23] X. Geng, H. Liu, OpenLLaMA: An Open Reproduction of LLaMA, 2023. URL: https://github.</p>
      <p>com/openlm-research/open_llama.
[24] M. Conover, M. Hayes, A. Mathur, X. Meng, J. Xie, J. Wan, S. Shah, A. Ghodsi,
P. Wendell, M. Zaharia, R. Xin, Free Dolly: Introducing the World’s First Truly
Open Instruction-Tuned LLM, 2023. URL: https://www.databricks.com/blog/2023/04/12/
dolly-first-open-commercially-viable-instruction-tuned-llm.
Example 3 (Continue): Repair and Grounding a Capital Theory
Step 2:
land,
in 
 ( ,  ) ∧  ( ,  )</p>
      <p>ABC finds there is a fault in having two capitals in
Scotwhich it suggests replacing “capitalOf” with “dummyPred”</p>
      <p>( , ) , and it is grounded as ‘cityOf‘:
⟹  = 
⟹  ( ℎ,
⟹  
⟹  (, )
( , )

)</p>
    </sec>
    <sec id="sec-6">
      <title>C. Examples of Grounding Results</title>
      <p>We compare the performance of grounding containing background information of two models
in Example 4, which illustrates that the grounding would be more reasonable with providing
background information.</p>
      <p>Example 4: A comparison of grounding answers of Example 2
• Without extra content: What is a possible entity, such that opus is broken wing of
the entity? Answer within 5 wordsa.
• With background information: Given that opus is super penguin. What is a
possible entity, such that opus is broken wing of the entity? Answer within 5
wordsa.</p>
      <p>– GPT-3.5 Turbo: Defective wing.
– T5 Large (NQ): Feathers are
– GPT-3.5 Turbo: X is injured
– T5 Large (NQ): Cannot fly
aNotice that the name of the penguin Opus is not capitalized in the prompt.</p>
      <p>We provide Example 5 as a comparison of grounding performance on diferent LLMs, in
which the answers are more accurate with larger model sizes. Despite our explicit instruction
to return answers within five words and request of the maximum output tokens as five, Dolly
2.0 and OpenLLaMA still have a high tendency to return a complete sentence.</p>
      <p>Example 5: A comparison of grounding answers of a repaired Capital Theory
Question: What is a possible entity such that edinburgh is cap of of the entity? Answer
within 5 words.</p>
      <p>• Dolly 2.0 3B: Edinburgh is the capital of
• Dolly 2.0 7B: The answer is the Edinburgh
• OpenLLaMA 3B: edinburgh is cap of
• OpenLLaMA 7B: The answer is Scotland.
• GPT-3.5 Turbo &amp; T5 XL (NQ): Scotland
• GPT-4: Scotland or United Kingdom.
• T5 Small &amp; Large : Edinburgh
• T5 Small (NQ): other social entity
• T5 Large (NQ): Kingdom of Scotland</p>
      <p>We experimented on the efect of the number of groundings generated from LLM. This
experiment lay in the diversity of the returned results. With open-ended questions like that in
Example 6, the answers returned reflect the multifaceted nature of potential responses. The
augmented number of returned answers not only aids in identifying high-quality grounding
but also empowers users to brainstorm a broader spectrum of possible groundings.</p>
      <p>Example 6: A comparison of answers of prompt from a repaired Tweety Theory
Question: Given that tweety is penguin. What is a possible entity such that polly is bird
of the entity, and In a FOL expression, if x is bird of the entity, then x is fly? Answer
within 5 words.a.</p>
      <p>• GPT-3.5 Turbo: “Airplane.”, “Flying creature like parrot”, “Fish”
• GPT-4: “Sky or Air could be”, “Possible entity is ’the”, “The possible entity: magical”
• Suggested Answer: “flying”</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Barr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Feigenbaum</surname>
          </string-name>
          ,
          <source>The handbook of artificial intelligence,</source>
          volume
          <volume>1</volume>
          ,
          <string-name>
            <surname>Butterworth-Heinemann</surname>
          </string-name>
          ,
          <year>1981</year>
          . URL: https://www.sciencedirect.com/science/article/ pii/B9780865760899500089. doi:https://doi.org/10.1016/B978-0
          <source>-86576-089-9</source>
          .
          <fpage>50008</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bundy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Representational change is integral to reasoning</article-title>
          ,
          <source>Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences</source>
          <volume>381</volume>
          (
          <year>2023</year>
          )
          <article-title>20220052</article-title>
          . URL: https://royalsocietypublishing.org/doi/abs/10.1098/rsta.
          <year>2022</year>
          .
          <volume>0052</volume>
          . doi:
          <volume>10</volume>
          .1098/rsta.
          <year>2022</year>
          .
          <volume>0052</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <source>Automating the Repair of Faulty Logical Theories</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sakama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Inoue</surname>
          </string-name>
          ,
          <article-title>An abductive framework for computing knowledge base updates</article-title>
          ,
          <source>Theory and Practice of Logic Programming</source>
          <volume>3</volume>
          (
          <year>2003</year>
          ). doi:
          <volume>10</volume>
          .1017/S1471068403001716.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. O.</given-names>
            <surname>Hansson</surname>
          </string-name>
          ,
          <article-title>Ten philosophical problems in belief revision</article-title>
          ,
          <source>Journal of Logic and Computation</source>
          <volume>13</volume>
          (
          <year>2003</year>
          ). doi:
          <volume>10</volume>
          .1093/logcom/13.1.37.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bundy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mitrovic</surname>
          </string-name>
          , Reformation:
          <string-name>
            <given-names>A</given-names>
            <surname>Domain-Independent Algorithm for Theory Repair</surname>
          </string-name>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Harnad</surname>
          </string-name>
          ,
          <source>The Symbol Grounding Problem</source>
          ,
          <year>1990</year>
          . URL: http://cogprints.org/3106/.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Lyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Havaldar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Apidianaki</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>CallisonBurch, Faithful Chain-of-Thought Reasoning</article-title>
          , arXiv preprint arXiv:
          <volume>2301</volume>
          .13379 (
          <year>2023</year>
          ). URL: http://arxiv.org/abs/2301.13379.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schuurmans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bosma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ichter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Chi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , Chainof-Thought
          <source>Prompting Elicits Reasoning in Large Language Models, Advances in Neural Information Processing Systems</source>
          <volume>35</volume>
          (
          <year>2022</year>
          )
          <fpage>24824</fpage>
          -
          <lpage>24837</lpage>
          . URL: http://arxiv.org/abs/2201. 11903.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ramesh</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hesse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , E. Sigler,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Litwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chess</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
          </string-name>
          , Language Models are
          <string-name>
            <surname>Few-Shot</surname>
            <given-names>Learners</given-names>
          </string-name>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          . URL: http://arxiv.org/abs/
          <year>2005</year>
          .14165.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Petroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rocktäschel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bakhtin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Riedel</surname>
          </string-name>
          ,
          <article-title>Language Models as Knowledge Bases?</article-title>
          , arXiv preprint arXiv:
          <year>1909</year>
          .
          <volume>01066</volume>
          (
          <year>2019</year>
          ). URL: https://github. com/pytorch/fairseq.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <article-title>How Much Knowledge Can You Pack Into the Parameters of a Language Model?</article-title>
          ,
          <source>EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference</source>
          (
          <year>2020</year>
          )
          <fpage>5418</fpage>
          -
          <lpage>5426</lpage>
          . URL: https://arxiv. org/abs/
          <year>2002</year>
          .08910v4. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .emnlp-main.
          <volume>437</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ceri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gottlob</surname>
          </string-name>
          , L. Tanca,
          <article-title>What you always wanted to know about Datalog (and never dared to ask)</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>1</volume>
          (
          <year>1989</year>
          )
          <fpage>146</fpage>
          -
          <lpage>166</lpage>
          . doi:
          <volume>10</volume>
          .1109/69.43410.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mpagouli</surname>
          </string-name>
          ,
          <article-title>Converting First Order Logic into Natural Language: A First Level Approach</article-title>
          , in: Current Trends in
          <source>Informatics: 11th Panhellenic Conference on Informatics, PCI</source>
          ,
          <year>2007</year>
          , pp.
          <fpage>517</fpage>
          -
          <lpage>526</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>