<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Explainable Artificial Intelligence and Reasoning in the Context of Large Neural Network Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefanie Krause</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Harz University of Applied Sciences</institution>
          ,
          <addr-line>Friedrichstraße 57-59, Wernigerode, 38855</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The current generation of artificial intelligence (AI) systems ofers tremendous benefits, but their efectiveness is limited by the inability of the machine to explain its decisions and actions to users. My dissertation will delve into the subject of explainable AI (XAI), exploring its various aspects and implications. I will focus on post-hoc local explainability of large AI models to provide human-readable explanations making users understand the automated decision-making of complex models. I aim to evaluate current large language models (LLMs) like ChatGPT or Llama on explainability and its implications, e.g., on education. My thesis also focuses on questions in the field of reasoning with LLMs, since reasoning is fundamental in LLMs for enhancing their understanding and generation of text, improving problem-solving capabilities, and facilitating natural human interaction. However, reasoning is a very dificult task for a computer and the capacities of LLMs regarding diferent reasoning tasks are not yet fully examined.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;explainable AI</kwd>
        <kwd>reasoning</kwd>
        <kwd>large language models</kwd>
        <kwd>education</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Motivation</title>
      <p>
        Reasoning stands as a core component of human intelligence, essential for problem-solving,
making decisions, and critical thinking. Recently, advancements in large language models
have made significant progress in the field of NLP, suggesting that these models might possess
reasoning capabilities, especially as they increase in size [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Nonetheless, the full capacity of
LLMs to reason efectively remains a subject of ongoing debate, which I want to explore further.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        At the moment LLMs like ChatGPT are dominating AI and achieved remarkable results in various
tasks. In the education sector for example, students use this new technology for assignment
writing among other tasks [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The main advantage of LLM is that one can easily generate
natural language explanations for any QA task. We aim for explainability without a loss in
performance and maybe even improve performance by using explanations during the training
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Explanations are a way to verbalize the reasoning that the models learn during training
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Rajani at al. developed a new dataset called Common Sense Explanations (CoS-E) and the
Commonsense Auto-Generated Explanations (CAGE) framework and improved the accuracy
of language models for commonsense reasoning [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Commonsense reasoning is a dificult
challenge for a computer to handle [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. There is a big gap between the logical approach with
deductive reasoning and the inductive, associative, and empirical nature of human reasoning,
rooted in past experiences. Intriguingly, LLMs lack explicit semantic knowledge, grammatical
structures, or logical rules essential for explicit reasoning, not to mention large-scale ontologies
found in logical knowledge bases like Adimen-SUMO [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. A potential solution could lie in
training neural networks to explicitly learn reasoning, possibly by focusing on certain sentence
forms as in syllogistic reasoning may be implemented with neural-symbolic cognitive reasoning
by specifically structured neural networks [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ]. LLMs can further be utilized to translate
a natural language problem into a symbolic formulation [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. A novel prompting technique,
known as chain-of-thought prompting, has emerged recently. This method aids LLMs in
tackling reasoning challenges by directing them to generate a series of intermediate steps before
providing the final answer [ 12]. [13] recently enhanced previous work by presenting REFINER,
a system designed to enhance LLMs by training them to produce intermediate reasoning steps
through interaction with a critic model that ofers automated feedback on the reasoning process.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Research Questions</title>
      <p>According to the previous context and motivation the primary question of my research is: To
what extent can users and developers benefit from post-hoc local explainability of
large AI models?</p>
      <sec id="sec-3-1">
        <title>In a recent paper, I answered the two following research questions:</title>
        <p>• Can LLMs like ChatGPT handle commonsense reasoning in question answering tasks
with near-human-level performance?
• Are LLMs like ChatGPT able to generate good, human-understandable explanations for
their decisions?
In further research, I aim to study the syllogistic reasoning abilities of LLMs as deductive
reasoning is very diferent from the inductive reasoning that is used for commonsense reasoning.
Moreover, the goal is to analyse the impact of few-shot learning and chain-of-thought as well
as the potential of Retrieval-Augmented Generation (RAG). RAG ofers a way to optimize the
output of an LLM with specific information without changing the underlying model itself.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Research Approach</title>
      <p>My objective is to rigorously test various LLMs across a carefully chosen set of reasoning
tasks, aiming to compare these models’ performances against that of humans. Furthermore,
I plan to examine explanations of LLMs for reasoning tasks firstly by humans and secondly
with an automated evaluation mechanism employing diverse scoring metrics to assess the
interpretability and coherence of these models’ reasoning capabilities. This comprehensive
evaluation strategy seeks to illuminate the strengths and inadequacies of LLMs in mimicking
human-like reasoning and understanding.</p>
      <p>In my initial study on this subject, I delved into commonsense reasoning, using the LLM
ChatGPT to evaluate 11 benchmark datasets. Employing a questionnaire, I compared ChatGPT’s
responses against those of human participants, additionally questioning their evaluations of
the explanations provided by the LLM. Later I broadened this inquiry by including additional
open-source LLMs, such as Llama-3 by Meta and Gemma by Google, aiming for a deeper
comprehension of LLMs’ reasoning competencies. This endeavour focuses on contrasting the
reasoning explanations generated by various LLMs for the same tasks.</p>
      <p>Beyond commonsense reasoning, my research intends to explore syllogistic reasoning with
diferent LLMs to analyse the field of deductive reasoning. Moreover, I’m intrigued by the
potential of employing chain-of-thought prompting in reasoning tasks, a method that could
further elucidate how LLMs navigate complex reasoning processes. This multifaceted approach
promises to shed light on LLMs’ reasoning abilities, paving the way for advancements.</p>
      <p>To mitigate issues like hallucinations and allow better customization and scalability across
various applications I aim to study RAG. I believe RAG can enhance the models’ ability to
provide precise, up-to-date, and contextually relevant information.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Preliminary Results</title>
      <p>In this section, I present the results of my paper Commonsense Reasoning and Explainable
Artificial Intelligence Using Large Language Models [14] presented at the European Conference
on Artificial Intelligence 2023.</p>
      <p>We analysed 11 benchmark datasets specifically curated to challenge solvers lacking
commonsense knowledge. We randomly select 30 examples from each dataset. These tasks span
diverse domains, including medicine, physics, and scenarios from daily life. Our evaluation
of ChatGPT’s capabilities using these QA benchmarks reveals a spectrum of performance
outcomes. Notably, ChatGPT’s weakest performance was observed on the CommonsenseQA
dataset, where it achieved an accuracy of 56.67%, while its strongest performance was recorded
on the Story Cloze Test, reaching an impressive accuracy of 93.33%. A detailed representation of
the performance on each of the eleven datasets is shown in Table 1. Over all datasets ChatGPT
answered with an accuracy of 73.33%, 77 tasks were answered incorrectly (23.33%), and we
did not get a valid response for 11 QA tasks (3.33%). Not valid means that ChatGPT does not
respond which answer option is correct and instead asks for further context information. We
conducted an error analysis and found that there are six kinds of problems where Chat- GPT
still struggles with:</p>
      <sec id="sec-5-1">
        <title>1. missing context</title>
        <p>2. comparative reasoning
3. subjective reasoning
4. slang, unoficial abbreviations, and youth language
5. social situations
6. medical domain
In our extensive questionnaire with 49 participants we found that the participants answered
73.72% of the 20 QA tasks correctly compared to ChatGPT’s 90.00% on the same questions.</p>
        <p>Figure 2 compares ChatGPT’s performance to that of surveyed participants across the datasets.
ChatGPT outperformed humans on six datasets, while humans excelled in four, notably
struggling with CommonsenseQA where ChatGPT also had its lowest performance. The biggest
performance gap was observed in the COPA and Cosmos QA datasets, with humans
outperforming ChatGPT by 26.47% in COPA and ChatGPT surpassed humans by 19.53% in Cosmos QA.
Interestingly, ChatGPT showed strength in Cosmos QA, which requires contextual
commonsense reasoning, despite humans significantly outperforming ChatGPT in COPA, which demands
an understanding of cause and efect, and selecting the most plausible option. The findings
suggest ChatGPT struggles with comparative reasoning where multiple plausible options exist,
hinting that traditional explicit reasoning approaches might fare better in such scenarios. We
further evaluated the explanations given by ChatGPT with the help of a questionnaire and
found that explanations were mostly rated “good” or “excellent” with 67.60% and only 42 times
very poor. Explanations were rated “fair” or better with 84.80%. See Figure fig:explanation for
more details.</p>
        <p>The study demonstrates that ChatGPT achieved a 73.33% overall accuracy rate on eleven QA
datasets, which require commonsense reasoning for correct answers. Despite certain issues,
ChatGPT managed to surpass our survey participants in six of ten datasets (excluding the
MedMCQA medical dataset), suggesting that LLMs like ChatGPT are approaching near-human
performance in commonsense reasoning within QA tasks. Additionally, the research also delved
into the explainability of LLMs, a critical facet in addressing the opacity of these black-box
systems. According to our questionnaire, most of ChatGPT’s explanations were rated “good”
or “excellent”, supporting our hypothesis that LLMs are capable of producing high-quality
explanations.</p>
        <p>A recent extension of this work [26], in which further LLMs have been analysed and a larger
human-centred study was conducted, is currently under review.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Next Steps</title>
      <p>My next steps for further research are:
1. Expand the analysis to more LLMs: Broaden the scope of the study to include a variety
of LLMs to evaluate their capabilities in commonsense reasoning. This will provide a
comparative understanding of diferent models’ strengths and weaknesses in this area.
2. Investigate syllogistic reasoning abilities: Conduct in-depth analysis of the syllogistic
reasoning abilities of LLMs. This involves assessing how well these models can perform
logical deductions based on premises, which is a critical aspect of human-like reasoning
and decision-making.
3. Assess the impact of chain-of-thought prompting: Explore how chain-of-thought
prompting influences the performance of LLMs. This approach, which involves prompting models
to outline their reasoning step-by-step, could enhance both the accuracy of responses
and the quality of explanations provided by LLMs.
4. Evaluate the quality of LLM-generated explanations: Systematically assess the quality
of explanations generated by LLMs. This could involve several criteria such as clarity,
completeness, and correctness. Understanding the explanatory capabilities of LLMs is
vital for their applicability in education, decision support, and other areas requiring
interpretability.
5. Study the influence of RAG to receive more accurate, customized and context-aware LLM
responses. The aim is to identify the strengths and limitations of RAG. The focus could
be again on both the explanation quality as all as the accuracy of information.
These steps will contribute to a deeper understanding of the capabilities and limitations of
LLMs, particularly in tasks requiring sophisticated reasoning and explanations.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>I like to thank my supervisors Frieder Stolzenburg and Ute Schmid, for their invaluable guidance
and support. Their expert advice and insightful feedback are crucial to this thesis.
[12] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al., Chain-of-thought
prompting elicits reasoning in large language models, Advances in neural information processing
systems 35 (2022) 24824–24837.
[13] D. Paul, M. Ismayilzada, M. Peyrard, B. Borges, A. Bosselut, R. West, B. Faltings, Refiner: Reasoning
feedback on intermediate representations, 2024. arXiv:2304.01904.
[14] S. Krause, F. Stolzenburg, Commonsense reasoning and explainable artificial intelligence using
large language models, in: European Conference on Artificial Intelligence, Springer, 2023, pp.
302–319.
[15] N. Mostafazadeh, M. Roth, A. Louis, N. Chambers, J. Allen, LSDSem 2017 shared task: The Story
Cloze Test, in: Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and
Discourse-level Semantics, 2017, pp. 46–51. URL: https://aclanthology.org/W17-0906.pdf .
[16] Y. Onoe, M. J. Q. Zhang, E. Choi, G. Durrett, CREAK: A dataset for commonsense reasoning over
entity knowledge, 2021. URL: https://arxiv.org/pdf/2109.01653.
[17] M. Chen, M. D’arcy, A. Liu, J. Fernandez, D. Downey, CODAH: An adversarially-authored question
answering dataset for common sense, in: Proceedings of the 3rd Workshop on Evaluating Vector
Space Representations for NLP, 2019, pp. 63–69. URL: https://www.jaredfern.com/publication/
codah/.
[18] S. Singh, N. Wen, Y. Hou, P. Alipoormolabashi, T.-L. Wu, X. Ma, N. Peng, COM2SENSE: A
commonsense reasoning benchmark with complementary sentences, in: Findings of the Association for
Computational Linguistics: ACL-IJCNLP 2021, Association for Computational Linguistics, 2021, pp.
883–898. URL: https://aclanthology.org/2021.findings-acl.78.
[19] L. Huang, R. Le Bras, C. Bhagavatula, Y. Choi, Cosmos QA: Machine reading comprehension with
contextual commonsense reasoning, in: Proceedings of the 2019 Conference on Empirical Methods
in Natural Language Processing and the 9th International Joint Conference on Natural Language
Processing (EMNLP-IJCNLP), Association for Computational Linguistics, 2019, pp. 2391–2401. URL:
https://aclanthology.org/D19-1243/.
[20] L. Du, X. Ding, K. Xiong, T. Liu, B. Qin, e-CARE: a new dataset for exploring explainable causal
reasoning, in: Proceedings of the 60th Annual Meeting of the Association for Computational
Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, 2022, pp. 432–446.</p>
      <p>URL: https://aclanthology.org/2022.acl-long.33/.
[21] P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, O. Tafjord, Think you have
Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, CoRR – Computing Research
Repository abs/1803.05457, Cornell University Library, 2018. URL: https://arxiv.org/abs/1803.05457.
[22] M. Sap, H. Rashkin, D. Chen, R. LeBras, Y. Choi, Social IQa: Commonsense reasoning about
social interactions, in: Proceedings of the 2019 Conference on Empirical Methods in Natural
Language Processing and the 9th International Joint Conference on Natural Language Processing
(EMNLP-IJCNLP), Association for Computational Linguistics, 2019, pp. 4463–4473. URL: https:
//aclanthology.org/D19-1454/.
[23] M. Roemmele, C. A. Bejan, A. S. Gordon, Choice of plausible alternatives: An
evaluation of commonsense causal reasoning., in: AAAI Spring Symposium: Logical
Formalizations of Commonsense Reasoning, 2011, pp. 90–95. URL: https://aaai.org/papers/
02418-choice-of-plausible-alternatives-an-evaluation-of-commonsense-causal-reasoning/.
[24] A. Pal, L. K. Umapathi, M. Sankarasubbu, MedMCQA: A large-scale multi-subject multi-choice
dataset for medical domain question answering, ACM Conference on Health (2022). URL: https:
//arxiv.org/pdf/2203.14371.
[25] A. Talmor, J. Herzig, N. Lourie, J. Berant, CommonsenseQA: A question answering challenge
targeting commonsense knowledge, 2018. URL: https://arxiv.org/pdf/1811.00937.
[26] S. Krause, F. Stolzenburg, From data to commonsense reasoning: The use of large language models
for explainable ai, arXiv preprint arXiv:2407.03778 (2024).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Gunning</surname>
          </string-name>
          , D. Aha,
          <article-title>DARPA's explainable artificial intelligence (XAI) program</article-title>
          ,
          <source>AI</source>
          magazine
          <volume>40</volume>
          (
          <year>2019</year>
          )
          <fpage>44</fpage>
          -
          <lpage>58</lpage>
          . https://doi.org/10.1609/aimag.v40i2.
          <fpage>2850</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lapuschkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wäldchen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Binder</surname>
          </string-name>
          , G. Montavon,
          <string-name>
            <given-names>W.</given-names>
            <surname>Samek</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-R. Müller</surname>
          </string-name>
          ,
          <article-title>Unmasking Clever Hans predictors and assessing what machines really learn</article-title>
          ,
          <source>Nature Communications</source>
          <volume>10</volume>
          (
          <year>2019</year>
          )
          <article-title>1096</article-title>
          . URL: https://www.nature.
          <source>com/articles/s41467-019-08987-4</source>
          ,
          <fpage>10</fpage>
          .1038/s41467-019-08987-4.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Barredo Arrieta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Díaz-Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Del</given-names>
            <surname>Ser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bennetot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tabik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barbado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gil-Lopez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Molina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Benjamins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Chatila</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Herrera</surname>
          </string-name>
          ,
          <string-name>
            <surname>Explainable Artificial</surname>
          </string-name>
          <article-title>Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible ai</article-title>
          ,
          <source>Information Fusion</source>
          <volume>58</volume>
          (
          <year>2020</year>
          )
          <fpage>82</fpage>
          -
          <lpage>115</lpage>
          .
          <fpage>10</fpage>
          .1016/j.infus.
          <year>2019</year>
          .
          <volume>12</volume>
          .012.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bommasani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Borgeaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yogatama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bosma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Metzler</surname>
          </string-name>
          , et al.,
          <article-title>Emergent abilities of large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2206.07682</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Krause</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. H.</given-names>
            <surname>Panchal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ubhe</surname>
          </string-name>
          ,
          <article-title>The evolution of learning: Assessing the transformative impact of generative ai on higher education</article-title>
          ,
          <source>arXiv preprint arXiv:2404.10551</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N. F.</given-names>
            <surname>Rajani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McCann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <article-title>Explain yourself! leveraging language models for commonsense reasoning</article-title>
          , arXiv preprint arXiv:
          <year>1906</year>
          .
          <volume>02361</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Siebert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stolzenburg</surname>
          </string-name>
          ,
          <article-title>Commonsense reasoning using theorem proving and machine learning</article-title>
          ,
          <source>in: Machine Learning and Knowledge Extraction: Third IFIP TC 5, TC 12, WG 8.4, WG 8</source>
          .9,
          <string-name>
            <surname>WG</surname>
          </string-name>
          <year>12</year>
          .9 International
          <string-name>
            <surname>Cross-Domain</surname>
            <given-names>Conference</given-names>
          </string-name>
          , CD-MAKE
          <year>2019</year>
          ,
          <article-title>Canterbury</article-title>
          ,
          <string-name>
            <surname>UK</surname>
          </string-name>
          ,
          <year>August</year>
          26-
          <issue>29</issue>
          ,
          <year>2019</year>
          , Proceedings 3, Springer,
          <year>2019</year>
          , pp.
          <fpage>395</fpage>
          -
          <lpage>413</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Álvez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lucio</surname>
          </string-name>
          , G. Rigau,
          <article-title>Adimen-sumo: Reengineering an ontology for first-order reasoning</article-title>
          ,
          <source>International Journal on Semantic Web and Information Systems (IJSWIS) 8</source>
          (
          <year>2012</year>
          )
          <fpage>80</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A. d.</given-names>
            <surname>Garcez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Broda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Gabbay</surname>
          </string-name>
          ,
          <article-title>Symbolic knowledge extraction from trained neural networks: A sound approach</article-title>
          ,
          <source>Artificial Intelligence</source>
          <volume>125</volume>
          (
          <year>2001</year>
          )
          <fpage>155</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Le</given-names>
            <surname>Bras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bhagavatula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Cosmos</surname>
            <given-names>QA</given-names>
          </string-name>
          :
          <article-title>Machine reading comprehension with contextual commonsense reasoning</article-title>
          , in: K. Inui,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          Wan (Eds.),
          <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>2391</fpage>
          -
          <lpage>2401</lpage>
          . URL: https://aclanthology.org/D19-1243. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1243.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Albalak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , Logic-lm:
          <article-title>Empowering large language models with symbolic solvers for faithful logical reasoning</article-title>
          ,
          <source>arXiv preprint arXiv:2305.12295</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>