<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Dual-Phase Models for Extracting Information and Symbolic Reasoning: A Case-Study in Spatial Reasoning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roshanak Mirzaee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Parisa Kordjamshidi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, Michigan State University</institution>
          ,
          <addr-line>MI</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>QA Benchmark: Story</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Question Answring Asked (NN) Relation</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Spatial reasoning over text is challenging as the models need to extract the direct spatial information from the text, reason over those, and infer implicit spatial relations. Recent studies highlight the struggles even large-scale language models encounter when it comes to performing spatial reasoning over text. In this paper, we explore the potential benefits of disentangling the processes of information extraction and reasoning in models to address this challenge. To explore this, we devise various models that disentangle extraction and reasoning (either symbolic or neural) and compare them with SOTA baselines with no explicit design for these parts. Our experimental results consistently demonstrate the eficacy of disentangling, showcasing its ability to enhance models' generalizability within realistic data domains.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Spatial Reasoning</kwd>
        <kwd>Spatial Role Labeling</kwd>
        <kwd>Disentangling Extraction and Reasoning</kwd>
        <kwd>Pretrained Language Models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1. Introduction
a) PLMs:
End-to-End model
#
and QA+SpRL annotation, showcasing the advantage of
the design of this model in utilizing QA and SpRL data
within explicit extraction layers and the data
preprocessing. Also, the better performance of this model compared
to PistaQ demonstrates how the end-to-end structure of
SREQA can handle the errors from the extraction part
while capturing some rules and commonsense knowledge
from ReSQ training data that are not explicitly supported
in the symbolic reasoner.</p>
      <p>Table 1 Table 3
Results on auto-generated datasets. We use the accuracy Result on ReSQ. *Further training on SpaRTUN. The
metric for both YN (Yes/No) and FR (Find Relations) questions. _ℎ refers to evaluation without further training on
ReSQ or CLEF training data.</p>
      <p>
        Ttroaitnraining tohnetehxetrcaocrtrieosnpomnoddiunlgeso,rwaeuxaidliaaprtytdhaetmastehtsr.oTughhe Story: abepthwoeteonoafnadroaopmicwtuitrhe wonhittheewwaalllsl a,tbwovoestinhgelebebdesd.s with a night table in
outcomes of our experiments provide noteworthy in- Question: Are the beds below the picture? Answer: Yes
sights: Story BERT F0a:c['tasp:ircigtuhrte(2',,'t1h),ebbeelodws(']2,,20:)[,'an']e,a1r:([2'a,0p)icture', 'the wall']
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Tables 1 and 2 show the performance of models Facts: GPT3 3F:ac['ttsw:oasbionvgele(5b,e3d),sa',b'tohvee(b5e,d6s)'.].,.5: ['a picture'], 6: ['the wall', 'the beds']
SopnatRwToUdNataasnedtsSwp aitrhtcQoAnt-rAoulletdo.eAxpseirti misesnhtoawlcno,nPdiistitoanQs, Queries: GBPETR3T bbeellooww((30 ,, 50))??oorrbbeellooww((30,,61))??
voauttipoenrsfoirnmtshiasllsPetLtMingbadseemlionnesstarnatdeStRhEatQdAis.eOnutarnogblsinerg- Reasoning: GBPETR3T bbeellooww((30 ,, 50))==TFraulsee,,bbeeloloww(3(0, ,61))==FFaalslsee→→AnAsnwsewre=rY=eNso
extraction and symbolic reasoning compared to PLMs Figure 2: An example of using BERT-based SpRL and GPT3
enhances the models’ reasoning capabilities, even with as information extraction in PistaQ on a ReSQ.
comparable or reduced supervision. This result suggests
that SpRL annotations are more efective in the PistaQ
pipeline than when utilized in BERT-EQ in the form of (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) Recent research[1] show that powerful LLMs
canQA supervision. Note that the BERT-EQ uses all the orig- not perform well on SQA task. Similarly, our experiments,
inal dataset questions and extra questions created from as shown in Tables 2 and 3 show the lower performance
the full SpRL annotations. of GPT3.5 compared to humans and our models PistaQ
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) We select ReSQ as an SQA dataset with realistic and SREQA. However, we show that harnessing LLMs’
settings and present the result of models on this dataset potentials in information extraction [8] can yield
signifin Table 3. The low performance of PistaQ is attributed icant benefits within the disentangled structure of
exto , first, the absence of integrating commonsense in- traction and reasoning. Figure 2 provides a comparison
formation in this model and, second, the errors in the between the BERT-based SpRL extraction modules and
extraction modules, which are passed to the reasoning GPT3.5 with  _ℎ prompting in PistaQ. It shows
modules. SREQA* surpasses the PLMs trained on QA that GPT3.5 extracts more accurate information, leading
to correct answers from the reasoning phase.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cahyawijaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wilie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lovenia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Do</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fung</surname>
          </string-name>
          ,
          <article-title>A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination</article-title>
          , and interactivity,
          <year>2023</year>
          . arXiv:
          <volume>2302</volume>
          .
          <fpage>04023</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mirzaee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Rajaby</given-names>
            <surname>Faghihi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ning</surname>
          </string-name>
          , P. Kordjamshidi,
          <article-title>SPARTQA: A textual question answering benchmark for spatial reasoning, in: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>4582</fpage>
          -
          <lpage>4598</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .naacl-main.
          <volume>364</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          . naacl-main.
          <volume>364</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Mollá</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Van Zaanen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
          </string-name>
          , et al.,
          <article-title>Named entity recognition for question answering (</article-title>
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Coheur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. V.</given-names>
            <surname>Lobo</surname>
          </string-name>
          ,
          <article-title>Named entity recognition in questions: Towards a golden collection</article-title>
          .,
          <source>in: LREC</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lapata</surname>
          </string-name>
          ,
          <article-title>Using semantic roles to improve question answering</article-title>
          ,
          <source>in: Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing</source>
          and
          <string-name>
            <surname>Computational Natural Language Learning (EMNLP-CoNLL</surname>
            <given-names>)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Prague, Czech Republic,
          <year>2007</year>
          , pp.
          <fpage>12</fpage>
          -
          <lpage>21</lpage>
          . URL: https://aclanthology. org/D07-1002.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H. R.</given-names>
            <surname>Faghihi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kordjamshidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Teng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Allen</surname>
          </string-name>
          ,
          <article-title>The role of semantic parsing in understanding procedural text</article-title>
          ,
          <source>arXiv preprint arXiv:2302.06829</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mirzaee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kordjamshidi</surname>
          </string-name>
          ,
          <article-title>Transfer learning with synthetic corpora for spatial role labeling and reasoning</article-title>
          ,
          <source>in: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Abu Dhabi, United Arab Emirates,
          <year>2022</year>
          , pp.
          <fpage>6148</fpage>
          -
          <lpage>6165</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .emnlp-main.
          <volume>413</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Long</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Geng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <article-title>Large language models are strong zero-shot retriever</article-title>
          ,
          <source>arXiv preprint arXiv:2304.14233</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>