<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards Reliable Compositional Behavior in QALD Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Maria Schmidt</string-name>
          <email>daschmidt@techfak.uni-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raoul Schubert</string-name>
          <email>raoul.schubert@uni-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Cimiano</string-name>
          <email>cimiano@techfak.uni-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Compositionality, Question Answering over Linked Data, Large Language Models</institution>
          ,
          <addr-line>Semantic Web</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Semantic Computing Group, CITEC, Technical Faculty, Bielefeld University</institution>
          ,
          <addr-line>Bielefeld</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <fpage>2</fpage>
      <lpage>6</lpage>
      <abstract>
        <p>Accompanying the Research Track paper “CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting”, we investigate how compositionality is approached in our compositional question answering over linked data (QALD) pipeline “NeoDUDES”. This way, we point out how some of the limitations of large language models (LLMs) w.r.t. compositional interpretation of QALD questions can be dealt with by combining LLMs with symbolic methods. In our demo, we show detailed intermediate results from the NeoDUDES pipeline, underlining how the strengths of neural and symbolic approaches can be combined in a fine-grained, compositional pipeline to tackle compositional tasks in a more reliable fashion.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The reasoning abilities of large language models (LLMs) and especially the abilities of LLMs to work
and reason in a compositional way have been investigated by numerous related works in recent years,
either by directly targeting compositionality [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ], or indirectly through various multi-step reasoning
tasks [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7">4, 5, 6, 7</xref>
        ]. Other works also investigated the abilities of LLMs w.r.t. compositionality from
a (complexity-) theoretical perspective [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref8 ref9">8, 9, 10, 11, 12</xref>
        ]. However, in contrast to our accompanying
Interpret Questions in a QALD Setting” [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], most related work does not deal with or focus on QALD
specifically. Therefore, we adapt the compositionality term of Zoltán G. Szabó [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to the QALD domain
and generate a corresponding benchmark dataset CompoST (“Compositional Systematicity Test”) to test
the abilities of LLMs to systematically recombine known parts to new SPARQL queries. Our evaluation,
summarizing over 400 experiments, raises substantial concerns w.r.t. the ability of LLMs to interpret
QALD questions in a systematic, compositional way, even when all necessary information to interpret a
question is given in the input. In line with, e.g., Dziri et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], this may indicate fundamental limitations
of LLMs when it comes to truly compositional tasks.
      </p>
      <p>
        In this paper, we further analyze the issues of LLMs with compositional tasks that have been raised
in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Furthermore, we propose first solutions to those problems by demonstrating how we deal
with compositionality in our compositional question answering over linked data (QALD) pipeline
“NeoDUDES”1 [
        <xref ref-type="bibr" rid="ref15">15, 16</xref>
        ]. This way, we show new avenues for future work to ensure reliable compositional
behavior by combining the strengths of symbolic and LLM-based approaches in a single QALD pipeline.
Finally, we leverage the fine-grained nature of our pipeline to present its intermediate results in a
corresponding demo2 and thus allow deep insights into the inner workings of the pipeline.
∗Corresponding author.
https://davidmschmidt.de/ (D. M. Schmidt); http://cimiano.de (P. Cimiano)
      </p>
      <p>CEUR</p>
      <p>ceur-ws.org</p>
      <p>
        This paper only highlights the parts of the NeoDUDES pipeline most relevant for compositionality.
For more detailed information about the pipeline, we refer the interested reader to [
        <xref ref-type="bibr" rid="ref15">15, 16</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Methods</title>
      <p>
        In CompoST [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], we focus on a sub-property of compositionality, namely systematicity. That means,
citing a classic example from Zoltán G. Szabó [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], a compositional system understanding both “brown
dog” and “black cat” should understand “brown cat” as well. Adapted to the QALD domain, we thus
expect a compositional system which correctly generates SPARQL queries for “What is the birth name
of Angela Merkel?” and “What is the birth place of Barack Obama?” to also generate a correct query
for, e.g., “What is the birth place of Angela Merkel?”. As we show in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], this property is violated
frequently by current LLMs. Thus, in this section, we focus on the question how one can approach
compositionality in a diferent, more robust way by the example of the NeoDUDES pipeline.
      </p>
      <p>An overview of the whole pipeline is given in Figure 1. In this paper, we mainly focus on the DUDES
composition as well as the way ambiguities are handled by the pipeline.</p>
      <p>The core of the NeoDUDES pipeline as well as its compositional backbone are Dependency-based
Underspecified Discourse Representation Structures (DUDES) , which are used to represent the meaning of
a question or parts of it. They are defined as follows:
Definition 1 (Dependency-based Underspecified Discourse Representation Structure [
A Dependency-based Underspecified Discourse Representation Structure (DUDES) is a triple ( , , )
17, 15]).
where:
•  ∈  ∪ {} is the main variable (also called referent marker or distinguished variable) where 
represents the absence of a main variable
•  = ( , ) is a Discourse Representation Structure (DRS) [17, 18, 19] with
(a) Entity
for
(b) Property DUDES for dbo:
(c) Composition of 2a and 2b using
dbr:Angela_Merkel
birthName
pairs (, " ")
with</p>
      <p>selection
and ( , )
selection pair (, " ") .
– set of variables  (also called discourse universe or referent markers)
– set of conditions  over variables 
•  is a set of selection pairs of the form ( , )</p>
      <p>with  being a variable from  and  being a marker
word for that variable with  representing the empty marker, i.e., no marker being connected to that
variable. Instead of writing  , the second tuple component can also just be left out.</p>
      <p>
        Two example DUDES are given in Figures 2a and 2b. Thus, DUDES represent the meaning of (parts
of) a question or sentence through logical formulas that roughly correspond to SPARQL triple patterns
in most cases. Additionally, the main variable and selection pairs are what makes this representation
compositional, as they are used for the composition operation of two DUDES. This operation has the
goal to compose the meaning of two DUDES, and thus two parts of a question, into one combined
meaning representation. This composition operation is defined as follows:
Definition 2 (DUDES Composition [
        <xref ref-type="bibr" rid="ref15">17, 15</xref>
        ]). Let  1 = ( 1,  1 = ( 1,  1),  1),  2 = ( 2,  2 =
( 2,  2),  2) be two DUDES with disjoint variable sets, i.e.  1 ∩  2 = ∅. The DUDES composition
operation ⊙ for substituting  1 into  2 using selection pair  = ( ∈ 
2, ) ∈  2 and resulting in a composed
DUDES   = (  ,   = (  ,   ),   ), written   =  1 ⊙  2, is defined as follows:
      </p>
      <p>= ( 2 ∪  1) ∖ 
  =  2[ ≔  1] ∪  1
  =  2[ ≔  1] ∪  1
  = {
 1 if  =  2
 2
else</p>
      <p>An example of the result of such a composition operation is given in Figure 2c. In practice, most
questions consist of more than two parts. Therefore, the composition operation is applied bottom-up
along a (slightly compacted) dependency tree in order to create a single DUDES representation from
multiple parts of a question. Both the original DUDES for all parts of the question as well as the final
resulting DUDES are presented in our demo for the respective given question.</p>
      <p>However, as the result of the above composition operation depends on both the used selection pair
as well as the direction in which the operation is applied, there may arise ambiguities when there are
multiple possibilities that cannot be further disambiguated. In those cases, we both want to avoid the
problem of state space explosion as well as losing potentially correct combinations by deciding for one
option too early and discarding the others.</p>
      <p>Therefore, the NeoDUDES pipeline applies an iterative approach for most steps, assembling one query
at a time without explicitly storing all possible combinations of, e.g., DUDES compositions in memory.
However, even though this limits the memory consumption, the number of possible combinations
remains large. In order to still get results in a reasonable time, a decision has to be made which
combinations are tried first and to which extent to tradeof memory for runtime.</p>
      <p>In the pipeline, this is mainly done in the Tree Scorer and SPARQL Selector components. The Tree
Scorer gets a set of trees to which diferent node merging heuristics have been applied. For example,
in the original dependency tree, the entity dbr:Angela_Merkel is split into two nodes corresponding
to “Angela” and “Merkel”, respectively. To improve those correspondences and facilitate ontology
h
t
ep 2
D
1
3
0.34
0.15
0.11</p>
      <p>Zero-Shot (Hard)</p>
      <p>Breadth
2
matching, various merging heuristics are applied. However, this gives us a number of candidate trees
that typically cannot all be processed at the same time. Therefore, the Tree Scorer assigns each tree a
score, aiming to measure how promising that tree is in terms of size and matched ontology resources
and thus determining an order for those trees in which they are further processed. These scores together
with the corresponding trees are as well part of the demo.</p>
      <p>Similarly, the iterative approach which follows after the Tree Scorer produces candidate SPARQL
queries one by one while avoiding to store multiple possible combinations at the same time to limit
memory usage. From these queries, one has to be chosen as the final output. As there are no clear rules
for what a good SPARQL query for a specific question is, we train an LLM to compare two candidate
queries w.r.t. a given question and choose the “better” one. These single comparisons of the LLM-based
SPARQL selection are then aggregated in diferent ways for diferent strategies to arrive at a final
decision. This way, the symbolic part of the approach shows its strength by trying diferent possibilities
in a structured and reliable way while an LLM-based approach deals with the more “fuzzy” task of
selecting a final query from a set of candidates. The SPARQL selection is also illustrated in the demo.</p>
      <p>All in all, this underlines how symbolic and neural components can work together to provide both,
reliable compositional behavior without losing possible combinations of the compositional parts, as
well as LLM-based optimizations and trained heuristics for scenarios where all available rules have been
applied but still some decisions need to be made. This shows promising avenues for future research.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Results and Discussion</title>
      <p>
        Revisiting the scores achieved by current LLMs in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], this underlines the need for new methods that
deal with compositionality in a more robust way. This gets especially clear when considering the scores
of the best-performing zero-shot approach on the hard CompoST dataset, presented in Figure 3, together
with the scores of the few-shot and fine-tuning approaches shown in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. These  1 scores, grouped by
breadth and depth of the respective SPARQL graph pattern, were achieved by Llama 3.3 [20], using
MIPRO prompt optimization with the heavy preset in combination with Chain of Thought prompting.
All experiments were conducted using the DSPy framework [21, 22]. Overall, a broad set of models
has been tested, namely Llama 3.3 (70B) [20], Phi-4 (14B) [23], Qwen2.5-Coder (7B) [24, 25], OLMo 2
(7B) [26] and GPT-4o-mini [27]. The prompting techniques included plain prompting, COPRO prompt
optimization as well as MIPRO prompt optimization, each tested with and without Chain of Thought
prompting. Further information on the conducted experiments as well as heatmaps for few-shot and
ifne-tuning can be found in the accompanying Research Track paper [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        In general, the experimental results of Schmidt et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] show that LLMs struggle with compositional
tasks, especially as the size of the questions gets further away from the data observed during training
although it was ensured that the training data contained all relevant information and “building blocks”
to construct the answer for the questions in the validation and test splits of the dataset. Even for
“selfcontained” experiments, containing all necessary information to solve the task in the input, the achieved
scores did not exceed 0.57 in terms of test macro  1 scores (achieved on the easy CompoST dataset
with Llama 3.3 using few-shot prompting together with MIPRO prompt and shot optimization with a
medium preset). This shows additional efort is needed whenever reliable compositional interpretation
of QALD questions is necessary. Some possibilities on how to achieve this have been outlined above.
      </p>
      <p>However, there are also limitations of the presented NeoDUDES pipeline. First, the pipeline relies on
the availability of a Lemon lexicon [28], covering all relevant verbalizations of used properties. Similarly,
e.g., rdfs:label data for our trie-based entity matcher or some other entity matcher has to be available.
Second, depending on how many combinations have to be tested before a suitable candidate is found,
the runtime of the NeoDUDES pipeline can be much longer than typical inference times of current
LLMs. Finally, the initial implementation efort of the pipeline was higher than the efort typically
necessary for, e.g., fine-tuning or prompt optimization for the QALD task.</p>
      <p>Nevertheless, the existing pipeline can now be easily adapted to new datasets or knowledge graphs.
This can be especially useful for small or domain-specific datasets which are not suficient for purely
LLM-based approaches either due to their size or because the respective knowledge graph or the style
of the questions deviates too much from the LLM training data. Moreover, an open modular pipeline
like the NeoDUDES approach provides a whole new level in terms of explainability and possibilities to
justify answers or fix errors that a purely LLM-based approach typically cannot ofer.</p>
      <p>In future work, we aim to test diferent ways to generate the required Lemon lexicon automatically,
using combinations of existing data sources (e.g., WordNet, Wikidata alias entries, inflection tools,
etc.) as well as LLM-based generation. Additionally, as the goal of the pipeline is to use symbolic and
neural approaches where they each work best, we plan to replace diferent parts of the pipeline with
LLMs for that specific sub-task and investigate how this compares to the performance of the symbolic
pipeline in terms of compositionality. Although preliminary results show promising performance of the
NeoDUDES pipeline on the CompoST dataset, we aim to provide a full evaluation in the future. Finally,
various performance optimizations and further parallelization is planned to improve the responsiveness
and runtime of the pipeline.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>
        To summarize, in this paper, we revisited the results of the accompanying Research Track paper [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
highlighting the weaknesses and limitations of LLMs when it comes to truly compositional tasks.
Motivated by these findings, we investigated how the NeoDUDES pipeline, a compositional approach
by design that combines the strengths of both symbolic and neural methods in a transparent modular
pipeline, approaches compositionality. An illustration of these aspects and advantages is also part of
the corresponding demo3.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work is partially funded by the Ministry of Culture and Science of the State of North
RhineWestphalia under grant no NW21-059A (SAIL).</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>The authors have not employed any Generative AI tools.
3Demo video: https://doi.org/10.5281/zenodo.16531345. Although the video only shows one example, the demo supports
generating these illustrations for any given input question.
lexical knowledge in a compositional QALD system, in: M. Alam, M. Rospocher, M. van Erp,
L. Hollink, G. A. Gesese (Eds.), Knowledge engineering and knowledge management, Springer
Nature Switzerland, Cham, 2025, pp. 102–122.
[16] D. M. Schmidt, M. F. Elahi, P. Cimiano, Lexicalization Is All You Need: Examining the Impact
of Lexical Knowledge in a Compositional QALD System, in: C. Badenes-Olmedo, I. Novalija,
E. Daga, L. Stork, R. G. Pillai, L. Dierickx, B. Kruit, V. Degeler, J. Moreira, B. Zhang, R. Alharbi,
Y. He, A. Graciotti, A. M. Tirado, V. Presutti, E. Motta (Eds.), Joint Proceedings of Posters, Demos,
Workshops, and Tutorials of the 24th International Conference on Knowledge Engineering and
Knowledge Management (EKAW-PDWT 2024), volume 3967 of CEUR Workshop Proceedings, CEUR,
Amsterdam, Netherlands, 2024.
[17] P. Cimiano, C. Unger, J. P. McCrae, Ontology-Based Interpretation of Natural Language, Synthesis</p>
      <p>Lectures on Human Language Technologies, Morgan &amp; Claypool Publishers, 2014.
[18] K. Hans, A theory of truth and semantic representation, Formal Methods in the Study of language
(1981).
[19] H. Kamp, U. Reyle, From discourse to logic: Introduction to modeltheoretic semantics of natural
language, formal logic and discourse representation theory, volume 42, Springer Science &amp; Business
Media, 2013.
[20] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur,
A. Schelten, A. Vaughan, et al., The Llama 3 Herd of Models, 2024. URL: http://arxiv.org/abs/2407.
21783. doi:10.48550/arXiv.2407.21783, arXiv:2407.21783 [cs].
[21] O. Khattab, A. Singhvi, P. Maheshwari, Z. Zhang, K. Santhanam, S. Vardhamanan, S. Haq, A. Sharma,
T. T. Joshi, H. Moazam, H. Miller, M. Zaharia, C. Potts, DSPy: Compiling declarative language
model calls into self-improving pipelines, The Twelfth International Conference on Learning
Representations, 2024.
[22] O. Khattab, K. Santhanam, X. L. Li, D. Hall, P. Liang, C. Potts, M. Zaharia,
Demonstrate-searchpredict: Composing retrieval and language models for knowledge-intensive NLP, arXiv preprint
arXiv:2212.14024 (2022).
[23] M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M.
Javaheripi, P. Kaufmann, J. R. Lee, Y. T. Lee, Y. Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. d.
Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y. Wu, D. Yu, C. Zhang, Y. Zhang, Phi-4
Technical Report, 2024. URL: http://arxiv.org/abs/2412.08905. doi:10.48550/arXiv.2412.08905,
arXiv:2412.08905 [cs].
[24] A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, others, Qwen2
technical report, arXiv preprint arXiv:2407.10671 (2024).
[25] B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Dang, others, Qwen2.</p>
      <p>5-coder technical report, arXiv preprint arXiv:2409.12186 (2024).
[26] T. OLMo, P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jordan,
N. Lambert, D. Schwenk, O. Tafjord, T. Anderson, D. Atkinson, F. Brahman, C. Clark, P. Dasigi,
N. Dziri, M. Guerquin, H. Ivison, P. W. Koh, J. Liu, S. Malik, W. Merrill, L. J. V. Miranda, J. Morrison,
T. Murray, C. Nam, V. Pyatkin, A. Rangapur, M. Schmitz, S. Skjonsberg, D. Wadden, C. Wilhelm,
M. Wilson, L. Zettlemoyer, A. Farhadi, N. A. Smith, H. Hajishirzi, 2 OLMo 2 Furious, 2025. URL:
http://arxiv.org/abs/2501.00656. doi:10.48550/arXiv.2501.00656, arXiv:2501.00656 [cs].
[27] OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J.
Altenschmidt, S. Altman, S. Anadkat, et al., GPT-4 Technical Report, 2024. URL: http://arxiv.org/abs/
2303.08774. doi:10.48550/arXiv.2303.08774, arXiv:2303.08774 [cs].
[28] J. P. McCrae, D. Spohr, P. Cimiano, Linking lexical resources and ontologies on the semantic web
with lemon, in: Proceedings of the 8th extended semantic web conference on The semantic web:
research and applications (ESWC), volume 6643, 2011, pp. 245–259.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Dziri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sclar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>West</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bhagavatula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Le</given-names>
            <surname>Bras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Hwang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sanyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Welleck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ettinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Harchaoui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <article-title>Faith and fate: limits of transformers on compositionality</article-title>
          ,
          <source>in: Proceedings of the 37th international conference on neural information processing systems</source>
          , Nips '
          <volume>23</volume>
          , Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <year>2023</year>
          . Number of pages:
          <volume>40</volume>
          Place: New Orleans, LA, USA tex.
          <source>articleno: 3081.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>O.</given-names>
            <surname>Press</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <article-title>Measuring and Narrowing the Compositionality Gap in Language Models</article-title>
          , in: H.
          <string-name>
            <surname>Bouamor</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Pino</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          Bali (Eds.),
          <source>Findings of the Association for Computational Linguistics: EMNLP</source>
          <year>2023</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Singapore,
          <year>2023</year>
          , pp.
          <fpage>5687</fpage>
          -
          <lpage>5711</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .findings-emnlp.
          <volume>378</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .findings-emnlp.
          <volume>378</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hupkes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dankers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mul</surname>
          </string-name>
          , E. Bruni, Compositionality Decomposed:
          <article-title>How do Neural Networks Generalise?</article-title>
          ,
          <source>Journal of Artificial Intelligence Research</source>
          <volume>67</volume>
          (
          <year>2020</year>
          )
          <fpage>757</fpage>
          -
          <lpage>795</lpage>
          . URL: https://jair.org/ index.php/jair/article/view/11674. doi:
          <volume>10</volume>
          .1613/jair.1.11674.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Backurs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bubeck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Eldan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gunasekar</surname>
          </string-name>
          , T. Wagner,
          <article-title>Unveiling Transformers with LEGO: a synthetic reasoning task</article-title>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2206.04301. doi:
          <volume>10</volume>
          .48550/ arXiv.2206.04301, arXiv:
          <fpage>2206</fpage>
          .04301 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Andreassen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gur-Ari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Michalewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Austin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bieber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lewkowycz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bosma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Odena</surname>
          </string-name>
          , Show Your Work:
          <article-title>Scratchpads for Intermediate Computation with Language Models</article-title>
          ,
          <year>2021</year>
          . URL: http://arxiv.org/abs/2112.00114. doi:
          <volume>10</volume>
          .48550/arXiv. 2112.00114, arXiv:
          <fpage>2112</fpage>
          .00114 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Welleck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hajishirzi</surname>
          </string-name>
          , Y. Choi,
          <source>NaturalProver: Grounded Mathematical Proof Generation with Language Models, Advances in Neural Information Processing Systems</source>
          <volume>35</volume>
          (
          <year>2022</year>
          )
          <fpage>4913</fpage>
          -
          <lpage>4927</lpage>
          . URL: https://proceedings.neurips.cc/paper_files/paper/2022/hash/ 1fc548a8243ad06616eee731e0572927-Abstract-Conference.html.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Saparov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>Language Models Are Greedy Reasoners: A Systematic Formal</surname>
          </string-name>
          <article-title>Analysis of Chain-of-</article-title>
          <string-name>
            <surname>Thought</surname>
          </string-name>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2210.01240. doi:
          <volume>10</volume>
          .48550/arXiv.2210.01240, arXiv:
          <fpage>2210</fpage>
          .01240 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W.</given-names>
            <surname>Merrill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sabharwal</surname>
          </string-name>
          ,
          <article-title>The parallelism tradeof: Limitations of log-precision transformers</article-title>
          ,
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>11</volume>
          (
          <year>2023</year>
          )
          <fpage>531</fpage>
          -
          <lpage>545</lpage>
          . URL: https://doi.org/ 10.1162/tacl_a_00562. doi:
          <volume>10</volume>
          .1162/tacl_a_
          <volume>00562</volume>
          , tex.eprint: https://direct.mit.edu/tacl/articlepdf/doi/10.1162/tacl\_a\_
          <volume>00562</volume>
          /2131191/tacl\_a\_
          <volume>00562</volume>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Peng</surname>
          </string-name>
          , H. Wu,
          <article-title>Theoretical limitations of multi-layer</article-title>
          <string-name>
            <surname>Transformer</surname>
          </string-name>
          ,
          <year>2024</year>
          . URL: https: //arxiv.org/abs/2412.02975, arXiv:
          <fpage>2412</fpage>
          .02975 [cs.LG].
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Zubić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Soldá</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sulser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Scaramuzza</surname>
          </string-name>
          ,
          <article-title>Limits of deep learning: Sequence modeling through the lens of complexity theory</article-title>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2405.16674, arXiv:
          <fpage>2405</fpage>
          .16674 [cs.LG].
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Towards revealing the mystery behind chain of thought: a theoretical perspective</article-title>
          , in: A.
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Globerson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Saenko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Hardt</surname>
          </string-name>
          , S. Levine (Eds.),
          <source>Advances in neural information processing systems</source>
          , volume
          <volume>36</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2023</year>
          , pp.
          <fpage>70757</fpage>
          -
          <lpage>70798</lpage>
          . URL: https://proceedings.neurips.cc/paper_files/paper/2023/ file/dfc310e81992d2e4cedc09ac47eff13e-Paper-Conference.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Strobl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Merrill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Angluin</surname>
          </string-name>
          ,
          <article-title>What formal languages can transformers express? A survey, Transactions of the Association for Computational Linguistics 12 (</article-title>
          <year>2024</year>
          )
          <fpage>543</fpage>
          -
          <lpage>561</lpage>
          . URL: http://dx.doi.org/10.1162/tacl_a_00663. doi:
          <volume>10</volume>
          .1162/tacl_a_
          <volume>00663</volume>
          , publisher: MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>D. M. Schmidt</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Schubert</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <article-title>Compost: A benchmark for analyzing the ability of llms to compositionally interpret questions in a qald setting</article-title>
          ,
          <source>in: The Semantic Web - ISWC 2025</source>
          , Springer Nature Switzerland, Cham,
          <year>2025</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2507.21257, (in press).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Z. G. Szabó,</surname>
          </string-name>
          <article-title>The case for compositionality</article-title>
          , in: M.
          <string-name>
            <surname>Werning</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Hinzen</surname>
          </string-name>
          , E. Machery (Eds.),
          <source>The oxford handbook of compositionality</source>
          , Oxford University Press,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>D. M. Schmidt</surname>
            ,
            <given-names>M. F.</given-names>
          </string-name>
          <string-name>
            <surname>Elahi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <article-title>Lexicalization is all you need: Examining the impact of</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>