<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>E. Damiano);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Evaluating Large Language Models on OWL Lite Reasoning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emanuele Damiano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Orciuoli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dipartimento di Scienze Aziendali - Management &amp; Innovation Systems, Università degli Studi di Salerno</institution>
          ,
          <addr-line>via Giovanni Paolo II, 132, Fisciano (SA)</addr-line>
          ,
          <country country="IT">Italia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This work investigates the capability of large language models (LLMs) to interpret OWL Lite ontologies and perform reasoning over them. We propose an evaluation framework based on the well-known LUBM ontology that is transformed into text and vectorized by using an embedding algorithm, enabling retrieval-augmented generation (RAG) to support query answering. Any previous knowledge of LLMs related to LUBM has been excluded by using adequate prompts, in order to rely exclusively on the information locally obtained through RAG. A set of 53 manually constructed queries is used to probe the models' ability to perform ontology-based inference aligned with the ontology axioms. Such queries vary in complexity (from level 1 to level 3) based on the number and depth of required logical inference operations. No local knowledge about OWL and ontology-based reasoning has been provided to the models; therefore, we are confident to evaluate emergent abilities in the realm of reading, interpreting, and reasoning on OWL Lite ontologies. The answers provided by the LLMs are compared against a gold standard to compute accuracy. Furthermore, we evaluate and compare the performance of diferent existing models within this setting to assess their relative efectiveness in OWL-based reasoning tasks. The obtained results ofer interesting insights into the reasoning potential of LLMs when grounded in symbolic ontological knowledge.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large Language Models</kwd>
        <kwd>Web Ontology Language (OWL)</kwd>
        <kwd>Ontology-based Reasoning</kwd>
        <kwd>Neurosymbolic Computing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and Motivations</title>
      <p>
        Large Language Models (LLMs) are increasingly popular tools capable of sophisticated natural language
processing, demonstrating their potential in diverse applications, including complex code generation and
combinatorial problem-solving [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this context, recent studies have sought to integrate LLMs with
symbolic reasoners [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to develop more robust and capable intelligent systems under the framework
of Neurosymbolic Computing [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The central question addressed by this research is the degree
to which LLMs can efectively process and leverage structured knowledge represented in formal
ontologies, as well as how various well-known LLMs accessible from platforms like Groq, Ollama,
and others compare in this capability. To this end, we delve into the abilities of LLMs in reading,
interpreting, and reasoning over OWL ontologies, which are foundational for semantic web technologies
and knowledge representation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Understanding how LLMs interpret and reason over structured
knowledge, such as OWL ontologies, is a fundamental challenge at the intersection of symbolic and
neural AI. OWL ontologies represent formal, machine-readable semantics that underpin critical domains
like healthcare, finance, legal systems, and scientific research. Demonstrating that LLMs can perform
correct inferences over such ontologies—especially under constrained setups like Retrieval-Augmented
Generation (RAG)—is not only a technical milestone but a conceptual leap toward integrating symbolic
reasoning with neural models[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This has direct implications for neurosymbolic computing, where
the goal is to combine the robustness of learning with the rigor of logic [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Moreover, proving that
LLMs can act as semantic reasoners supports the development of more explainable AI systems, as
ontological reasoning is inherently interpretable [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. It also contributes to Generative eXplainable AI
(GenXAI) by showcasing how generative models can produce grounded, semantically valid responses
rooted in structured knowledge[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Ultimately, this research can redefine how we build intelligent
systems—moving from purely data-driven responses to semantically-aware and logically-consistent
behavior, with transformative potential across knowledge-intensive applications. The paper presents a
methodology to evaluate the aforementioned ability by leveraging LUBM1, a well-known benchmark
originally used to assess the performance of triple stores, and constructing a set of queries (written
in natural language) with increasing dificulty levels that require LLMs to make inferences over the
LUBM ontology in order to provide the correct answers. In this paradigm, our study focuses on OWL
Lite, a simplified subset of the Web Ontology Language, chosen for its balance between expressiveness
and simplicity, making it suitable for assessing the foundational understanding of ontologies by LLMs.
Specifically, we provide a prompt constraining the models to consider only the LUBM ontology provided
locally through RAG and any existing knowledge about OWL Lite and ontology-based reasoning.
In this way, we can evaluate models’ emergent abilities needed to read, interpret, and reason on
OWL Lite ontologies, i.e., the ability of LLMs to act as reasoners. The evaluation activities focus on
both quantitative and qualitative aspects. The quantitative evaluations are mainly based on accuracy
measures when comparing LLMs’ answers to the gold standard during query execution. The qualitative
evaluations, on the other hand, involve analyzing the LLMs’ reasoning processes and evaluating their
plausibility and faithfulness [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] by extracting and inspecting a self-evaluation from the LLMs’
answers.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>Recent studies have explored the ability of Large Language Models (LLMs) to interpret and reason over
ontological knowledge expressed in formal languages. Two notable directions are particularly relevant
to our work: the evaluation of symbolic knowledge implicitly learned by LLMs and the assessment of
their reasoning capabilities over structured ontologies.</p>
      <p>
        Authors of [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] introduce ONTOLAMA, a benchmark designed to evaluate whether LLMs encode
semantic subsumption relations derived from OWL ontologies. Their methodology probes whether
concepts of the form  ⊑  are entailed by the model when translated into natural language templates.
Crucially, their approach relies on zero-shot evaluation without external knowledge access and focuses
exclusively on subsumption. While insightful in measuring internalized semantic knowledge, it does
not test the model’s ability to reason dynamically over a structured ontology.
      </p>
      <p>
        In contrast, the work [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] proposes a comprehensive evaluation of LLMs’ understanding of DL-Lite
ontologies, examining tasks such as syntax checking, concept and role subsumption, instance checking,
property characteristics, and query answering. Their framework involves prompting LLMs directly
with formal axioms and evaluating their reasoning behavior without any external knowledge retrieval.
The study demonstrates that while LLMs can handle simple axiomatic patterns, their performance
degrades significantly when transitivity or larger ABoxes are involved.
      </p>
      <p>
        Our work difers significantly from both approaches. Rather than probing the LLM’s internalized
ontology knowledge or reasoning over embedded axioms, we leverage a retrieval-augmented generation
(RAG) architecture to provide the model with access to an OWL Lite ontology—specifically the LUBM
benchmark—in TURTLE (TTL) format[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. We encode the ontology using dense text embeddings and
enable the model to retrieve relevant axioms during inference. The reasoning is evaluated through a
curated set of natural language queries, stratified by inferential complexity. This allows us to assess
whether the LLM can simulate an OWL reasoner when given symbolic knowledge dynamically at
inference time, rather than relying on prior training.
      </p>
      <p>From a neurosymbolic AI perspective, our approach demonstrates the feasibility of grounding LLMs
in external ontological structures to support symbolic reasoning, thus bridging the gap between neural
text processing and logical inference. Unlike prior works, we evaluate operational reasoning using</p>
      <sec id="sec-2-1">
        <title>1https://swat.cse.lehigh.edu/projects/lubm/</title>
        <p>realistic ontological content, reflecting practical challenges in knowledge-based systems and explainable
AI.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>The methodology adopted to evaluate the ability of LLMs to read, understand, and make inferences
over OWL Lite ontologies consists of two main phases. The first phase is to provide the OWL Lite
ontology to the LLM and guide it to answer user’s questions leveraging on both its knowledge about
OWL Lite and the specific available OWL Lite ontology.</p>
      <p>As reported in Fig. 1 and detailed in Fig. 2, an important aspect of the first phase involves prompt
engineering. Such a prompt excludes previous knowledge about LUBM, which might conflict with
the provided ontology, thereby enabling a fair evaluation of the models’ ability to leverage both
their understanding of OWL Lite, based on their prior knowledge, and their capability to apply such
knowledge when working on specific OWL-based schemes like LUBM. The second phase consists of
concretely assessing the aforementioned ability by comparing the performance of several models and
analyzing quantitatively and qualitatively such performance. This is realized by executing a benchmark
composed by the existing LUBM ontology and a new set of natural language questions (crafted by
the authors of this work) requiring the execution of OWL Lite inferences over LUBM. The underlying
idea is that if LLMs are able to answer questions requiring OWL Lite inferences, after limiting their
knowledge to OWL Lite language and LUBM ontology, they demonstrate (within some limits) that they
can read, understand, make inferences, and in some sense act as a symbolic reasoner. Such results are
important for future works targeted at designing neurosymbolic systems.</p>
      <sec id="sec-3-1">
        <title>3.1. OWL Lite</title>
        <p>OWL Lite is a simplified sublanguage of the Web Ontology Language (OWL) designed for taxonomies
and basic constraints, balancing expressiveness with computational tractability. It corresponds to the
SHIF(D) description logic, enabling structured knowledge representation while maintaining decidability.
The key inference types supported by OWL Lite are listed in Tab. 1. More expressive ontology languages
(e.g., OWL 2) will be considered in further works.</p>
        <p>
          Description
Determines subclass and subproperty relationships. Enables
deduction of implicit class/property hierarchies using TBox reasoning (e.g.,
if  ⊑  and  ⊑ , then  ⊑ )[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ][
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          Identifies when two classes or properties are equivalent, including
handling of owl:sameAs for individuals, leading to entailment of
all facts about one individual to all equivalents[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ][
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>
          Assigns individuals (ABox) to classes based on asserted and inferred
property values and class definitions[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ][
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          Detects logical contradictions within the ontology, such as violations
of cardinality (only 0 or 1 allowed in OWL Lite), property constraints,
or incompatible class assertions[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ][
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          Supports inference with transitive, symmetric, functional, and
inverse properties. For example, infers new relationships via property
characteristics (e.g., if  is transitive and  ,   then  )[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ][
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          Infers class membership of individuals based on property domain
and range declarations (e.g., if   and  has domain , then  is
inferred to be an instance of )[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. LUBM Ontology</title>
        <p>
          The Lehigh University Benchmark (LUBM) is a widely used benchmark for evaluating the performance
of Semantic Web knowledge base systems, particularly those supporting OWL reasoning[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. The LUBM
ontology is designed to model a university domain, providing a structured framework for representing
entities such as students, faculty, organizations, and academic programs.
        </p>
        <p>
          The ontology defines a comprehensive set of classes and relationships relevant to university life.
Notably, it includes 43 classes and 32 properties, with 25 object properties (relationships between
classes) and 7 datatype properties (attributes with literal values). Key classes include Person, Student,
Employee, Dean, TeachingAssistant, Organization, Program, University, Work, Course,
Unit, Stream, and Graduate Course. Relationships such as author, member, degreeFrom, masterDegreeFrom,
and takesCourse connect these classes, reflecting real-world interactions in an academic environment[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>A distinctive feature of the LUBM ontology is its use of OWL Lite language constructs, including
inverse properties (inverseOf), transitive properties (Transitive Property), some-value
restrictions (someValuesFrom), and intersections (intersecti onOf). This allows for a moderate level of
expressivity while maintaining computational tractability, making the ontology well-suited for
benchmarking systems with varying reasoning capabilities. The LUBM benchmark is accompanied by scalable
synthetic datasets that represent universities and their constituents, enabling controlled experiments
and repeatable performance evaluations. The benchmark also includes a set of 14 extensional queries
that test a variety of reasoning and retrieval tasks, ranging from simple class membership checks to
complex relationship traversals. In summary, the LUBM ontology serves as a robust, standardized test
bed for assessing the scalability, reasoning, and querying capabilities of Semantic Web systems, with a
particular focus on OWL-based knowledge bases.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Making LLM Awareness of LUBM Ontology</title>
        <p>
          The mechanism by which the LUBM ontology was provided to the LLM is mainly based on RAG
(Retrieval-Augmented Generation) [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. RAG is an advanced technique that enhances LLMs by
dynamically integrating external, up-to-date information into their responses. Unlike traditional LLMs,
which rely solely on static training data, RAG first retrieves relevant documents or data from external
sources (e.g., databases, document repositories, or web pages) based on the user’s query. The retrieved
information is then combined with the original prompt and fed into the LLM, enabling it to generate
more accurate, contextually relevant, and factually grounded answers [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. This approach not only
improves the model’s performance and reduces the risk of hallucinations (incorrect or fabricated
information), but also allows LLMs to access domain-specific or proprietary knowledge without the need
for costly retraining or model updates [18]. In particular, Fig. 1 shows the workflow by means, in the
proposed approach, the LLM answers the user’s query layering on the LUBM ontology. More in detail,
the document provided in RAG mode is the LUBM ontology processed through text embedding and
stored into a Vector DB, namely PGVector2. Therefore, when the query arrives at the system, it is
vectorized through text embedding and used to search the knowledge base (LUBM ontology). The
search result is then attached, as context, to the original prompt (with the user’s question) and sent to
the LLM, which in turn answers the question.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Evaluation Approach</title>
        <p>In order to assess and compare the ability of LLMs to read, understand, and make inferences on OWL
Lite ontologies, a comprehensive and manually annotated set of challenging queries was designed. Such
a set is made up of 53 queries in natural language, emulating the definition of individuals and reasoning
over the LUBM ontology. Tab. 1 outlines the diferent facets of reasoning under assessment. Categorizing
the queries in this way enables a more comprehensive evaluation of results, helping to identify model
lfaws and biases based on the types of inference involved. An additional layer of categorization is
provided by the level of each query, which ranges from 1 to 3 and reflects the complexity of the logical
reasoning required to answer it. level 1 includes queries that involve direct reasoning, requiring only a
single inference step. level 2 comprises queries that necessitate two concatenated inference operations,
representing moderately complex reasoning. Finally, level 3 encompasses queries that demand advanced
reasoning, involving three concatenated inference steps. The distribution of queries across these levels
in the evaluation set is as follows: level 1 – 55%, level 2 – 38%, and level 3 – 7%. The following element
portrays an example of two items utilized in this evaluation:
"id": 13,
"premise": "X is an undergraduate student.",
"query": "Is X taking a teaching course?",
"correct answer": "Yes",
"level": 2,
"type": "instance-class equivalence"
"premise": "X is part of Y. X is part of Z.",
"query": "Is Y part of Z?",
"correct_answer": "No",
"level": 1,
"id": 29,
"type": "transitive property"
},....]</p>
        <p>Each evaluation item is composed of a premise, i.e., a statement that encodes an assertion about an
individual within the ontology, serving as the foundation for the query. As stated in the following
sections, these two fields can be considered together, forming a full_query field to pass as input to the
model for the evaluation. To ensure a robust evaluation, a diverse set of LLMs was utilized, varying both
in size and training methodology, allowing analyzing how these factors afect the ability to develop
diferent reasoning over an ontology. In this evaluation setup, unlike classical OWL reasoners, the
models were instructed to operate under a Closed World Assumption, meaning that any information
not explicitly stated in the ontology is considered false. Future work could extend the evaluation to the
Open World Assumption (OWA) to assess the models’ ability to distinguish whether a query can be
definitively answered with the available information. The models were configured to return a structured
response consisting in their final answer to the query (only replying with "yes" or "no") and their
reasoning process. The models’ coherence was then evaluated, i.e., whether the reasoning process was
consistent with the final answer.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Prompting</title>
        <p>An important part of obtaining reliable results from the LLMs consists of providing an input prompt
clearly explaining the task. The prompt was designed following common principles of prompt
engineering, such as clarity, specificity, chain of thoughts, and role assignment [ 19]. More specifically, the
defined prompt can be divided into system prompt, instructions, and user’s prompt. The system prompt,
used to recall the general knowledge of the models about OWL. Moreover, the instructions guide models
to read the LUBM ontology, retrieved from the knowledge base (provided in RAG modality), and apply
their knowledge about OWL on such ontology to answer the user’s questions. The instructions try
to avoid pre-training knowledge that possibly conflicts with the provided LUBM ontology. This also
ensures, to a certain extent, a correct evaluation of reading, comprehension and inference of models.
Lastly, instructions put also importance on aligning the reasoning process of the model with their
response, and on the structured generation of the response as a JSON object as depicted in Fig. 2.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimentation and Evaluation</title>
      <sec id="sec-4-1">
        <title>4.1. Settings</title>
        <p>As it was anticipated in section 3.4, the results were compared on a diverse set of LLMs for exploring the
diferences in the results according to model size, architecture, or training process. Only open-source
or open-weights models were considered in the set. This includes diferent version of Llama models,
spanning from smaller (Llama 3.1 8b)3 to bigger size in parameters and newer versions (Llama</p>
        <sec id="sec-4-1-1">
          <title>3https://huggingface.co/meta-llama/Meta-Llama-3-8B</title>
          <p>Model
Llama 3.1 8b
Llama 3 70b
Llama 3.3
70b versatile
Llama4-Maverick
3 70b4, Llama 3.3 70b versatile5). Also, the Llama4-Maverick6 model was included for
evaluating the performance of a Mixture of Experts architecture. The qwen-qwq-32b7 model was
added for evaluating a mid-size model incorporating Reinforcement Learning in its training process.
Lastly, we also included a model based on the Llama architecture, trained via distillation using the
Deepseek-R1 model8. Tab. 2 outlines the main features of the selected models.</p>
          <p>The phi-data framework9 was utilized for creating agents with well defined roles and contextual
knowledge. The OWL Lite ontology in textual format was chunked and embedded into vectors
utilizing the PGvector10 PostgreSQL extension for performing similarity searches. The agents were then
connected to the LLMs through the Groq11 API and Ollama12 for running local models. All experiments
were conducted on a machine equipped with 16 GB of RAM and an NVIDIA GeForce RTX 4050 Laptop
GPU with 6 GB of dedicated VRAM.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Experiments</title>
        <p>For each model in Tab. 2, an agent was created by incorporating the system prompt and instructions
as described in Section 3.5. Each agent was instructed to skip the internal step-by-step reasoning (via
the reasoning parameter), making them equivalent to a plain LLM invocation with RAG, where the
context is simply retrieved and appended to the prompt. The manually annotated item in the set of
queries was passed to the agent to obtain the structured response, which consisted of a discrete answer
(yes or no) and a rationale (referred to as "thoughts") explaining the reasoning behind the answer. To
ensure robustness in the evaluation process, a recursive function was employed for each query. This
function attempted to generate a valid structured response, and in the event of a failure (e.g., malformed
output or parsing errors), it automatically retried the generation process up to five times. This was
necessary due to occasional inconsistencies in the model’s output format. Eventually, the metric used
to evaluate performance was classification accuracy, computed by comparing the discrete answers to
the ground-truth annotations.
4meta-llama/Llama-3.1-70B
5https://console.groq.com/docs/model/llama-3.3-70b-versatile
6https://console.groq.com/docs/model/llama-4-maverick-17b-128e-instruct
7https://huggingface.co/Qwen/QwQ-32B
8https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B
9https://www.phidata.app/
10https://github.com/pgvector/pgvector
11https://console.groq.com/home
12https://ollama.com/</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Results</title>
        <p>Moreover, Fig 3 reports common errors related to the same inference types, i.e., queries incorrectly
answered by at least two models. Queries based on the transitive property lead to the highest number
of errors, suggesting models’ dificulties in applying transitive reasoning. Fig 4 reports the level of
complexity of the commonly misclassified queries, stressing that the models’ classification for these
inference types was problematic even for level 1. The figure reports that the 24.5% of level 1 and 25%
of level 2 queries were wrongly classified by more than one model. This even distribution indicates
that the errors depended more on the inference type than on the level of complexity. Also, it is notable
that there are no common level 3 errors, probably mainly due to their low representation in the set
considered.</p>
        <p>Fig. 5 presents an analysis of the errors model by model, with a focus on the underlying inference
types. While the models exhibit diferent patterns of error—indicating varied strategies for addressing
the queries—some commonalities emerge that warrant further examination.</p>
        <p>As anticipated by Fig. 3, the models’ struggles with ’transitive-property’, since it was the most present
inference type among the frequent errors, and 5 out of 6 models report at least one error of this kind.
A closer look to the models’ explanations for a query of level 1 of such a type provides details about
the models’ biases in the logical pattern. The considered query is analyzed as an example: "X is part
of Y. X is part of Z. Is Y part of Z?". The qwen-qwq-32b and llama3-70b wrongly answered with
’yes’ and presented the common patterns in their explanations. Both the models misunderstood the
logic of the transitive property, incorrectly assuming that if X is part of Y and X is part of Z, we can
conclude that Y is part of Z as well. This logical flaw is repeated over queries of this type, showing an
over extension of the reasoning, as the models are not relying just on axioms of the provided ontology,
but they apply common associative reasoning. The Llama4-Maverick model is the one providing the
best explanations, correctly recognizing the logic of transitive property and its caveats: if X is part of Y
and Y is part of Z, then X is part of Z. However, the question is asking whether Y is part of Z, given that X is
part of Y and X is part of Z. The transitivity of ’subOrganizationOf’ does not directly answer this question.</p>
        <p>To determine if ’Is the Head Of’ is a subproperty
of ’Member Of’, I need to check the subproperty
axioms in the LUBM ontology. Specifically, I’ll
look for a subPropertyOf axiom that relates
’Is the Head Of’ to ’Member Of’.</p>
        <p>To answer this question, we must first identify the properties
’Member Of’ and ’Is the Head Of’ in the LUBM ontology.</p>
        <p>Then, we should determine if ’Is the Head Of’ is defined
as a subproperty of ’Member Of’. This can be done by checking
the property hierarchies defined in the LUBM ontology.</p>
        <p>If ’Is the Head Of’ is indeed a subproperty of ’Member Of’,
it means that any instance related by ’Is the Head Of’
is also related by ’Member Of’. This relationship is
crucial for ontological reasoning and inference.</p>
        <p>The LUBM ontology defines the ’headOf’ property
as a subproperty of ’worksFor’, and ’worksFor’ is a
subproperty of ’memberOf’. Therefore, ’headOf’ is
indirectly a subproperty of ’memberOf’.</p>
        <p>To determine if ’Is the Head Of’ is a subproperty of ’Member Of’,
I need to check if there is a subproperty
axiom in the LUBM ontology that defines
’Is the Head Of’ as a subproperty of ’Member Of’.</p>
        <p>To answer this question, I need to check if ’Is the Head Of’
is a subproperty of ’Member Of’ in the LUBM ontology.</p>
        <p>This can be done by searching for the rdfs:subPropertyOf
axiom that relates ’Is the Head Of’ to ’Member Of’.</p>
        <p>To determine if ’Is the Head Of’ is a subproperty
of ’Member Of’, we need to analyze their
definitions and relationships within the LUBM ontology.
discrete
answer
no
no
yes
no
no
no</p>
        <p>A similar pattern applies for the ’instance-class equivalence which appears at least once as an error
in all the models considered. The error analysis is enlarged by examining one of the trickiest queries in
the set, being correctly classified only by the Llama4-Maverick model: ’Member Of’ is a property Is ’Is
the Head Of’ a subproperty of ’Member Of’?</p>
        <p>Models’ responses are reported in Tab. 4. In particular, Llama4-Maverick identifies the correct
properties and is able to reason considering its sub-properties. For the other models, it is notable that
although the reasoning appears generally correct, it is limited to the task description. This also appears
too vague in the case of llama3.1-8b and too verbose in the case of llama3.3-70b-versatile. This pattern
emerges consistently across misclassified queries, regardless of inference type and level of complexity,
depicting a lack of coherence between the rationale and the final answer. This is especially true for the
smallest model -llama3.1-8b- while it appears with less frequency in the newest version of llama
-llama4-Maverick- indicating greater logical understanding and coherence from a model with an
updated training and a Mixture of Experts architecture.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Future Works</title>
      <p>This work investigated how LLMs can read, understand, and perform logical inferences over an OWL
Lite ontology. The well-known LUBM ontology was employed as a reference symbolic model due
to its coverage of diverse inference types, including inverse properties, transitive properties,
somevalue restrictions, and class intersections. The LLMs’ knowledge was constrained to only the supplied
LUBM ontology, excluding prior notions from their training data. Their inference ability was then
evaluated through a dedicated set of manually annotated binary queries, allowing the assessment of
their capacity to understand and generalize the concepts expressed in the ontology and apply them to
new contexts and individuals. The proposed approach consists of two phases. The first phase makes
the OWL Lite ontology accessible to the LLM by transforming it into a dense vector representation.
The evaluation query is also transformed, and using similarity scores, the relevant context is retrieved
and assembled into a prompt for obtaining both the discrete answer and self-explanation from the LLM.
The second phase involves extensive results evaluation, considering the inference types supported by
OWL Lite. The results evaluation in Section 4.3 identified Llama4-Maverick as the best one in terms
of accuracy and quality of reasoning, evaluated through the models’ self-explanations. The results also
suggested that models are more afected by the type of reasoning than by query complexity. In particular,
’transitive-property’, was identified as the most problematic inference type. Overall, this work provides
insights into LLM reasoning when grounded in symbolic ontological knowledge, thereby contributing
to the development of semantically aware AI systems. Demonstrating that LLMs can perform reasoning
following rules and constraints expressed by an ontological schema underscores their potential to
generate grounded, explainable, and logically consistent responses. This aligns with the broader vision
of Generative eXplainable AI (GenXAI), moving beyond purely data-driven outputs toward AI systems
that are both interpretable and anchored in structured domain knowledge. Considering this work as
foundational, several future research directions are identifiable. The evaluated query set and the LLMs
input instructions are a fundamental part of the approach as they shape the models’ behavior and the
basis for the evaluation. Two important assumptions were made in this work to reduce the overall
complexity. The queries’ answers were limited to yes/no; future works may explore the behavior of
the models in a multiclass setting. Similarly, the Open World Assumption (OWA) was not considered.
For a comprehensive evaluation of LLMs as logical reasoners, it is important to assess their ability to
distinguish between queries that can be definitively answered and those that cannot be determined
under OWA.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used OpenAI GPT-4 for grammar and spelling check.
After using this service, the authors reviewed and edited the content as needed and take full responsibility
for the publication’s content.
[18] W. Zhang, J. Zhang, Hallucination mitigation for retrieval-augmented large language models: A
review, Mathematics 13 (2025). URL: https://www.mdpi.com/2227-7390/13/5/856. doi:10.3390/
math13050856.
[19] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al., Chain-of-thought
prompting elicits reasoning in large language models, Advances in neural information processing
systems 35 (2022) 24824–24837.
[20] A. Vats, R. Raja, V. Jain, A. Chadha, The evolution of mixture of experts: A survey from basics to
breakthroughs, Preprints (August 2024) (2024).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Docherty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Cooper</surname>
          </string-name>
          ,
          <article-title>Materials science in the era of large language models: a perspective, Digital Discovery 3 (</article-title>
          <year>2024</year>
          )
          <fpage>1257</fpage>
          -
          <lpage>1272</lpage>
          . URL: https://doi.org/10.1039/d4dd00074a. doi:
          <volume>10</volume>
          .1039/d4dd00074a.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chu-Carroll</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Burnham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Melville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nachman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Özcan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ferrucci</surname>
          </string-name>
          ,
          <article-title>Beyond llms: Advancing the landscape of complex reasoning</article-title>
          , arXiv (Cornell University) (
          <year>2024</year>
          ). URL: http://arxiv.org/abs/2402.08064. doi:
          <volume>10</volume>
          .48550/arxiv.2402.08064.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Albalak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , Logic-lm:
          <article-title>Empowering large language models with symbolic solvers for faithful logical reasoning</article-title>
          , arXiv (Cornell University) (
          <year>2023</year>
          ). URL: https: //arxiv.org/abs/2305.12295. doi:
          <volume>10</volume>
          .48550/arxiv.2305.12295.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Antoniou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. v.</given-names>
            <surname>Harmelen</surname>
          </string-name>
          , Web ontology language: Owl, Handbook on ontologies (
          <year>2009</year>
          )
          <fpage>91</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Beverley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Franda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Karray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maxwell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Benson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Smith</surname>
          </string-name>
          , Ontologies, arguments, and
          <article-title>large-language models (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.</given-names>
            <surname>Gibaut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pereira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Grassiotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Osorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gadioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Munõz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gomes</surname>
          </string-name>
          , C. dos
          <string-name>
            <surname>Santos</surname>
          </string-name>
          ,
          <article-title>Neurosymbolic ai and its taxonomy: a survey, arXiv</article-title>
          (Cornell University) (
          <year>2023</year>
          ). URL: https: //arxiv.org/abs/2305.08876. doi:
          <volume>10</volume>
          .48550/arXiv.2305.08876.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gaur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sheth</surname>
          </string-name>
          ,
          <article-title>Building trustworthy neurosymbolic ai systems: Consistency, reliability, explainability, and safety</article-title>
          ,
          <source>AI</source>
          Magazine
          <volume>45</volume>
          (
          <year>2024</year>
          )
          <fpage>139</fpage>
          -
          <lpage>155</lpage>
          . URL: https://doi.org/10.1002/aaai.12149. doi:
          <volume>10</volume>
          .1002/aaai.12149.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <article-title>Explainable generative ai (genxai): A survey, conceptualization</article-title>
          , and research agenda,
          <source>Artificial Intelligence Review</source>
          <volume>57</volume>
          (
          <year>2024</year>
          )
          <fpage>289</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Mondorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Plank</surname>
          </string-name>
          ,
          <article-title>Beyond accuracy: Evaluating the reasoning behavior of large language models - a survey, arXiv</article-title>
          (Cornell University) (
          <year>2024</year>
          ). URL: http://arxiv.org/abs/2404.
          <year>01869</year>
          . doi:
          <volume>10</volume>
          . 48550/arxiv.2404.
          <year>01869</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Tanneru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lakkaraju</surname>
          </string-name>
          ,
          <article-title>Faithfulness vs. plausibility: On the (un)reliability of explanations from large language models</article-title>
          , arXiv (Cornell University) (
          <year>2024</year>
          ). URL: https://arxiv. org/abs/2402.04614. doi:
          <volume>10</volume>
          .48550/arXiv.2402.04614.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Horrocks</surname>
          </string-name>
          ,
          <article-title>Language model analysis for ontology subsumption inference</article-title>
          ,
          <year>2023</year>
          . URL: https://arxiv.org/abs/2302.06761. arXiv:
          <volume>2302</volume>
          .
          <fpage>06761</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <article-title>Can large language models understand dl-lite ontologies? an empirical study</article-title>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2406.17532. arXiv:
          <volume>2406</volume>
          .
          <fpage>17532</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Beckett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          , E. Prud'hommeaux, G. Carothers, Rdf
          <volume>1</volume>
          .1 turtle, World Wide Web Consortium (
          <year>2014</year>
          )
          <fpage>18</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heflin</surname>
          </string-name>
          ,
          <article-title>Lubm: A benchmark for owl knowledge base systems</article-title>
          ,
          <source>Web Semant</source>
          .
          <volume>3</volume>
          (
          <year>2005</year>
          )
          <fpage>158</fpage>
          -
          <lpage>182</lpage>
          . URL: https://doi.org/10.1016/j.websem.
          <year>2005</year>
          .
          <volume>06</volume>
          .005. doi:
          <volume>10</volume>
          .1016/j.websem.
          <year>2005</year>
          .
          <volume>06</volume>
          .005.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bellatreche</surname>
          </string-name>
          , G. Fokou,
          <string-name>
            <given-names>M.</given-names>
            <surname>Baron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khouri</surname>
          </string-name>
          , Ontodbench:
          <article-title>Novel benchmarking system for ontology-based databases</article-title>
          , in: R. Meersman,
          <string-name>
            <given-names>H.</given-names>
            <surname>Panetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dillon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rinderle-Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dadam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pearson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferscha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. F.</given-names>
            <surname>Cruz</surname>
          </string-name>
          (Eds.),
          <source>On the Move to Meaningful Internet Systems: OTM 2012</source>
          , Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2012</year>
          , pp.
          <fpage>897</fpage>
          -
          <lpage>914</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dou</surname>
          </string-name>
          , T.-Y. Ho,
          <string-name>
            <given-names>P.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Trustworthiness in retrieval-augmented generation systems: A survey</article-title>
          ,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2409.10102.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>A survey on retrieval-augmented text generation for large language models</article-title>
          ,
          <source>ArXiv abs/2404</source>
          .10981 (
          <year>2024</year>
          ). URL: https://api.semanticscholar.org/CorpusID:269188036.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>