<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GSS: Graph Semantic Summarization Using Answer-Centered Explanatory Subgraphs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Valantis Zervos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Spyridon Doukeris</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kiki Miniadou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giannis Vassiliou</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sophia Sideri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haridimos Kondylakis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CSD, University of Crete</institution>
          ,
          <addr-line>Heraklion</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Electrical and Computer Engineering</institution>
          ,
          <addr-line>HMU, Heraklion</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>FORTH-ICS</institution>
          ,
          <addr-line>Heraklion</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>LIPADE, Université Paris Cité</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <abstract>
        <p>Knowledge graphs support large-scale structured knowledge representation, but their size and semantic complexity make query-driven access and summarization challenging. Public resources such as DBpedia and Wikidata contain billions of triples, complicating the extraction of contextually meaningful information for a given query. We present GSS (Graph Semantic Summarization), a query-driven framework for constructing answer-centered explanatory subgraphs (ACE) from large knowledge graphs. Given a natural language query, GSS combines large language models with symbolic SPARQL querying to identify salient entities, retrieve relevant triples, and assemble compact subgraphs centered on answer-bearing entities while preserving explanatory context. We evaluate GSS on the QALD-9-plus benchmark using gold-standard answers to query DBpedia and Wikidata. The evaluation combines answer-containment, ranking, and intrinsic graph-quality metrics to assess both correctness and explanatory structure. Experiments with two LLMs, show that GSS retrieves correct answers and produces semantically coherent summaries that enhance contextual understanding.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>A knowledge graph is a structured representation of real-world entities, such as people, places, objects,
and abstract concepts, and the semantic relationships between them. By modeling information as nodes
(entities) and edges (relations), knowledge graphs provide contextualized, machine-readable data that
supports reasoning, semantic search, and intelligent applications such as question answering systems
and AI assistants.</p>
      <p>
        Motivation. Knowledge graphs have become a fundamental infrastructure for organizing and accessing
structured knowledge across a wide range of domains, including the Semantic Web, digital libraries,
scientific data management, and intelligent assistants. Modern knowledge graphs often contain millions
or billions of triples, enabling expressive semantic querying but also making efective information access
increasingly challenging. Users are rarely interested in exhaustive query results; instead, they seek
concise, interpretable answers accompanied by contextual information that explains and situates those
answers within the broader graph. This need is especially pronounced in interactive and query-driven
settings, where summaries must be generated on demand and adapted to diverse information needs.
Problem Statement. Knowledge Graph Summarization (KGS) addresses this challenge by producing
compact and informative representations of large and complex knowledge graphs. Instead of exposing
users to raw triples, KGS aims to reduce information overload by selecting the most relevant entities and
relationships while preserving the semantic structure necessary for understanding [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This problem is
especially critical in scenarios such as question answering, exploratory search, and decision support,
where users require not only correct answers but also explanatory context that situates those answers
within the broader graph. The challenge becomes even more pronounced when summaries must be
generated dynamically in response to diverse natural language queries, without relying on historical
usage patterns or predefined templates like previous works.
Formalization. Formally, given a knowledge graph  with a set of triples  and a user query
, the objective is not only to retrieve a subset of triples  ⊂  that directly answer the query,
but also to identify a set of relevant triples  ⊂  that together form a subgraph  ⊂  . This
subgraph should be compact, semantically coherent, and centered around the ACE triples, providing
explanatory context that supports interpretation of the answer. For example, given the question
 = “What is the capital of France?”, the goal is to construct a subgraph  that includes the ACE triple
&lt;Paris, capitalOf, France&gt; along with additional contextual triples such as &lt;Paris, type, City&gt; or &lt;France,
locatedIn, Europe&gt;, as illustrated in Figure 1.
      </p>
      <p>
        Usefullness. From the perspective of large language models (LLMs), answer-centered explanatory
subgraphs provide a principled mechanism for grounding natural language generation in explicit
symbolic evidence. By constraining the model’s input to a compact, query-focused subgraph centered
on ACE triples, the reasoning space is reduced to semantically relevant entities and relations, improving
controllability and mitigating spurious correlations that may arise from unstructured or excessively
large contexts. This structured grounding enables hybrid neuro-symbolic setups in which LLMs are
responsible for semantic interpretation and language generation, while correctness and contextual
consistency are enforced through graph-based retrieval [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. As a result, the generated answers can
be more faithfully aligned with the underlying knowledge graph, supporting transparent reasoning,
reduced hallucination, and explainable outputs that can be traced back to explicit triples and short
relational paths within the subgraph
Limitations of Existing Approaches. Existing approaches to knowledge graph summarization
typically rely on graph-theoretic measures, handcrafted heuristics, predefined summary templates, or
historical query workloads [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. While efective in constrained settings, these methods often struggle to
adapt to diverse natural language queries or to accurately capture user intent. Moreover, approaches
that depend on interaction logs or query histories [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] assume the availability of large and representative
workloads—an assumption that rarely holds for long-tail entities, emerging knowledge graphs, or
specialized domains. As a result, generating meaningful summaries for rare entities or previously
unseen information needs remains an open challenge.
      </p>
      <p>Large language models (LLMs) can address these limitations. LLMs can interpret natural language
queries and can extract semantically salient queries from user input, without requiring prior query logs.
When combined with symbolic querying over knowledge graphs, LLMs can support the construction of
query-focused summaries that balance correctness with explanatory richness.</p>
      <p>Our Approach. Motivated by these observations, we propose GSS (Graph Semantic Summarization),
a query-driven framework for constructing answer-centered explanatory subgraphs (ACE) from large KGs.
GSS integrates LLM-based keyword extraction with symbolic SPARQL querying to identify ACE triples
and organize surrounding contextual information into compact, query-focused summaries. Rather
than relying on historical query logs or predefined templates, GSS operates on demand, producing
summaries that are anchored around correct answers while enriching them with semantically relevant
context.</p>
      <p>We demonstrate GSS in real KGs using DBpedia and Wikidata. The system is evaluated with the
QALD-9-plus benchmark, leveraging gold-standard answer entities as anchor points to test whether
generated summaries contain the correct answers and structure additional explanatory information
around them. We report answer-containment and ranking metrics alongside intrinsic graph-quality
measures, capturing both correctness and explanatory coherence.</p>
      <p>Contributions. The main contributions of this work are summarized as follows:
• We propose GSS, a query-driven knowledge graph summarization framework that constructs ACE
subgraphs by integrating large language models with symbolic SPARQL querying over DBpedia and
Wikidata.
• We introduce a hybrid scoring and ranking strategy that combines LLM-derived entity importance
with embedding-based semantic similarity to prioritize answer-relevant and contextually meaningful
triples.
• We present an evaluation methodology that treats summaries as ACE, assessing answer containment
and intrinsic graph properties such as path consistency, density, and the presence of bridge entities.
• We empirically evaluate GSS on the QALD-9-plus benchmark, demonstrating that the framework
retrieves correct answers while producing compact, semantically coherent subgraphs that enhance
contextual understanding.</p>
      <p>The remaining part of this paper is structured as follows: Section 2 presents related work, whereas
our approach is presented in Section 3. Section 4 evaluates our approach, whereas Section 4.2 discusses
the main findings. Finally, Section 5 concludes this paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Following the taxonomy of Cebirić et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], semantic graph summarization methods can be broadly
classified into quotient and non-quotient approaches. Quotient-based methods merge nodes with similar
types or structural properties to produce compact, schema-level summaries, but are inherently
coarsegrained and unsuitable for query-specific or instance-level information needs. In contrast, non-quotient
methods construct summaries as subgraphs of the original knowledge graph, retaining selected entities
and relations based on relevance or task-specific criteria. Since GSS generates query-focused subgraphs
that preserve original entities, it falls into the class of structural, non-quotient summarization approaches.
Comprehensive surveys of the area are provided in [
        <xref ref-type="bibr" rid="ref1 ref4">1, 4</xref>
        ].
      </p>
      <p>Early non-quotient summarization work includes the RDF Sentence Graph model of Zhang et al. [5],
which constructs summaries based on overlapping RDF sentences guided by user-defined weights and
preferences. While expressive, this approach requires manual parameter tuning and assumes explicitly
specified relevance signals, limiting its applicability in open or large-scale settings.</p>
      <p>Subsequent work explored personalized ontology summarization. Queiroz-Sousa et al. [6] rank
ontology concepts using structural measures and connect them via the Broaden Relevant Paths algorithm,
relying on user-defined constraints to guide personalization. However, these methods operate primarily
at the schema level and require explicit user input, making them less suitable for ad-hoc or natural
language queries.</p>
      <p>
        More recent approaches incorporate usage information. GLIMPSE [7] maximizes inferred user utility
based on representative query sets, but depends on users providing such queries and sufers from
scalability issues on large graphs. Workload-based methods such as WBSUM [8] and iSummary [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
infer relevance from query logs to construct schema-level or selective summaries. While efective when
rich workload data is available, these methods depend on large, representative query logs that are often
unavailable or heavily skewed, limiting their efectiveness for non-popular or long-tail entities.
      </p>
      <p>GSS addresses these limitations by adopting an on-demand, query-driven approach. Instead of relying
on historical workloads or explicit user parameters, GSS leverages large language models to extract
salient entities directly from a natural language query and uses symbolic SPARQL querying to assemble
a compact, semantically coherent subgraph. This design eliminates the need for query logs, enabling
efective summarization for rare entities and previously unseen information needs.</p>
    </sec>
    <sec id="sec-3">
      <title>3. The GSS Approach</title>
      <p>Problem Setting. Let  = (, ) be a knowledge graph, where  is the set of entities,  the set
of predicates, and  ⊆  ×  ×  is the set of RDF triples. Let  denote the set of all triples in a
. Given a natural language query , the goal of GSS is to construct a subgraph  = (, ) such
that  ⊆  ,  ⊆  , and  is both compact and semantically aligned with the intent of . In this
work, we focus on summaries that are explicitly centered around answer entities and enriched with
explanatory context.</p>
      <p>
        Definition 1 (ACE Subgraph). Given a query  and a knowledge graph , an answer-centered
explanatory (ACE) subgraph is a subgraph  ⊆  that: (i) contains at least one ACE triple for ,
whenever such triples exist in , and (ii) includes additional triples that are semantically connected
to the answer entities and provide explanatory context for interpreting the answer.
Definition 2 (Entity Importance). Given a query , an entity importance function  :  → [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]
assigns to each entity a score reflecting its relevance to .
      </p>
      <p>Algorithm 1 Graph Semantic Summarization (GSS)
Definition 3 (Triple Importance). Given a triple  = (, , ) and an entity importance function ,
the importance-based score of  is defined as imp() = max{(), ()}, capturing the intuition
that a triple is important if it involves at least one highly relevant entity.</p>
      <p>Definition 4 (Semantic Similarity). Let (·) be a sentence embedding function that maps text to a vector
space. The semantic similarity between a query  and a triple  is defined as sim() = cos((), ()),
where () is computed from a textual serialization of the triple.</p>
      <p>
        Definition 5 (Triple Relevance). The overall relevance of a triple  to a query  is defined as a convex
combination () =  ·  imp() + (1 − ) ·  sim(), where  ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] controls the trade-of between
entity-centric importance and semantic alignment.
      </p>
      <sec id="sec-3-1">
        <title>3.1. The GSS Algorithm</title>
        <p>Algorithm 1 presents the overall procedure executed by GSS for a single query, producing an ACE
subgraph. In the sequel, we explain Algorithm 1 step by step, following the pipeline shown in Figure 2.</p>
        <p>
          Given a natural language query , GSS first extracts a set of salient entity and concept mentions 
using an LLM (line 1). These mentions are assigned importance scores through the function  (line 2),
which reflects their relevance to the query intent as defined above. The output of this step is a mapping
from KG entities to normalized importance values in [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ]. The extracted mentions are then grounded
to the knowledge graph by resolving them to KG identifiers, yielding the set  (line 3). Using these
identifiers, GSS retrieves an initial set of triples 0 that involve the resolved entities (line 4), ensuring
that ACE triples are considered early in the process.
        </p>
        <p>To provide explanatory context around answer candidates, GSS performs a controlled one-hop
expansion over the neighborhood of 0, producing an additional set of triples 1 (line 5). The candidate
pool  = 0 ∪ 1 (line 6) balances contextual coverage with compactness by restricting expansion
depth.</p>
        <p>For each candidate triple  = (, , ) ∈  (line 7), GSS computes two complementary relevance
signals. The importance-based score imp() (line 8) is derived from the entity importance values ()
and (), capturing the degree to which the triple involves answer-relevant entities. In parallel, the
semantic similarity score sim() (line 9) measures the embedding-based semantic alignment between
the query  and the textual representation of .</p>
        <p>These signals are combined into a single relevance score () =  ·  imp() + (1 − ) ·  sim()
(line 10), where  is a user-configurable parameter controlling the relative emphasis on entity importance
versus semantic similarity. In our experiments,  is fixed across queries to ensure deterministic and
reproducible behavior.</p>
        <p>After scoring all candidate triples (lines 7–11), GSS selects the top-ranked subset  ′ under the specified
size budget  (line 12). The entity set of the output subgraph is defined as the set of subjects and objects
appearing in  ′ (line 13). Finally, duplicate or equivalent triples are removed to reduce redundancy
(line 14), and the resulting ACE subgraph  = (,  ′) is returned (line 15).</p>
        <p>Example 1. We illustrate the behavior of GSS using the running example query:  =
“What is the capital of France?”.</p>
        <p>Entity Extraction and Importance Assignment. In Line 1 of Algorithm 1, GSS utilizes LLMs to
identify salient entities from the query. In this case, the extracted set is  = {France, capital}. Then,
Line 2 assigns normalized importance scores reflecting query relevance. The explicitly mentioned entity
France receives the highest importance score (France) = 1.0, while the concept capital is treated as a
strongly implied relation indicator (capital) = 0.8.</p>
        <p>URI Resolution and Initial Triple Retrieval. In Line 3, entities are resolved to knowledge graph
identifiers. For DBpedia, France is mapped to dbr:France. Using this identifier, GSS retrieves an initial
set of triples 0 (Line 4) that involve dbr:France:
0 = ⎨⎧ ⟨⟨FPraarnics,ec,laopciattaeldOIf,nF,rEaunrcope⟩e,⟩, ⎬⎫</p>
        <p>⎩ ⟨France, populationTotal, 67,000,000⟩, . . . ⎭
Among these, the triple ⟨Paris, capitalOf, France⟩ is an ACE triple.</p>
        <p>One-Hop Expansion for Explanatory Context. To enrich the answer with explanatory context,
GSS performs a one-hop expansion over 0, producing an additional set 1 (Line 5). This expansion
retrieves triples connected to entities appearing in 0, such as:
1 = ⎨⎧ ⟨⟨PPaarriiss,, ltoycpae,tCeidtIyn⟩,,France⟩, ⎬⎫</p>
        <p>⎩ ⟨Europe, type, Continent⟩, . . . ⎭
The candidate triples are formed as  = 0 ∪ 1 (Line 6). We empirically observed that further
expansions rapidly increase noise without improving explanatory coherence.</p>
        <p>Triple Scoring. For each triple  ∈  , GSS computes two relevance signals. The importance-based
score imp() (line 8) is derived from the importance values of the entities participating in the triple. For
example: imp(⟨Paris, capitalOf, France⟩) = max{(Paris), (France)} = 1.0. In contrast,
a peripheral triple such as ⟨France, populationTotal, 67,000,000⟩ receives a lower
importancebased score, as it does not directly support the query intent.</p>
        <p>Following, GSS computes the semantic similarity score sim() (Line 9) by embedding both the query
 and a textual serialization of  into a vector space and computing cosine similarity. The final score
() (Line 10) is computed as a combination of importance and semantic similarity.
Ranking and Output Subgraph Construction. After scoring all triples, GSS selects the top-ranked
subset  ′ under the size budget  (Line 12). In this example,  ′ contains:
⎧⎨ ⟨Paris, capitalOf, France⟩, ⎬⎫</p>
        <p>⟨Paris, type, City⟩,
⎩ ⟨France, locatedIn, Europe⟩ ⎭</p>
        <p>The resulting entity set  (Line 13) includes Paris, France, and Europe. After deduplication
(Line 14), the final output is an ACE subgraph that is anchored around the correct answer Paris and
augmented with minimal yet semantically coherent contextual information explaining its role and
location.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Properties of the GSS Algorithm</title>
        <p>Obviously, the explanatory subgraph produced by GSS depends only on the input query  and the
underlying knowledge graph , and does not rely on historical query logs. We now discuss key
properties of GSS that characterize its behavior and suitability for ACE summarization. These properties
follow directly from the design of Algorithm 1
Lemma 1 (Answer Anchoring). Whenever an ACE triple exists in  for a query , GSS prioritizes its
inclusion in the resulting explanatory subgraph.</p>
        <p>Proof Sketch. ACE triples involve entities mentioned or implied by the query. Such entities receive high
importance scores , which propagates to an important triple. In addition, ACE triples typically exhibit
strong semantic similarity to the query. As a result, their combined relevance score is high, ensuring
that they are ranked early and selected during the top- selection step, provided they exist in .
Lemma 2 (Long-Tail Coverage). GSS can construct ACE subgraphs for entities that have not appeared in
prior queries, as long as they are present in the knowledge graph .</p>
        <p>Proof Sketch. Because GSS does not rely on historical query logs or usage statistics, entity relevance is
inferred directly from the semantic content of the input query via LLM-based extraction. As long as an
entity is present in  and can be resolved during URI grounding, it can participate in triple retrieval
and expansion. Consequently, entities from the long tail of the knowledge graph are treated on equal
footing with frequently queried entities.</p>
        <p>Lemma 3 (Bounded Summary Size). For fixed configuration parameters, the size of the explanatory
subgraph produced by GSS is upper-bounded by a constant independent of ||.</p>
        <p>Proof Sketch. Algorithm 1 enforces explicit limits on both expansion depth (one-hop expansion) and
the number of retained triples via the size budget . These parameters bound the number of triples
that can be included in the output regardless of the total size of . Therefore, the size of the resulting
explanatory subgraph is independent of || and depends only on configuration choices.
Lemma 4 (Explanatory Coherence). All non-answer entities in the output subgraph are connected to at
least one answer entity by a path of length at most two.</p>
        <p>Proof Sketch. All triples in the output subgraph originate from the initial retrieval step, which includes
ACE triples, or from the one-hop expansion around those triples. Consequently, any non-answer entity
appears either directly in an ACE triple or in a triple adjacent to such an entity. This guarantees that
every non-answer entity is connected to at least one answer entity by a path of length at most two,
ensuring explanatory coherence.</p>
        <p>Computational Complexity. Let  be the number of entities extracted from the query,  the average
degree of an entity in , and  the number of retrieved triples after expansion. Entity extraction
and scoring are constant with respect to ||. Triple retrieval and one-hop expansion require ( · )
SPARQL pattern matches. Semantic similarity computation requires () embedding comparisons.
Ranking and selection run in ( log ) time. Overall, the algorithmic complexity of GSS with respect
to the knowledge graph is ( ·  +  log ) , excluding the inference latency of the LLM model. Since
, , and  are bounded by configuration parameters, GSS operates independently of the total size of
the knowledge graph, making it suitable for large-scale deployments.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <p>The objective of the evaluation is to assess whether GSS produces efective Answer-Centered Explanatory
(ACE) subgraphs, i.e., compact graph structures that (i) contain correct answer entities for a given
query and (ii) organize additional contextual information around them in a semantically coherent and
interpretable manner. To this end, we adopt a multi-layer evaluation framework that jointly assesses
answer faithfulness and explanatory structure.</p>
      <p>Evaluation Setup. All experiments are conducted on CPU-only hardware using pre-trained
transformer models to ensure reproducibility. We evaluate GSS under two LLM configurations, namely
gemini-2.5-flash-lite and llama-3.1-8b-instant. For each query, GSS produces a ranked
list of candidate RDF triples. To simulate a realistic summarization setting while maintaining
computational feasibility, only the top  = 15 ranked triples per query are retained. These candidates are
subsequently reranked using a cross-encoder model (cross-encoder/ms-marco-MiniLM-L-6-v2),
which jointly encodes the natural language query and a textual representation of each triple. The
reranked triples constitute the final ACE subgraph evaluated for each query.</p>
      <p>Datasets and Benchmarks. We evaluate GSS on DBpedia and Wikidata using the QALD-9-plus
benchmark. QALD-9-plus provides natural language questions paired with gold-standard SPARQL
queries and answer bindings, but does not define gold-standard summaries or contextual graphs. As
a result, QALD cannot be used to directly assess the correctness of a generated summary as a whole.
Instead, gold answer entities are treated as anchor points, and the evaluation focuses on whether the
produced subgraphs (i) contain correct answer entities and (ii) organize additional information around
them in a coherent explanatory structure. Gold answer entities are extracted from the QALD JSON
ifles by retaining the corresponding DBpedia and Wikidata resource URIs. Questions whose answers
consist exclusively of literals or empty bindings are excluded to ensure compatibility with graph-based
analysis.</p>
      <p>Evaluation Metrics. Our evaluation metrics are organized into two conceptual layers, reflecting the
dual objective of ACE summarization.</p>
      <p>Layer 1: Answer Faithfulness. These metrics assess whether the generated subgraphs contain correct
answer entities and whether they are ranked early:
• Precision, Recall, and F1-score, computed over predicted entities.
• Mean Reciprocal Rank (MRR), based on the rank of the first ACE triple.
• Hits@k, measuring whether at least one gold answer entity appears within the top- ranked triples.
These metrics answer the question: Does the summary contain the correct answer entities, and are they
prioritized early in the ranking?
Layer 2: Explanatory Structure and Answer Containment. These metrics evaluate how efectively
the subgraph organizes contextual information around answer entities:
• Path Consistency, capturing whether non-gold entities participate in coherent paths leading to an
answer entity.
• Graph Density, measuring the structural compactness of the subgraph as the ratio of observed
triples to the maximum possible number of triples between its entities; higher values indicate more
tightly connected summaries.
• Bridge Entities, defined as non-gold entities that lie on at least one path connecting another entity
in the subgraph to a gold answer entity, capturing intermediate entities that support explanatory
connectivity.
• AvgDist, measuring the average shortest-path distance from non-gold entities to the nearest gold
answer entity (lower is better). Let  = (, ) be the generated subgraph for a query,  ⊆  
the set of gold answer entities, and  =  ∖  the non-gold entities. AvgDist is defined as:
1</p>
      <p>∑︁
AvgDist =</p>
      <p>
        min dist (, )
|| ∈ ∈
Together, these metrics assess whether the generated subgraphs function as explanations rather than
mere containers of answer entities. For summaries consisting of a single triple, several structural metrics
become degenerate by construction and should therefore be interpreted with care.
Competitors. We compare GSS against iSummary [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a workload-aware approach that constructs
query-focused summaries using historical query logs. Unlike GSS, iSummary is deterministic and
always returns a triple patterns appearing in the queries. As a result, iSummary is optimized for answer
retrieval rather than explanation, serving as a strong baseline for answer faithfulness but a limited one
for explanatory structure.
Path Consistency
Graph Density
Bridge Entities
AvgDist (↓ better)
# Avg. Triples
      </p>
      <sec id="sec-4-1">
        <title>4.1. Results</title>
        <p>efective alignment between the semantic intent of the query and the extracted graph content. In
contrast, iSummary achieves higher Precision, MRR, and Hits@k scores. This is expected, as iSummary
deterministically returns one or more triple patterns from an input query, which places the answer at
rank one and optimizes ranking-based metrics. However, it significantly limits the and explanatory
context, as no additional entities are introduced.</p>
        <p>These limitations are reflected in the explanatory metrics. GSS achieves lower AvgDist values,
indicating that contextual entities are tightly centered around gold answer entities through short
explanatory paths. Moreover, the presence of bridge entities shows that GSS has intermediate nodes
that introduce explanations rather than returning single answers. By contrast, iSummary does not have
explanatory behavior due to the absence of multi-hop structure, resulting in higher AvgDist values and
zero bridge entities.</p>
        <p>Ablation Study. To examine GSS with respect to its relevance scoring, we conduct an ablation study by
changing the weighting parameter  that controls the trade-of between entity importance and semantic
similarity. Results are shown in Table 2 for  ∈ {0, 0.25, 0.5, 0.75, 1} . When  = 0 , it relies exclusively
on semantic similarity, producing semantically related triples but weak summaries, as evidenced by
higher AvgDist values and a reduced number of bridge entities. In contrast,  = 1
prioritizes the
importance of the entity alone, leading to smaller summaries with limited explanatory content.</p>
        <p>Intermediate values consistently outperform the two corner cases. Values in the range  ∈ [0.5, 0.75]
have the best balance between answer faithfulness and explanatory structure, minimizing AvgDist while
maximizing the presence of bridge entities. This behavior confirms that efective ACE summarization
requires considering both entity importance and semantic alignment. Based on these observations, we
ifx  = 0.6 for the experiments of Table 1.</p>
        <p>Overall, the results highlight a distinction between answer retrieval and answer-centered explanation.
While iSummary excels at identifying correct answers, GSS produces compact yet structured subgraphs
that contextualize answers within semantically coherent explanations.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Discussion</title>
        <p>The evaluation highlights a clear trade-of between answer retrieval and explanation generation.
Methods such as iSummary, which deterministically return a single triple, are highly efective at surfacing
correct answers and therefore perform well on ranking-based metrics such as MRR and Hits@k.
However, these methods inherently lack the ability to contextualize answers, resulting in poor recall and the
absence of explanatory structure.</p>
        <p>In contrast, GSS is explicitly designed to produce answer-centered explanatory subgraphs. By
incorporating contextual entities and multi-hop connections, GSS sacrifices strict answer compactness
in favor of interpretability and semantic coherence. This design choice is reflected in higher recall, lower
AvgDist values, and the consistent presence of bridge entities, indicating that supporting information is
meaningfully organized around the answer.</p>
        <p>Importantly, several structural metrics become degenerate for single-triple summaries and should
not be interpreted as indicators of explanation quality in such cases. Metrics such as AvgDist, path
consistency, and bridge entities are therefore critical for distinguishing between answer-centric retrieval
and explanation-oriented summarization.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>We introduced GSS, a query-driven approach for generating Answer-Centered Explanatory (ACE)
subgraphs over large knowledge graphs. Unlike traditional answer retrieval or extractive summarization
methods, GSS produces compact yet structured subgraphs that not only contain correct answers but
also provide meaningful contextual explanations. Through a comprehensive evaluation on DBpedia and
Wikidata using the QALD-9-plus benchmark, we demonstrated that GSS efectively balances answer
faithfulness with explanatory structure. While baseline methods such as iSummary achieve strong
performance on ranking-based metrics, they fail to provide explanatory context due to their single-triple
output. In contrast, GSS consistently generates multi-hop subgraphs that position contextual entities
close to gold answers, improving interpretability and supporting downstream analysis. These results
suggest that answer-centered explanation should be treated as a distinct objective from answer retrieval,
requiring dedicated evaluation metrics and summarization strategies.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>The authors have not employed any Generative AI tools.
[5] X. Zhang, G. Cheng, Y. Qu, Ontology summarization based on rdf sentence graph, in: WWW, 2007.
[6] P. O. Queiroz-Sousa, A. C. Salgado, C. E. S. Pires, A method for building personalized ontology
summaries, J. Inf. Data Manag. 4 (2013) 236–250.
[7] T. Safavi, C. Belth, L. Faber, D. Mottin, E. Müller, D. Koutra, Personalized knowledge graph
summarization: From the cloud to your pocket, in: ICDM, 2019.
[8] G. Vassiliou, G. Troullinou, N. Papadakis, H. Kondylakis, Wbsum: Workload-based summaries for
RDF/S kbs, SSDBM, 2021.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Š.</given-names>
            <surname>Čebirić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Goasdoué</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kondylakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kotzinos</surname>
          </string-name>
          , I. Manolescu, G. Troullinou,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zneika</surname>
          </string-name>
          ,
          <article-title>Summarizing semantic graphs: a survey</article-title>
          ,
          <source>The VLDB journal 28</source>
          (
          <year>2019</year>
          )
          <fpage>295</fpage>
          -
          <lpage>327</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Skrlj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koloski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pollak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lavrac</surname>
          </string-name>
          ,
          <article-title>From symbolic to neural and back: Exploring knowledge graph-large language model synergies</article-title>
          ,
          <source>CoRR abs/2506</source>
          .09566 (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Vassiliou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alevizakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Papadakis</surname>
          </string-name>
          , H. Kondylakis,
          <article-title>isummary: Workload-based, personalized summaries for knowledge graphs</article-title>
          ,
          <source>ESWC</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Kellou-Menouer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kardoulakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Troullinou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Kedad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Plexousakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kondylakis</surname>
          </string-name>
          ,
          <article-title>A survey on semantic schema discovery</article-title>
          ,
          <source>VLDBJ</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>