<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Clinical Oncology 40 (2022) 9058-9058.
doi:10.1200/JCO.2022.40.16\_suppl.9058.
[28] D. P. Carbone</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.3115/v1/P15-1067</article-id>
      <title-group>
        <article-title>VISE: Validated and Invalidated Symbolic Explanations for Knowledge Graph Integrity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Disha Purohit</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yashrajsinh Chudasama</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Torrente</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria-Esther Vidal</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hospital Universitario Puertade Hierro-Majadahonda</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>L3S Research Center</institution>
          ,
          <addr-line>Hannover</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Leibniz University Hannover</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>TIB-Leibniz Information Centre for Science and Technology</institution>
          ,
          <addr-line>Hannover</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>1</volume>
      <fpage>687</fpage>
      <lpage>696</lpage>
      <abstract>
        <p>Knowledge graphs (KGs) are naturally capable of capturing the convergence of data and knowledge, thereby making them highly expressive frameworks for describing and integrating heterogeneous data in a coherent and interconnected manner. However, based on the Open World Assumption (OWA), the absence of information within KGs does not indicate falsity or non-existence; it merely reflects incompleteness. The process of inductive learning over KGs involves predicting new relationships based on existing factual statements in the KG, utilizing either numerical or symbolic learning models. Recently, Knowledge Graph Embedding (KGE) and symbolic learning have received considerable attention in various downstream tasks, including Link Prediction (LP). LP techniques employ latent vector representations of entities and their relationships in KGs to infer missing links. Furthermore, as the quantity of data generated by KGs continues to increase, the necessity for additional quality assessment and validation eforts becomes more apparent. Nevertheless, state-of-the-art KG completion approaches fail to consider the quality constraints while generating predictions, resulting in the completion of KGs with erroneous relationships. The generation of accurate data and insights is of vital importance in the context of healthcare decision-making, including the processes of diagnosis, the formulation of treatment strategies, and the implementation of preventive actions. We propose a hybrid approach, VISE, which adopts the integration of symbolic learning, constraint validation, and numerical learning techniques. VISE leverages KGE to capture implicit knowledge and represent negation in KGs, thereby enhancing the predictive performance of numerical models. Our experimental results demonstrate the efectiveness of this hybrid strategy, which combines the strengths of symbolic, numerical, and constraint validation paradigms. VISE implementation is publicly accessible on GitHub (https://github.com/SDM-TIB/VISE).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Knowledge Graphs</kwd>
        <kwd>Symbolic Learning</kwd>
        <kwd>SHACL Constraints</kwd>
        <kwd>Numerical Learning</kwd>
        <kwd>Explainability</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>EXPLIMED - First Workshop on Explainable Artificial Intelligence for the medical domain - 19-20 October 2024, Santiago
de Compostela, Spain
* Corresponding author.</p>
      <p>Male
hasAgeCategory</p>
      <p>CurrentSmoker</p>
      <p>hasGender
Known False Facts
ageCategory</p>
      <p>Male
Young
hasRelapse</p>
      <p>Yes
Unknown False Facts
hasCancer</p>
      <p>Breast</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Knowledge Graphs (KGs) are rich structured data model that represents real-world information
in the form of entities and relations that efectively merge data and knowledge through factual
statements [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. However, KGs are not complete based on the Open World Assumption
(OWA) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] principle. The process of inductive learning over KGs encompasses a variety of
techniques for the acquisition of knowledge within KGs that will facilitate the completion
of KGs. Inductive learning is crucial for detecting missing links in KGs; it includes deducing
patterns and relationships from the existing KG. Established approaches can learn symbolic
or numerical representations of KGs’ patterns, which correspond to the fundamental building
blocks for inferring missing links [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], thus, completing KGs efectively.
      </p>
      <p>
        KGE methods project entities and relations from a KG into a lower-dimensional vector space
while preserving their semantic significance. Existing KGE approaches [
        <xref ref-type="bibr" rid="ref5 ref6 ref7 ref8">5, 6, 7, 8</xref>
        ] have
demonstrated promising results in various knowledge acquisition tasks, including link prediction,
entity recognition, relation extraction, etc. Training KGE models typically involves ranking
observed (positive) instances higher than unobserved (negative) instances. However, since KGs
only provide positive instances, it becomes essential to generate negative instances [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] that
can enable the model to learn intricate and valuable semantics. As illustrated in Figure 1, the
potential for prediction under the incomplete nature of KGs is demonstrated by the example of
a lung cancer patient and the relationships between various characteristics of this patient in
the KG. The four quadrants are depicted on the right of Figure 1, which includes predictions
or missing facts that can be classified as either known true facts, known false facts, unknown
true facts, or unknown false facts. The category of Known True Facts represents the facts of the
patient that are already present in the KG. For example, we know that the patient is male. The
category of Known False Facts refers to facts that are known but are not true. The categories
of Unkown True Facts and Unkown False Facts are the missing facts that are often predicted by
symbolic learning or numerical learning. Conversely, traditional KGs do not explicitly
represent negated facts or relationships. Instead, they concentrate on representing positive facts or
relationships between entities in the KG. While this strategy simplifies the representation and
querying processes, it also excludes the representation of negated facts in KGs, which impairs
the performance of downstream tasks, for example, link prediction (LP) for KG completion.
Inductive learning techniques struggle to learn from only positive data in the KGs, resulting in
poor predicting performance. For instance, knowing the positive facts ⟨Patient X,
hasAgeCategory, Young⟩, and ⟨Immunotherapy, hasDrug, Vinorelbine⟩, a KGE model could predict ⟨Patient
X, hasRelapse, No Relapse⟩. Nonetheless, the latent vector representation of the entities and
their relationships are not self-explanatory. Extracting explanations eficiently for the inductive
abilities remains an outstanding research challenge.
      </p>
      <p>
        The problem of explaining the LP has received significant attention in critical domains like
healthcare. Various approaches [
        <xref ref-type="bibr" rid="ref7">7, 10, 11, 12</xref>
        ] attempt to understand the inner mechanism of
such inductive learning techniques, but they are unable to capture the insights of the model
behavior with negated facts. We follow Rossi et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] vocabulary and extract explanations
for LP problems. The necessity and suficiency of explanations can be characterized in several
ways. For instance, the addition of a set of facts to a knowledge graph (KG) for an entity can
lead to the model making a prediction, whereas the absence of a set of facts cannot.
In Figure 2, an exemplar sub-graph depicts the task of predicting a missing tail entity ⟨Patient
1, patientDrug, Nivolumab⟩. If the known facts about the head entity Patient 1, i.e., ⟨Patient 1,
hasStage, IIIA⟩, and ⟨Patient 1, hasSmokingHabit, CurrentSmoker⟩ are removed from the training
graph, the model’s predicted tail changes. Hence, the model relies on these necessary facts to
forecast Nivolumab, a plausible tail entity. In suficient scenario, for instance, adding the fact
⟨Patient 1, treatmentType, Immunotherapy⟩ and ⟨Patient 1, hasStage, IIIA⟩ to the training graph,
can lead the model to predict their drug as Nivolumab.
      </p>
      <p>
        Several studies demonstrate that generating high-quality negatives is a dificult but critical step
in improving KGE. As a result, negative sampling (NS) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] has become an essential component
of knowledge representation learning, considerably improving the performance of KGE models
through efective negative selection. The current inductive learning approaches, such as
symbolic learning [13, 14] and numerical learning [15, 16], fail to consider the validity and invalidity
of constraints when anticipating missing links. This results in the addition of connections
to KG graphs that do not meet domain requirements. Constraints can be validated using the
Shapes Constraint Language (SHACL)- W3C standardized shape constraint language. SHACL
constraints are symbolic constraints that provide explanations for the validity and integrity of
data in a KG. SHACL constraints serve as a set of rules or guidelines, defining the permissible
shapes that data instances can take within the graph. These constraints validate the content,
and relationships of entities, ensuring compliance with predefined standards or expectations.
In essence, SHACL constraints ofer a symbolic framework for evaluating the correctness and
coherence of data, contributing to the overall quality and reliability of the KG.
Our approach VISE tackles the challenge of KG completion by introducing a hybrid approach
that utilizes symbolic learning, symbolic constraints validation, and numerical learning to avail
the best of all the paradigms. VISE enhances the capabilities of KGE models by incorporating
symbolic learning inferences and constraints validation, thereby further transforming the input
KG by rewriting the relationships to specify negations in the KG. Thus, VISE helps numerical
learning, i.e., KGE models excel in predictive performance empowering KG completion.
Additionally, extracting two types of rationales necessary and suficient facts for LP tasks.
The rest of the paper is organized as follows: Section 2 motivates the KG completion problem
and defines the basic concepts of inductive learning. Section 3 presents the problem statement
SHACL Constraints for Medical Protocols
Nivolumab is NOT typically used to treat patients with EGFR positive gene mutations.
      </p>
      <p>Patient1 - Validates constraints</p>
      <p>Patient2 - Invalidates constraints
Male</p>
      <p>Old
hasGender hasAgeCategory hasSmokingHabit
and defines our approach, VISE, using a hybrid design pattern. Section 4 evaluates the approach
and benchmarks. Section 5 reports the results of the experimental study. Section 6 discusses the
state of the art. Finally, section 8 presents the conclusions and future work.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Motivation and Background</title>
      <p>This section uses an example to illustrate the problem of KG completion and the basic concepts
necessary to understand the approach presented in this paper.</p>
      <sec id="sec-3-1">
        <title>2.1. Motivating Example</title>
        <p>The motivation for our work arises from the fact that the KG completion methods do not
consider symbolic constraints to ensure KG integrity. The addition of missing relationships to
KGs for the completion of incomplete KGs without ensuring that the links added to satisfy the
domain constraints may correspond to unknown false facts. The state-of-the-art KG completion
approaches are deficient in considering the SHACL validation results in order to avoid
completing KGs with spurious relations. Figure 2 illustrates the lung cancer use case presented in
the current work. Domain experts (e.g., oncologists, medical doctors, or medical researchers)
specify clinical guidelines or protocols, for example, it is recommended that the drug Nivolumab
should be avoided for lung cancer patients mutated with EGFR Positive biomarker. These
recommendations are defined in terms of SHACL constraints to determine whether or not patients are
adhering to the clinical guidelines. The outcomes of performing SHACL constraints show if a
lung cancer patient validates or invalidates the constraints, i.e., if a patient mutated with EGFR
Positive is treated with Nivolumab drug. Therefore, SHACL validation reports verify the data
utilized by the KG completion procedures to complete the incomplete KGs, ensuring integrity.
Figure 2 shows the lung cancer KG that utilizes a set of variables to describe the main
characteristics of a lung cancer patient. These include the patient identifier (also known as the
electronic health record of a patient), gender, age, cancer stage (also known as the cancer
stage), smoking habits (also known as the smoking habit), lung cancer biomarkers, drugs and
treatments given to the patients. The OWA principle is used for representing KG in real-world
scenarios. KG completion approaches such as symbolic learning and numerical learning are
used to complete the missing relationships between entities of the KG. Symbolic learning, allows
the capture of explicit patterns from the KGs and the generation of Horn rules to derive insights
from the KGs. For example, as shown in Figure 2, a Horn Rule: lc:hasStage(X, IIIA),
lc:treatmentType(X, Immunotherapy):- lc:patient(X, Nivolumab) states that if a
patient has stage IIIA and receives immunotherapy, then it is most likely that the patient is being
treated with the drug Nivolumab. Numerical Learning, i.e., KGE models predicted missing links
by describing entities and their relations in a low-dimensional vector space. For example, by
taking into account the patient’s neighborhood, predicting the drug that a patient can receive.
Figure 2 showing the Patient 1 in green, validating the constraints since the patient is not
EGFR-mutated and hence can take Nivolumab following the clinical guidelines. Patient 2, shown
in red invalidates the constraints, i.e., does not adhere to the clinical guidelines. KG completion
approaches, such as symbolic and numerical learning, still predict that the patient should be
given Nivolumab, as they fall short in assessing whether the predicted missing links validate or
invalidate the clinical guidelines given by the domain experts. As shown, numerical learning
approaches such as TransH and RotatE predict patients (e.g., Patient 2) not adhering to clinical
guidelines with higher rank and score.</p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Preliminaries</title>
        <p>
          This section introduces basic preliminaries to understand our approach, i.e., shape, constraints,
shape evaluation, SHACL, knowledge graph embedding, Horn rule, heuristic-based negative
edges, support, confidence, PCA confidence, Hits@K, MRR, necessary and suficient explanation.
More details about these preliminary concepts in [
          <xref ref-type="bibr" rid="ref1">1, 13, 17, 18</xref>
          ].
        </p>
        <p>Knowledge Graphs. A knowledge graph (KG) is a directed edge-labeled graph  = (, , ),
where Con is a set of countable infinite constants.  ⊆ Con is a set of nodes,  ⊆ Con is a set of
edge labels, and  ⊆  ×  ×  is a set of edges.</p>
        <p>Constraints. A constraint corresponds to a rule that imposes restrictions on the values taken
for target nodes in  with a given edge .</p>
        <p>Shapes. A shape corresponds to a conjunction of constraints that a set of nodes in a knowledge
graph must satisfy. A shape  is inductively defined as follows:
::=T represents the value True;
| ∆  nodes belongs to the set of nodes  ;
|   a node satisfies the Boolean condition cond;
| 1 ∧ 2 is conjunction of shape 1 and shape 2;
| ¬ represents the negation of shape ;
| → {, } is cardinality on outward edges with label  to nodes satisfying ;
 and  are natural numbers.</p>
        <p>Shape Schema. A shapes schema is defined as a tuple ∑︀ = (, ,  ), where:
•  is a set of shapes;
•  is a set of shape labels;
•  : S →  is a total function from labels to shapes.</p>
        <p>Shape Target. Given a shapes schema Σ = ( , ,  ) and a directed edge-labelled graph
KG = (, , ),  (,  ) corresponds to the subset of nodes in  which are targets of  ∈  .
Shape Schema Evaluation. Given a shapes schema Σ = ( , ,  ) and a directed edge-labelled
graph KG = (, , ), a node  ∈  . Given a shape  ∈  , the shape evaluation function
[], ∈ {0, 1} states the results of evaluating  in a node  from  in KG.
• [ ]KG, = 1
• [∆  ]KG, = 1 if  ∈ 
• [ ]KG, = 1
• [1 ∧ 2], = min{[1],, [2],}
• [¬], = 1 − [],
• [→ {, }], = 1 if min ≤ |{ (, , ) ∈  | [], = 1}| ≤ max
Shape Schema Validation. Given a shapes schema Σ = ( , ,  ) and a directed edge-labelled
graph KG = (, , ), A node  ∈  validates  , i.e.,  |=  , if [], = 1 for all  ∈  and
 in  (,  ). KG satisfies the shape schema Σ = ( , ,  ), if for all  in  ,  |=  .
Example 2.1. A shape schema ∑︀ of lung cancer patients is given as follows:
• ∑︀ = (, ,  ),
•  = {→:ℎ {1, 1}, ∧, {→:ℎ {1, 1} →: 
{1,* }},
•  = { :  ,  :  },
•  ( :  ) = {→ℎ {1, 1}} ∧ {→:ℎ {1, 1},
•  ( :  ) = →:  {1,* }}.</p>
        <p>The evaluation of the shape schema ∑︀ = (, ,  ) of lung cancer patients represented in KG,
validates the nodes of patients with only one gender and age. Additionally, each lung cancer patient
should receive at least one treatment.</p>
        <p>Shapes Constraint Language (SHACL). SHACL [19] is the World Wide Web Consortium
(W3C) recommendation language for the declarative specification of integrity constraints over
RDF KGs. A SHACL shape represents a set of constraints that apply over the same entities; it
can refer to another shape, to represent constraints between entities of two types.
Knowledge Graph Embedding (KGE). Given a directed edge-labeled graph, KG = (V, E, L) and
set of vectors Γ . A KGE of KG is a pair of mappings ( ,  ) such that
•  :  → Γ , i.e.,  () maps a entity  in  to a vector in Γ , and
•  : L → Γ , i.e.,  () maps a directed edge  to a vector in Γ .</p>
        <p>A score function :  × ×  → R is used to measure the plausibility of candidate triples
represented in low-dimensional vector space, triples t = ⟨, , ⟩ with the higher score
 ( (),  (),  ()) values conveys better plausibility. The objective of KGE is to learn the
embeddings in ( , ) that maximize the plausibility of positive edges in + and minimize the
plausibility of negative edges in − . The set of positive edges, +, corresponds to the edges in
 . The set of negative edges, − , corresponds to the edges in  × L×  ∈/ +.
Hits@K. Given a tail prediction (, ?) over the directed edge-labeled graph  = (, , ),
the model predicts a list of  entities that might be related to the tail entity. @ determines
the fraction of plausible entities that appear in the top  predictions.</p>
        <p>|NumberOfPlausibleEntities ≤ |
@ = (1)

Mean Reciprocal Rank (MRR). Given the prediction problem either from a head or tail
perspective over the directed edge-labeled graph  = (, , ), the model ranks the plausible
entities and calculates the reciprocal rank of the plausible entity for each predictive task.

  = 1 ∑︁</p>
        <p>1
=1 
(2)
Horn Rule. A Horn rule is a logical implication defined as follows: Body ⇒ Head. The body of
the rule is comprised of predicate facts. The head is a predicate fact of a single atom. All the
variables in the Head are terms of at least one predicate fact in the Body. Every two predicate
facts in Body share at least one variable. We say a rule  : 1 ∧ 2 ∧ · · · ∧  =⇒ (, )
where Head represents (, ) and Body is 1 ∧ 2 ∧ 3 ∧ ... ∧ .</p>
        <p>Entailment of a Mined Rule [20]. Given a directed edge-labeled graph KG= (V,E,L) and a
mined rule  :  ⇒ , the entailed facts of R corresponds to the instantiations of
the predicate fact in Head on substituting the variables in Body, i.e., positive instantiations of
the conjunction of predicates in Body. That implies ∀  2 such that, Body[Z:=V2] is a positive
predicate fact, and Head[Z:=V2] corresponds to an entailed fact of . We can defined a predicate
fact as positive entailed fact +(), if [ :=  2] = (, ′) ∈ +.
Support of a Horn Rule. Given a directed edge-labeled graph KG= (V,E,L) and a mined rule
 :  ⇒ , the support of  indicates the number of positive entailed facts of Head.
Confidence of a Horn Rule. Given a directed edge-labeled graph KG= (V,E,L) and a mined rule
 :  ⇒ , the confidence of  is defined as a proportion of the positive predicate
facts of Head that are positive entailed facts based on .</p>
        <p>Heuristic-based Negative Edges (hE− ). Given a directed edge-labeled graph KG= (V,E,L) and
a mined rule  :  ⇒ , where Head is (, ). A heuristic-based negative edges ℎ−
corresponds to the set of instantiations (, ′) that do not belong to , but
• exists (, ) ∈ +(),
• (, ′) is entailed by Body.</p>
        <p>() =
()
|+ ∪ ℎ− |
A Necessary and Suficient Explanation. Given a directed edge-labeled graph KG= (V,E,L)
and a mined rule  :  ⇒ , where Head is (, ). Given a score function  :  × × 
→ R is used to measure the plausibility of predicted triples in KG= (V,E,L).</p>
        <p>• A necessary explanation corresponds to a set of predicate facts (, ) ∈ +, if removed
from  leads to a decrease in score function  .
• A suficient explanation corresponds to a set of predicate facts (, ) ∈/ +, if added to
, leads to an increase in score function  .</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Our Approach</title>
      <p>This section states the problem addressed in this paper and introduces the VISE framework,
which integrates symbolic learning, constraint validation, and numerical learning to create
more explainable, and reliable systems. The objective is to create a framework that is designed
to consider the semantics of symbolic systems.</p>
      <p>ℎ− () = {(, ′)|(, ′) ∈/ +() ∧ (, ) ∈ +() ∧ p(s,o’) is entailed by Body}
The set of ℎ− comprises triples to be predicted following this heuristic.</p>
      <p>
        PCA Confidence score of a Horn Rule. Given a directed edge-labeled graph KG= (V,E,L) and
a mined rule  :  ⇒ , where Head is (, ). The Partial Completeness Assumption
(PCA) score of  corresponds to the ratio of () to the cardinality of the union of +
and ℎ− . PCA confidence score quantifies the number of triples of the form (, ′) from −
that can be deduced following the heuristics edges. PCA (R) score ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ], where the score
indicates the amount of triples can be inferred.
(3)
(4)
3.1. Problem Statement
Consider edge-labeled graph KG = (V, E, L), such that each node  ∈  represents an entity, and
each  ∈  represents a unique relation between the entities. Let ∑︀ = (, ,  ) be a shape
schema over , and  (, , ′) be a scoring function quantifying the plausibility of a triple
(, , ′). The problem of link prediction over KG, i.e., a tail prediction ⟨, , ?⟩, such that  is the
subject entity in  and predicate  in  corresponds to the optimization problem of identifying
an entity ′ that produces the most plausible candidates for the incomplete triple (, , ′) and
 and ′ validate ∑︀ = (, ,  ).
      </p>
      <p>′ = arg min  (, , ) ∧  |=  ∧  |=</p>
      <p>∈
The aim is to find the most plausible entities  by inferring heuristic-based negative edges ℎ−
based on the positive edges + in KG and validates the shape schema ∑︀, i.e., (, , ′) |=  .</p>
      <p>Input:
KG
Input:
KG
Input:
KG
Input:
KG</p>
      <p>Mining
Model
Generate:
Mining Rules</p>
      <p>Generate:
Training Dataset</p>
      <p>Symbolic Learning</p>
      <p>Symbolic</p>
      <p>Model
Infer: Partial</p>
      <p>Completeness
Assumption Heuristic</p>
      <p>Symbol:
Valid and Invalid</p>
      <p>Predicted Links</p>
      <p>Symbol:
Transformed KG</p>
      <p>Data:
Embeddings</p>
      <p>Symbol:
Predicted Links</p>
      <sec id="sec-4-1">
        <title>3.2. The VISE Framework</title>
        <p>VISE encompasses a hybrid approach that showcases the impact of considering the semantics of
the symbolic system over the numerical learning approaches. VISE follows the hybrid design
pattern as illustrated in Figure 3, strategically combining numerical learning with symbolic
learning and constraints validation methods.</p>
        <p>Symbolic learning is applied to the input KG, resulting in the generation of logical rules and PCA
heuristic-based edges. The learned heuristic-based edges serve as prior knowledge, improving
numerical learning approaches such as KGE models combined with constraints validation and
KG transformation. During the process of symbolic learning, VISE utilizes extracted horn rules
in conjunction with PCA Confidence in order to infer heuristic-based negative edges. The mined
rules are subsequently employed to generate predictions regarding the missing relationships in
the input KG. These predictions are based on logical inference, which is used to calculate the
entailment of the mined rules. SPARQL queries are employed to infer the entailment of mined
rules and construct heuristic-based negative edges (hE− ).</p>
        <p>The predictions generated by the symbolic learning system in conjunction with the input
KG are then fed to the Constraints Validation and KG Transformation component, where the
predicted links are evaluated to determine whether they validate or invalidate the SHACL
constraints. Furthermore, the generated validation report is utilized to transform or rewrite the
SHACL Constraints for Medical Protocols
Nivolumab is NOT typically used to treat patients with EGFR positive gene mutations.
Symbolic Learning (Body:- Head)
lc:hasStage(X, IIIA), lc:treatmentType(X,Immunotherapy)
:lc:patientDrug(X, Nivolumab)</p>
        <p>Male
(c) KG enrichment with predictions generated using (d) KG enrichment with predictions of symbolic
symbolic learning rules rules and SHACL validation
KG to contain the information resulting from the constraints validation. The transformed KG is
then provided as input to the numerical learning models, i.e., KGE models, during the training
phase. This is achieved by processing the data into a low-dimensional space. The process of
numerical learning is capable of predicting missing links, thereby completing the KGs with
missing links that validate the constraints at a higher rank and with a greater probability of
accuracy. Tranformation or rewriting of KGs before giving as input to the numerical learning
component transforms the KGs to contain negated facts, allowing the KGE model to learn in
all the four quadrants as shown in Figure 1, enhancing the performance of the models and
empowering KG completion. Several studies demonstrated the need for negated facts in KGs
to boost the performance of KGE models. VISE employs a two-fold rewriting process. First, it
evaluates the links predicted by symbolic learning using constraints. Second, depending upon
the validation report of the predicted link. If the patient in the lung cancer KG invalidates the
constraint, the links that resemble the patient characteristics in the KG are added with negation.</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.3. Running Example</title>
        <p>As discussed in subsection 2.1, a SHACL constraint stipulates that a patient who has undergone
a mutation involving the EGFR gene should not be treated with the drug Nivolumab. Symbolic
learning, on the other hand, has identified a rule that states that if a patient is in the advanced
stage of lung cancer (stage IV) and has undergone immunotherapy treatment, there is a higher
probability that the patient will receive the drug Nivolumab. PCA heuristics enabled symbolic
learning to predict that the patient should receive Nivolumab.</p>
        <p>The SHACL validation revealed that this patient violated the constraint if the predicted link
was added to the KG. In the transformation process, VISE is capable of adding a fact to the
KG indicating that the patient should not receive Nivolumab. As a result, the transformed KG
explicitly represents positive and negative facts. This new way of modeling statements further
enhances the performance of the models by assigning high values of plausibility score ′ to
triples that are likely to be true and low scores to triples that are likely to be false.
To illustrate, a running example demonstrates the process of transforming the KG to explicitly
add negated facts to the KG, thereby further enhancing numerical learning and performance.
Figure 4 shows the SHACL constraints based on the clinical guidelines, horn rules mined over
the original KG, and a sub-figure that resembles the original lung cancer KG. Figure 5a shows
the Baseline 1 approach involves inputting the original KG into numerical learning, which is
then processed by KGE models without the addition of inferred facts from symbolic learning
for KG enrichment. Figure 5b illustrates the transformation of KG predicates implemented over
the original KG (Baseline 2), thereby demonstrating the necessity to rewrite the negated facts
to the KG for enhancement of numerical learning. For instance, the predicate may be altered
to either EGFR_Positive if a patient is positively mutated for EGFR or EGFR_Negative if a
patient is negatively mutated for EGFR mutation. The results of the transformation of KG are
presented in Section 5 and demonstrate the impact on the performance of KGE models.
Figure 5c shows the Baseline 3 approach that involves enriching the KG with symbolic learning.
The KG is enriched with symbolic learning rules that incorporate PCA heuristics to generate
heuristics-based negative (hE− ) edges. As discussed before, the rule can predict that if a stage
IV lung cancer patient receives immunotherapy treatment then that patient is more likely to
receive drug Nivolumab which is added as a fact as shown in the Figure 5c with patientDrug
Nivolumab in the KG given as input to KGE models. The Baseline 3 approach is in alignment
with the state-of-the-art methodology of SPaRKLE [20]. Furthermore, Figure 5d displays Baseline
4, which highlights the influence of constraint validation. To show the need for constraint
validation, inferred facts resulting from symbolic learning were added to the KG, which is only
fed into KGE models once a patient validates the constraint. As shown in Figure 5d, the patient
in red invalidates the constraint. As a result, the inferred fact patientDrug Nivolumab is not
included for that patient. Figure 6 illustrates the transformation of KG implemented in VISE. In
VISE, we aim to emphasize the significance of the hybrid approach in considering the impact of
symbolic systems. The transformation of KG is achieved by explicitly incorporating the inferred
facts predicted by symbolic rules, thereby validating the constraints for the predicted facts.
To illustrate, in the figure referenced in the text, the patient in green validates the constraint.
Consequently, the rewriting in the transformed KG includes biomarkerEGFR_Negative,
EGFR_Negative, patientDrug_Nivolumab, and Nivolumab. For the patient in red,
invalidates the constraint. Consequently, the following transformation is performed as follows
EGFR_Positive, EGFR_Positive, patientDrug_NoNivolumab, and NoNivolumab. The
results of the transformation of VISE KG are presented in 5, which demonstrates the impact
on the performance of KGE models. Figure 5a, Figure 5b, Figure 5c, and Figure 5d present the
benchmarks (resp., Baseline 1, Baseline 2, Baseline 3, and Baseline 4) utilized in the experiments.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Evaluation</title>
      <p>The aforementioned section outlines our proposed framework VISE and its components. In this
section, we report the experiment settings, benchmark description, observed results, involved
baselines, and models. We empirically assess the efectiveness of VISE in the LP problem over
the Lung Cancer KG. VISE provides comprehensive explanations: necessary and suficient for
LP task. For instance, an LP task can be "Whether a lung cancer patient is in Relapse?". Thus,
given a head entity and relation predicting the tail entity, i.e., ⟨Patient X, hasRelapse, ?⟩. The
empirical evaluation aims to answer the following research questions: RQ1) What is the impact
of negated facts on the KGE model’s performance and its explainability? RQ2) How do symbolic
rules and constraints enhance the explanations of the KGE model’s behavior?
Benchmark. We evaluate VISE approach on three anonymized Lung Cancer KGs: 1, 2,
and 3. Table 1 shows the statistics of all the benchmarks. The Lung Cancer KG comprises
medical records about a lung cancer patient from heterogeneous data sources. Each medical
record describes the characteristics of a patient sufering from lung cancer. The medical
characteristics include a cancer stage (e.g., Stage IVB), age, gender, smoking habit (e.g., Current Smoker),
type of mutation (e.g., EGFR Negative), recommended drug for treatment (e.g., Vinorelbine), the
occurrence of relapse (e.g., Relapse or Progression or No Relapse), and types of treatment (e.g.,
Immnunotherapy) for curing the cancer. The prediction problem is a link prediction to predict
the Relapse of a lung cancer patient, which can be Relaspe or No Relaspe. We utilize SHACL
constraints as medical protocols that recommend when a drug should be prescribed according
to a patient’s mutations; we defined one shape schema with four diferent SHACL constraints,
for instance, a constraint stating that "If a patient mutated with EGFR negative should not take
Afatinib, and if a patient mutated with EGFR positive should not take Nivolumab".
Baselines. We evaluate and compare four baselines for our VISE approach. Baseline 1 includes
the evaluation of the state-of-the-art KGE models for the KG completion. Baseline 2 reveals
the evaluation of transformed KG with KGE models. Baseline 3 utilizes the hybrid approach,
SPaRKLE [20], which employs symbolic learning techniques to enhance the performance
of KGE models. Baseline 4 combines SPaRKLE with the results of patients who satisfy the
medical protocols. VISE approach integrates the fusion of SPaRKLE with the transformed
KG including validation and violation results to enhance the performance of KGE models in
LP tasks. The current implementation utilizes various state-of-the-art KGE models from the
PyKEEN [21] pipeline, which includes TransE [16], TransD [22], TransH [23], and RotatE [24].
We conducted an ablation study to tune hyperparameters for KGE models based on benchmark
KGs. Translation-distance space models, including TransE, TransD, and TransD, translate the
head entity’s geometric embedding space with a given relation closer to the tail entity. RotatE, a
popular model for learning embeddings in Euclidean space, has attracted attention for learning
symmetric, asymmetric, 1-1, 1-N, N-1, and M-to-N relationships. Table 2, Table 3 and Table 4
demonstrates the comparison between baselines and VISE approach for KG completion.
Implementation. VISE is implemented in a virtual machine on Google Colab with 40 GiB
VRAM and 1 GPU NVIDIA A100-SMX4, with CUDA version 12.2 (Driver 535.104.05) using
Python 3.9. The source code of VISE approach, the benchmark KGs, and the trained KGE
models are publicly available in our GitHub repository 1. Figure 3 depicts the hybrid design
pattern, integrating inductive learning with symbolic learning techniques. Symbolic learning
includes logical horn rules () and SHACL constraints (). Symbolic learning is performed
over the input KG, resulting in rules, heuristic-based edges, and SHACL validation. Thus, the
inferred heuristic edges with validation results are utilized as implicit knowledge to enhance
inductive learning, i.e., KGE models. The predictions generated from the symbolic rules and
constraints materialized in the input KG and fed as input to inductive learning. The benchmark
KGs are divided into 80-20 train-test splits. The model’s eficacy in the LP problem is evaluated
using Hits@K and MRR. Both metrics have values between 0 and 1, and higher conveys better.
VISE relies on [20] and [25] for symbolic learning methods. Furthermore, our approach is
model-agnostic and compatible with other symbolic and inductive learning approaches.
1https://github.com/SDM-TIB/VISE</p>
    </sec>
    <sec id="sec-6">
      <title>5. Results</title>
      <p>In this empirical study, we evaluate the eficacy of numerical inductive and symbolic learning
approaches in terms of the evaluation metrics proposed by Akrami et al. [26]. These empirical
studies aim to address the research questions RQ1 in Section 5.1 and RQ2 in Section 5.2.</p>
      <sec id="sec-6-1">
        <title>5.1. Impact of Negated Facts on KGE Model Behavior</title>
        <p>We report the efectiveness of VISE approach, focusing on KGE models- TransE, TransD, TransH,
and RotatE in the context of lung cancer relapse prediction problems. The comprehensive
analysis revealed a robust performance compared to baselines. KGE models are trained over the
diferent benchmark KGs, i.e., positive edges +, to predict missing links. The evaluation report
presented in Table 2, 3, and 4 are obtained using the optimized hyperparameters provided by
the PyKEEN pipeline. The impact of negated facts is assessed with Hits@1, Hits@3, Hits@5,
Hits@10, and MRR in KG completion. TransE, a basic translation model, emerged as performing
worst in all baselines with benchmarks respectively. Nevertheless, highlighting the limitations
of TransE in modeling 1-N relationships leads to poor performance, particularly in predicting
the correct tail at the topmost position. TransH model results support the claim in [23], that it
outperforms TransE and TransD models. In 1 and 2, TransH performance contributes to
promising results in capturing complex geometric relationships with score values ranging from
0.413 to 0.865. TransD, which uses relation-specific projections to translate the embedding
space, yields slightly lower values than TransH and TransE. However, RotatE indicates the
best performance in all the testbeds except in 1. In 2 and 3, the values of Hits@1
range from 0.489 to 0.887. We can observe that the evaluation of benchmark KGs in diferent
experimental testbeds, VISE outperforms compared to the other baseline approaches. The
experimental evaluation comprises 100 testbeds per KGs, amounting to a total of 300 testbeds.
In summary, the evaluation results underline the robust performance of TransH and RotatE for
KG completion in lung cancer relapse prediction tasks.</p>
        <p>However, the rationale behind the inner workings of KGE models may be dificult to understand.
The experimental results demonstrate the need for explanations and assistance to understand
KGE model behavior. VISE shows improved KGE model performance and provides two types
of post hoc explanation for the prediction problem. In VISE approach, KGE models showed
marginally better performance compared to Baseline 1. We categorize our explanations as
necessary and suficient . The heuristic-based negative edges (ℎ− ) generated by symbolic
learning demonstrate the importance of enhancing the performance of VISE. The addition of
ℎ− edges to KG has been deemed a suficient explanation, as evidenced by the improved
performance of the KGE model in terms of Hits@K and MRR. For example, Table 5 displays
examples of mined Horn rules that were chosen based on the SHACL constraints, i.e., clinical
guidelines used to infer the ℎ− edges. Moreover, the removal of these edges from the KG
resulted in a notable decline in performance, which can be attributed to the necessity of these
facts, i.e., necessary facts to explain the prediction performance thereby answering RQ1.
5.2. Efectiveness of Symbolic Rules and Constraints on LP task
The Horn rules mined by AMIE [13] over LC KG are used to help doctors screen for and identify
persons who are at high risk of acquiring lung cancer. Mined rules are examined in terms of
biomarkers, medications, and therapies, and ranked according to the PCA confidence score. The
efectiveness of VISE is evaluated in terms of the impact of validating constraints for the missing
link being predicted by the symbolic learning technique. As described in Section 3 the
heuristicbased negative edges (hE− ) are predicted using the Partial Completness Assumption (PCA)
heuristics from the input KG. The PCA Confidence of a Horn rule, which indicates the amount
of incompleteness in a knowledge graph (KG), is employed to infer new links and predictions.
These predictions are validated by applying the SHACL constraints to determine the validity
of the inferred links. The results demonstrated in Table 5 indicate the amount of valid and
invalid predictions produced by the symbolic learning techniques. Table 5 shows examples
of the symbolic rules, for example, stage(?a, IV), treatment(?a, Immunotherapy) ⇒
drug(?a, Nivolumab) stating that if a stage IV lung cancer patient received Immunotherapy
treatment then it is more likely that the patient receives Nivolumab is with the PCA Confidence
score of 0.833. As mentioned before, the heuristics-based negative edges (hE− ) or predictions
are validated using SHACL constraints, and Table 5 shows the number of valid (#) and invalid
(#) links for each of the LC KGs used as a benchmark in VISE.</p>
        <p>Furthermore, the symbolic rules are used to represent the studies reported in the literature.
Table 6 provides examples of mined Horn rules and supporting literature. Consequently, the
impact of symbolic rules and constraints utilized to explain the KGE models is demonstrated,
thereby enabling an answer to be provided to the research question RQ2.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Related Work</title>
      <p>
        The integration of symbolic and numerical learning into KGs enhances their utility and
interpretability. Symbolic techniques employ rules and logic to identify missing relationships,
whereas numerical approaches utilize low-dimensional vector spaces to discern connections
between entities in large KGs. Symbolic constraint validation, in isolation, identifies
inaccuracies in KGs that can be employed to assess data quality. It is of vital importance to explain
the predictions in the healthcare domain, as this helps domain experts in decision-making,
identifying the interactions between drugs and their side efects, and patient diagnosis.
The majority of KGE methods employ triples from KGs as input, with the embeddings being
trained using vector space assumptions (e.g., translational, neural network, complex space) [15].
Furthermore, the embeddings are obtained to perform the link prediction task [21], as outlined
in Rivas et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Rivas et al. propose a neuro-symbolic perception for drug treatment response
to enhance the link prediction capabilities of KGE models by deducing implicit knowledge using
datalog rules. Akrami et al. [26] present a study that employs a realistic and updated assessment
of various KG completion techniques. The objective is to establish their usefulness in improving
KG completeness and quality. The findings of the study, as presented in work [ 26], indicate that
the embedding models may have been biased toward learning reverse relations for LP due to
the presence of data redundancy and Cartesian product relations.
      </p>
      <p>Furthermore, it was demonstrated that simple models, such as symbolic learning approaches,
outperform numerical models when data contains reverse relations or data redundancy.
Consequently, we aim to showcase the combination of symbolic and numerical methodologies that
frequently result in enhanced performance and outcomes in a variety of activities, including
KG completion in our proposed approach VISE. Moreover, the validation of constraints over
the predicted links from symbolic learning approaches can assist in identifying whether the
predicted links validate or invalidate the constraints. This process can enhance the system’s
performance by providing information about the constraint validation for LP tasks. While both
symbolic and numerical techniques have advantages and disadvantages, a hybrid approach
that combines them can mitigate shortcomings while leveraging the complementary benefits of
these KG completion methods to further empower KGs.</p>
      <p>In a related study, Lajus et al. [13] present a symbolic learning technique that captures the
co-occurrence of relationships, rules, and logical dependencies within KGs. Among numerous
KG completion approaches, this method employs the OWA to extract association rules from
drug(?a, Nivolumab) ⇐
treatment(?a, Immunotherapy),
treatment(?a, Intravenous_Chemotherapy)
biomarker(?a, EGFR_Negative) ⇐
relapseProgression(?a, Progression),
drug(?a, Pembrolizumab)
biomarker(?a, EGFR_Negative) ⇐
biomarker(?a, ALK_Negative),
treatment(?a, Radiotherapy_To_Bone)</p>
      <p>Statements</p>
      <p>Lung cancer patients in stage IV[27]
and receive Immunotherapy treatment
are more likely to revive Nivolumab[28].</p>
      <p>Non-small cell lung cancer patients
receive Nivolumab[28, 29] as first-line</p>
      <p>Chemotherapy and Immunotherapy
treatments for progression-free survival.</p>
      <p>EGFR[30] negative lung cancer
patients are more likely to experience
progression and are treated with</p>
      <p>pembrolizumab[31] drug.</p>
      <p>Lung cancer patients mutated with</p>
      <p>ALK negative and receives treatment
radiotherapy[32, 33] to bone are more likely
to be also mutated with EGFR[30] negative.</p>
      <p>KGs. AMIE [13] enhances the quality and completeness of KGs by deducing missing linkages
and connections, while also considering semantics. AnyBURL [14] (Anytime Bottom-Up Rule
Learning) is a state-of-the-art system that entails initiating specific instances within the KGs
and subsequently generalizing them to generate more encompassing logical rules that can be
applied across the KGs. Khajeh Nassiri et al. [34] propose a symbolic learning technique that
emphasizes the use of logical rules, including numerical predicates. This method enables KGs
to recognize and mine correlations between numerical values, measures, and other quantitative
properties, resulting in a more expansive and precise representation of real-world knowledge.
Chudasama et al. [35] demonstrated in one of the related studies that SHACL technologies may
be utilized to evaluate data over KGs for quality assessment, as well as in predictive modeling
analysis to improve model interpretability. SHACL technologies can be used to validate data and
then used with ensemble approaches, such as Random Forest and Decision Trees, to interpret the
behavior of Machine Learning (ML) models, which can help understand the outcomes generated
by prediction models. Rabbani et al. [36] employs a technique to extract validating shapes from
large KGs. Furthermore, an eficient SHACL validation engine [ 25] shows the best performance
in planning and executing SHACL shape schema to determine whether entities from KGs comply
with specific medical protocols. VISE is system-agnostic, allowing straightforward integration
with any existing KGE model or symbolic system. To achieve integration, the mining horn rules
for symbolic systems, as well as the computation of PCA and prediction scores, and a set of
SHACL constraints can be utilized for validation.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Discussion</title>
      <p>VISE framework demonstrates the efectiveness of considering PCA heuristics for LP tasks
and generates explanations. However, the proposed approach has limitations in terms of
incorporating the semantics of KGs. By employing symbolic reasoning, implicit facts can
be deduced, which can be utilized to enhance the neighborhood of an entity. Consequently,
considering the semantics of KGs would provide a comprehensive picture of the scalability of
each KGE model in real-world use cases. Furthermore, investigating the computational overhead
observed for each model to capture complex relationships can also be conducted in future studies.
The experimental results demonstrate that KGE models do not fully account for the contextual
knowledge of entities, such as entity validation in the context of medical protocols. Nevertheless,
our approach, VISE, is domain-agnostic and can be used to enhance and explain the behavior of
KGE models in LP tasks. The mining of rules, validation of entities, and training of the KGE
models scale with the size of KGs. Thus, exploiting the minimal neighborhood with specific
rules for negative sampling will aid in solving the scalability issues. Lastly, state-of-the-art
KGE models are commonly employed for a range of downstream tasks. However, their latent
vector representations lack self-interpretability. Consequently, future studies may benefit from
leveraging the enriched contextual information considering the semantics of KGs to enhance
the explanations generated by Large Language Models (LLMs). This could prove valuable in
critical domains such as healthcare, facilitating more eficient decision-making processes.</p>
    </sec>
    <sec id="sec-9">
      <title>8. Conclusions and Future Works</title>
      <p>VISE avails the advantages of the PCA heuristic, which improves predictions regarding missing
links. Constraint validation helps include additional symbolic system semantics in numerical
learning approaches. Empirical evidence indicates that the integration of symbolic learning
approaches with constraint validation can enhance the performance of KGE models, particularly
when prior knowledge is taken into account. Consequently, VISE exemplifies the advantages
of integrating symbolic and numerical methodologies into a hybrid or neuro-symbolic AI
system. This integration allows academics and practitioners to combine these two conceptual
AI approaches, thereby achieving accurate solutions for KG completion.</p>
      <p>Moreover, VISE demonstrated the necessity of rewriting the KGs to include negative edges
using SHACL constraints’ results rather than randomly generating negative samples for the
numerical learning approaches. It is important to note that hybrid approaches do come with
limitations. This work provides evidence that hybrid methods necessitate the integration of
various components, which can result in increased computational complexity. The processing
of symbolic systems may result in the mining of rules and the inference of triples that are
unnecessary for numerical models, which may not utilize them. This opens the door for future
research to eficiently execute hybrid systems and to fully leverage their benefits.</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by TrustKG- Transforming Data in Trustable Insights
with grant P99/2020 and the EraMed project P4-LUCAT (GA No. 53000015).
A review, CoRR abs/2402.19195 (2024). URL: https://doi.org/10.48550/arXiv.2402.19195.
doi:10.48550/ARXIV.2402.19195. arXiv:2402.19195.
[10] Y. Chudasama, Exploiting semantics for explaining link prediction over knowledge
graphs, in: C. Pesquita, H. Skaf-Molli, V. Efthymiou, S. Kirrane, A. Ngonga, D.
Collarana, R. Cerqueira, M. Alam, C. Trojahn, S. Hertling (Eds.), The Semantic Web: ESWC
2023 Satellite Events - Hersonissos, Crete, Greece, May 28 - June 1, 2023, Proceedings,
volume 13998 of Lecture Notes in Computer Science, Springer, 2023, pp. 321–330. URL: https:
//doi.org/10.1007/978-3-031-43458-7_50. doi:10.1007/978-3-031-43458-7\_50.
[11] H. Zhang, T. Zheng, J. Gao, C. Miao, L. Su, Y. Li, K. Ren, Data poisoning attack against
knowledge graph embedding, in: IJCAI, 2019.
[12] D. Purohit, M. Vidal, Mining symbolic rules to explain lung cancer treatments, in:
C. Pesquita, H. Skaf-Molli, V. Efthymiou, S. Kirrane, A. Ngonga, D. Collarana, R. Cerqueira,
M. Alam, C. Trojahn, S. Hertling (Eds.), The Semantic Web: ESWC 2023 Satellite Events
- Hersonissos, Crete, Greece, May 28 - June 1, 2023, Proceedings, volume 13998 of
Lecture Notes in Computer Science, Springer, 2023, pp. 69–74. URL: https://doi.org/10.1007/
978-3-031-43458-7_13. doi:10.1007/978-3-031-43458-7\_13.
[13] J. Lajus, L. Galárraga, F. Suchanek, Fast and Exact Rule Mining with AMIE 3, in: The</p>
      <p>Semantic Web, 2020.
[14] C. Meilicke, M. W. Chekol, D. Rufinelli, H. Stuckenschmidt, Anytime bottom-up rule
learning for knowledge graph completion, in: IJCAI-19, 2019. doi:10.24963/ijcai.
2019/435.
[15] M. Ali, M. Berrendorf, C. T. Hoyt, L. Vermue, M. Galkin, S. Sharifzadeh, A. Fischer, V. Tresp,
J. Lehmann, Bringing light into the dark: A large-scale evaluation of knowledge graph
embedding models under a unified framework, CoRR (2020).
[16] A. Bordes, N. Usunier, A. Garcia-Durán, J. Weston, O. Yakhnenko, Translating embeddings
for modeling multi-relational data, NIPS’13, Curran Associates Inc., Red Hook, NY, USA,
2013, p. 2787–2795.
[17] J. E. Labra-Gayo, H. García-González, D. Fernández-Alvarez, E. Prud’hommeaux,
Challenges in RDF Validation, Springer International Publishing, 2019, p. 121–151. URL:
http://dx.doi.org/10.1007/978-3-030-06149-4_6. doi:10.1007/978-3-030-06149-4_6.
[18] A. Rossi, D. Firmani, P. Merialdo, T. Teofili, Explaining link prediction systems based
on knowledge graph embeddings, in: Proceedings of the 2022 International Conference
on Management of Data, SIGMOD/PODS ’22, ACM, 2022. URL: http://dx.doi.org/10.1145/
3514221.3517887. doi:10.1145/3514221.3517887.
[19] H. Knublauch, D. Kontokostas, Shapes Constraint Language (SHACL), W3C
Recommendation, 2017. URL: https://www.w3.org/TR/2017/REC-shacl-20170720/.
[20] D. Purohit, Y. Chudasama, A. Rivas, M.-E. Vidal, Sparkle: Symbolic capturing of knowledge
for knowledge graph enrichment with learning, in: Proceedings of the 12th Knowledge
Capture Conference 2023, K-CAP ’23, Association for Computing Machinery, New York,
NY, USA, 2023, p. 44–52. URL: https://doi.org/10.1145/3587259.3627547. doi:10.1145/
3587259.3627547.
[21] M. Ali, M. Berrendorf, C. T. Hoyt, L. Vermue, S. Sharifzadeh, V. Tresp, J. Lehmann, PyKEEN
1.0: A Python Library for Training and Evaluating Knowledge Graph Embeddings, Journal
of Machine Learning Research (2021). URL: http://jmlr.org/papers/v22/20-825.html.
I. Desideri, M. Loi, A. Reginelli, P. Tassone, P. Correale, The role of brain radiotherapy
for egfr- and alk-positive non-small-cell lung cancer with brain metastases: a review, La
Radiologia medica 128 (2023). doi:10.1007/s11547-023-01602-z.
[33] A. Wrona, R. Dziadziuszko, J. Jassem, Combining radiotherapy with targeted therapies
in non-small cell lung cancer: focus on anti-egfr, anti-alk and anti-angiogenic agents,
Translational Lung Cancer Research 10 (2021). URL: https://tlcr.amegroups.org/article/
view/49653.
[34] A. K. Nassiri, N. Pernelle, F. Saïs, REGNUM: generating logical rules with numerical
predicates in knowledge graphs, in: The Semantic Web - 20th International Conference,
ESWC 2023, Hersonissos, Crete, Greece, May 28 - June 1, 2023, Proceedings, volume 13870,
Springer, 2023, pp. 139–155. doi:10.1007/978-3-031-33455-9\_9.
[35] Y. Chudasama, D. Purohit, P. D. Rohde, M.-E. Vidal, Enhancing interpretability of
machine learning models over knowledge graphs, in: N. Keshan, S. Neumaier, A. L.
Gentile, S. Vahdati (Eds.), Proceedings of the Posters and Demo Track of the 19th
International Conference on Semantic Systems co-located with 19th International
Conference on Semantic Systems (SEMANTiCS 2023), Leipzig, Germany, September 20
to 22, 2023, volume 3526 of CEUR Workshop Proceedings, CEUR-WS.org, 2023. URL:
https://ceur-ws.org/Vol-3526/paper-05.pdf.
[36] K. Rabbani, M. Lissandrini, K. Hose, Extraction of validating shapes from very large
knowledge graphs, Proc. VLDB Endow. 16 (2023) 1023–1032. URL: https://www.vldb.org/
pvldb/vol16/p1023-rabbani.pdf. doi:10.14778/3579075.3579078.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          , E. Blomqvist,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cochez</surname>
          </string-name>
          , C. d'Amato, G. de Melo,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gutierrez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kirrane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E. L.</given-names>
            <surname>Gayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Neumaier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Ngomo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Rashid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmelzeisen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sequeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zimmermann</surname>
          </string-name>
          , Knowledge Graphs,
          <source>Synthesis Lectures on Data, Semantics, and Knowledge</source>
          , Morgan &amp; Claypool Publishers,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gutierrez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Sequeda</surname>
          </string-name>
          , Knowledge graphs,
          <source>Commun. ACM</source>
          <volume>64</volume>
          (
          <year>2021</year>
          )
          <fpage>96</fpage>
          -
          <lpage>104</lpage>
          . doi:
          <volume>10</volume>
          . 1145/3418294.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Aisopos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jozashoori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Niazmand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Purohit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rivas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sakor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Iglesias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vogiatzis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Menasalvas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Vigueras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gómez-Bravo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Torrente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. H.</given-names>
            <surname>López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Pulla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dalianis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Triantafillou</surname>
          </string-name>
          , G. Paliouras,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vidal</surname>
          </string-name>
          ,
          <article-title>Knowledge graphs for enhancing transparency in health data ecosystems</article-title>
          ,
          <source>Semantic Web</source>
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>943</fpage>
          -
          <lpage>976</lpage>
          . URL: https://doi.org/10.3233/SW-223294. doi:
          <volume>10</volume>
          .3233/SW-223294.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Loyer</surname>
          </string-name>
          , U. Straccia,
          <article-title>Any-world assumptions in logic programming</article-title>
          ,
          <source>Theoretical Computer Science</source>
          <volume>342</volume>
          (
          <year>2005</year>
          )
          <fpage>351</fpage>
          -
          <lpage>381</lpage>
          . URL: https://www.sciencedirect.com/science/article/pii/ S030439750500304X. doi:https://doi.org/10.1016/j.tcs.
          <year>2005</year>
          .
          <volume>04</volume>
          .005.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Akrami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Re-evaluating embedding-based knowledge graph completion methods</article-title>
          ,
          <source>in: CIKM</source>
          ,
          <year>2018</year>
          . doi:
          <volume>10</volume>
          .1145/3269206.3269266.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rivas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Collarana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Torrente</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-E. Vidal</surname>
          </string-name>
          ,
          <article-title>A neuro-symbolic system over knowledge graphs for link prediction</article-title>
          ,
          <source>Semantic Web Journal. Special Issue on Neuro-Symbolic Artificial Intelligence and the Semantic Web</source>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          . doi:
          <volume>10</volume>
          .3233/SW-233324.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rossi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Firmani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Merialdo</surname>
          </string-name>
          , T. Teofili,
          <article-title>Explaining link prediction systems based on knowledge graph embeddings</article-title>
          ,
          <source>in: SIGMOD</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Knowledge graph embedding: A survey of approaches and applications</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>29</volume>
          (
          <year>2017</year>
          )
          <fpage>2724</fpage>
          -
          <lpage>2743</lpage>
          . URL: http://dx.doi.org/10.1109/TKDE.
          <year>2017</year>
          .
          <volume>2754499</volume>
          . doi:
          <volume>10</volume>
          .1109/tkde.
          <year>2017</year>
          .
          <volume>2754499</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Madushanka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ichise</surname>
          </string-name>
          ,
          <article-title>Negative sampling in knowledge graph representation learning:</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>