<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Refining SemOpenAlex Concept Ontology: A Constraint-Aware Approach via Knowledge Graph Embeddings and SKOS Constraints</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Özge Erten</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shervin Mehryar</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bo Xiong</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Remzi Çelebi</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christopher Brewster</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Science Group, TNO</institution>
          ,
          <addr-line>Kampweg, Soesterberg</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Artificial Intelligence, University of Stuttgart</institution>
          ,
          <addr-line>70569, Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Data Science, Maastricht University</institution>
          ,
          <addr-line>Paul-Henri Spaaklaan 1, 6229 GT, Maastricht</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The continuous growth in scientific publications has led to an increasing demand for eficient solutions in managing vast amounts of scholarly information. SemOpenAlex, an academic article Knowledge Graph (KG), aims to organize the scientific papers by tagging them with the concepts representing the relevant topics. The concepts are hierarchically organized using the Simple Knowledge Organization System (SKOS) vocabulary. However, this concept hierarchy contains noise resulting from Natural Language Processing (NLP) concept extraction. This paper proposes a link prediction based method to reduce the noise within SemOpenAlex concept hierarchy. The method utilizes informal SKOS consistency definitions to create negative triples which violate the definitions, combined with randomly generated negatives. The primary objective here is to integrate true negative samples into the knowledge graph embedding model during the learning process. This study contributes to refining SKOS-based KGs by enhancing the semantic quality of information within the KG.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Knowledge graph</kwd>
        <kwd>Link prediction</kwd>
        <kwd>SKOS vocabulary</kwd>
        <kwd>negative sampling</kwd>
        <kwd>custom negative sampling</kwd>
        <kwd>constraintaware embeddings</kwd>
        <kwd>constraint-aware negative sampling</kwd>
        <kwd>Knowledge graph embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        According to the online citation index Web of Science, over 6 million articles were published between
2018 and 2022 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The increasing number of publications, as shown in Figure 1, contribute to a rich
pool of scholarly knowledge. Since scholarly knowledge is rapidly increasing and evolving; accessing
to it is often insuficient [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For instance, from researchers’ perspective, the volume of publications
turns an eficient literature search and discovery of relevant articles and topics into a labor-intensive
task. One popular approach to managing this knowledge involves the use of Knowledge Graphs (KGs).
Beside representational capabilities, KGs ofer the support of dividing topics and their sub-fields [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Recently, SemOpenAlex KG [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] was developed to organize and represent scientific publications and
related information across diverse domains. This KG is structured around entities such as titles,
authors, abstracts and article texts. It also includes a concept hierarchy that categorizes publications
under diferent topics. This concept hierarchy was built using the Simple Knowledge Organization
System (SKOS) standard – a common vocabulary representing controlled vocabularies or knowledge
organization systems in a machine-readable way. Some popular knowledge bases facilitate it to
organize their information. For example, AGROVOC[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], an agricultural thesaurus, uses SKOS to
represent its semantic relations. GACS Core [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], an agricultural conceptual scheme developed to
enhance the consistency of agricultural research, also uses SKOS.
2 million
article
      </p>
      <p>
        However, the extraction and organization of SemOpenAlex concepts involve Natural Language
Processing (NLP) techniques and heuristics, often introduce noise into the extracted knowledge. This
leads to a loss of quality and reliability in the KG. Manual eforts in cleaning noise from KGs are
time-consuming and laborious. In this research, we aim to improve the quality of SemOpenAlex
concepts with an automated method [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Specifically, this work focuses on enhancing the data quality
of SemOpenAlex by improving the quality of SKOS connections in the concept hierarchy. To achieve
this, our approach first removes inconsistent relations in SemOpenAlex KG and subsequently predicts
more accurate ones. We apply a customised negative sampling method within Knowledge Graph
Embedding (KGE) techniques for link prediction. Namely, traditional KGEs commonly assume local
completeness in the KG, employing a ’Close World’ perspective when generating negative triples.
This approach can lead to the generation of negative triples that are actually correct, as emphasized
in Jain et al.(2021)’s study [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. To address this limitation, we propose a constraint-aware negative
sampling technique that leverages informal SKOS constraints for triple corruptions. Additionally,
we provide evidence that the predicted SemOpenAlex concept connections align with established
standards, by validating the predictions with the Unified Medical Language System (UMLS) ontology.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        In this context, Jain et al. (2021) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] point out a weakness in common methods for learning from
Knowledge Graphs (KGs). They argue that these methods assume KGs are locally complete, however
this can not guarantee truly incorrect negative samples for the training. To address this, the authors
introduce a method called ReasonKGE. ReasonKGE assesses whether the learning process aligns
with the description logic used. If not, it identifies predictions that violate the logic and includes
them, along with the similarly generated triples, as negative samples for the next training iteration.
The authors experiments conducted indicate that ReasonKGE produces more accurate predictions
compared to standard methods. On the other hand, Alam et al. (2020) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] highlight a common
challenge in KGE models related to the random generation of negative samples. Commonly, KGE
models follow the Close World Assumption (CWA). They treat unseen triples in KGs as unknown
and rely on existing data. To overcome this limitation, the authors introduce a negative sampling
approach that considers triple afinity. Initially, the method calculates the distance of a candidate
triple for corruption using cosine similarity, and compares it with the remaining triples in the KG.
Subsequently, an afinity function incorporates this distance and the comparison to find the fitness of
each entity for the corruption task. As its output, the function produces a set of entities with their
likelihood scores for corruption. Random selection of these entities forms a batch of negative samples
from the candidate triple. Ultimately, the proposed method enhances the performance of KGE models,
particularly in terms of time eficiency, for link prediction tasks as demonstrated in their experiments.
Another work is conducted by Yao et al. (2022) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] which aims to enhances the quality of negative
samples in KGE models. Their approach takes into account triples with similar context in the KG,
and generates more meaningful negative samples. For example, given the positive triple (Steve Jobs,
FounderOf, Apple Inc.), a higher-quality negative sample would be (Jerry Yang, FounderOf, Apple
Inc.) rather than a less contextually relevant one like (Yahoo, FounderOf, Apple Inc.). The method
assesses the quality of a negative sample based on its closeness to the positive entity during training.
After, it selects the most similar entity for corruption. The authors conducted experiments comparing
their approach to state-of-the-art methods in link prediction tasks, and their method demonstrated
better performance [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8, 9, 10</xref>
        ]. While previous works concentrate on general rules or KG structures
when generating custom negative samples, our approach distinguishes itself by emphasizing semantic
simplicity and employing specific rules. This approach improves the reliability of generating accurate
negative samples.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>This paper focuses on refining noisy SKOS relations in the SemOpenAlex scholarly KG. We formulate
the problem as a KG completion task with the aim to improve the concept hierarchy. We apply a
custom negative sampling method for KG completion models. This method is based on the World
Wide Web Consortium (W3C) informal SKOS constraint definitions. We use Knowledge Graph
Embedding (KGE) and Graph Neural Network (GNN) methods to measure the predictive accuracy of
our approach.</p>
      <p>
        Specifically, we experiment with baseline KGE models TransE [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ], DistMult [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and
QuatE [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] along with a GNN model. TransE is a translation type model which embeds vector
representations for entities with their relation to the translation of the entities in low dimensional
vector space. More specifically, the embedding of an tail entity is expected to be in close proximity to
the embedding of a head entity, augmented by their relationship vector. For instance, in the triple
(Steve Jobs, FounderOf, Apple Inc.), "Steve Jobs" is the head entity, "Apple Inc." is the tail entity,
and they are connected by the "FounderOf" relation. In TransE, the location of "Steve Jobs" plus
"FounderOf" is expected to be close to "Apple Inc." in the embedding space. RESCAL [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] models each
relation as a matrix and models each triple as a three-way interaction between the head, relation
and tail entity. However, RESCAL has quadratic number of relational parameters. DistMult [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
simplifies RESCAL by restricting matrices representing relations as diagonal matrices. Due to the
simple diagonal matrix relational modeling, DistMult cannot model symmetric relations. QuatE [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
embeds entities and relations as quaternions–4D hypercomplex numbers with one real component
and three imaginary components. For each triple, QuatE first rotates the head quaternion number by
performing a Hamilton Product with the corresponding relation quaternion. The plausibility of the
triple is then measured by the quaternion inner product between the rotated head entity and the tail
entity. QuatE has been proven to subsume DistMult and ComplEx. These methods produce negative
samples by randomly mixing head and tail entities. Since they follow the Close World Assumption,
there is a possibility for generating negative samples which are actually correct [
        <xref ref-type="bibr" rid="ref11 ref13 ref14 ref8">8, 11, 13, 14</xref>
        ].
      </p>
      <p>
        We first applied SKOS inconsistency definitions, formulation 1 and formulation 2, to identify groups
of triples that did not satisfy Constraint-1 or Constraint-2. More specifically, the formulation 1, implies
that there should not be entity  that is broader than entity , and entity  that is related to entity ,
as it would create an inconsistency. Likewise, the formulation 2 states that for all entities a, b, and c,
if  is broader than  and  is related to , then it should not be the case that  is broader than  to
maintain consistency [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>∀,  :  : (, ) ∧  : (, ) → ⊥
Constraint-2:
∀, ,  : ( : (, ) ∧  : (, ))</p>
      <p>→ ¬ : (, )</p>
      <p>
        Second, we cleaned the KG from the links that conflict with Constraint-1 and Constraint-2 to
KG completion step. For the link prediction, our method utilised KGE and GNN models. Briefly,
KGE is a machine learning technique which is used for KG applications including KG completion
and recommender systems. KGE representation reduces the complexity of the graph structure by
mapping KG relations and entities to a low-dimensional vector space. Namely, the KGE model maps
KG triples into the embedding space by taking into account positive and negative triples. During the
training process, model should learn to score positive triples higher than negative ones. Through an
iterative learning, the model improves its understanding of the semantic and structural patterns
present in the KG. On the other hand, GNNs enhance entity-centric KGE models by incorporating
neighboring information and graph structure during model training. The main goal of a GNN model
is to embed entities while considering the encoding of their neighbors in the graph [
        <xref ref-type="bibr" rid="ref17 ref18 ref8">8, 17, 18</xref>
        ].
      </p>
      <p>
        Third, we addressed the negative sampling issue by implementing a customized approach to
improve the quality of negative samples. We achieved this by using SKOS inconsistency definitions
which are described on the W3C Spec webpage1. Initially, we defined SKOS inconsistency definitions
as first-order logic formulas for having structured inconsistency representations. Then, we generated
corrupted triples by using these definitions on the KG triples during KGE model training as additional
negative samples. Lastly, we explored how an increased quantity of noisy samples afects KGE link
prediction accuracy on the SemOpenAlex concept hierarchy[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>In the final step, we trained KGE models using custom constraint-aware negative sampling. The
primary goal is to demonstrate that this method produces predictions that refine previously identified
erroneous connections. Since we initially remove all triples contributing to failed groups, it becomes
crucial to understand which triples are causing conflicts. To achieve this, we leveraged the Unified
Medical Language System (UMLS) ontology. Specifically, we contextually matched Medicine and its
sub-concepts in SemOpenAlex with UMLS concepts to create a proofing test set. This allows us to
distinguish which triples align with SKOS constraints and which ones introduce noise. We assume
the presence of a parent-child relation in UMLS proves their hierarchical relationship, so we add such
triples into the test set.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Setup</title>
      <p>To generate a smaller and more manageable dataset from SemOpenAlex, we extracted concepts
related to Medicine and its sub-concepts, as well as their relations using skos:related and skos:broader
relations. Initially, the dataset contained 345.119 triples with 51.885 diferent concepts. Following that,
we identified and removed 1.456 conflicting triples for Constraint-1 and 1.044 triples for Constraint-2
from the dataset. Figure 4 illustrates an example inconsistent group of triples. The inconsistencies
are used as negative samples as previously detailed out. For each training triplet, we perform
experiments with no negative sampling, a negative sample drawn at random, or using the proposed
negative sampling according to inconsistency constraints. By introducing these semantic negatives
during training, we aim to improve the model’s understanding about the dataset.</p>
      <p>In order to further test our methodology, we utilize the conflicting triples as a proof set, by focusing
on the skos:broader relation and extracting entities that have a matching UMLS concept. For instance
as shown in Figure 4, the SemOpenAlex concept Pediatrics (C187212893) was part of a conflict group
that matched with the UMLS Pediatrics (C0030755) concept. After identifying the match, we added
(Pediatrics, skos:broader, Medicine) triple into the test set since this UMLS pair can be a proof of
hierarchical connection between Pediatrics and Medicine concepts. We extended this matching
process for all Constraint-1 and Constraint-2 conflicts to obtain 204 triples for Constraint-1 and 203
triples for Constraint-2, ultimately used as a test set. Table 1 provides the details on how the data is
set up for training, validation, and testing the models.</p>
      <p>
        We considered three diferent scenarios under which the negative sampling’s efect is measured. In
the first scenario, no negative sampling is used in order to establish the capabilities of each model to
learn embeddings solely based on the positive samples. In the second scenario, a random negative
sampling component is added in order to establish a baseline for comparison. In the third scenario,
the proposed constraint-aware method is evaluated against the baseline methods. To compensate
for the imbalance in the number of negative samples, we further bootstrap this case at a ratio of
10% as follows [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. In every tenth iteration during training, we enforce a full batch to include only
constraint-based negative samples as our mixing strategy [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        The embedding methods are trained for 100 epochs without early stopping and embedding
dimensions of size 30 for TransE and DistMul, and size 50 for QuatE and GNN for best performance as
reported in the original publication. The parameters are set empirically and according to the settings
in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. During the training process, we use an adaptive stochastic optimizer with a batch
size of 10,000 for training and hold out a validation set of size 1,000 (9-to-1 ratio), as in Table 1. The
learning rate is set to 0.005 for all methods. The QuatE loss function is regularized with  = 0.05.
The GNN uses two convolution layers with an intermediate ReLu activation function. The code is
accessible on our Github: https://github.com/ozyygen/predict-KGE-SKOS.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Result and Discussion</title>
      <p>We conduct experiments with TransE, DistMult, QuatE KGE models and a GNN to evaluate the eficacy
of the proposed negative sampling on link prediction accuracy. Link prediction is one of the key
components in knowledge completion and therefore the refinement of SemOpenAlex concept hierarchy.
We train the models via the training process described above, followed by a 10-fold validation. The
trained and validated models are then assessed on the UMLS-verified triples. The validation set results
are shown in Table 2 and the test set results are shown in Table 3. We compare the performance in each
case according to Hits@K, Mean Rank (MR), and Mean Reciprocal Rank (MRR). The Hits@K metric
measures the precision of the model, while the MR and MRR are reported for generalizability purposes.</p>
      <p>From Table 2, it can be observed that including negative sampling given the same number of
training epochs improves the performance. The results indicate that including negative samples
(random or constraint-based) improves the training accuracy across the board. For instance, for either
TransE or DistMult, hits@10 is consistently above 0.70 using either constraint sets. The improvement
is more palpable in the case of TransE and Hits@10 metric, where an increase of around 0.60 points
is observed for the first set of SKOS constraints and an increase 0.73 for the second set of SKOS
constraints, over the baseline. This observation highlights the inherent diference in the constraints
used and the efect this phenomenon can have on the ‘translation’ based methods such as TransE.
Unlike TranE, the baseline Hits@10 results of DistMult already are competitive at 0.70 and 0.74
due to its algorithmic benefits. Nevertheless, still some improvements can be achieved using the
proposed negative sampling. The improvements over the base line are 0.16 points in the case of the
ifrst set of SKOS constraints and 0.03 points in the case of the second set of SKOS constraints. It
scores 0.08 points better using the second set of SKOS constraints, while the first set of constraints
appears to have a neutral efect. Similar to TransE, the choice of negative samples and whether the
ifrst or second constraint set is used, conclusively make a diference. The GNN and QuatE are widely
adopted methods in knowledge embedding and reasoning tasks due to their expressive power. As
expected, the baseline validation performance without negative sampling using either method is
remarkable, with Hits@10 results of 0.78 and 0.8 from the GNN as well as 0.9 and 0.87 from QuatE,
under the first set of SKOS constraints and the second set of SKOS constraints respectively. The slight
diference observed in performance under either constrain set is due to the structural diferences in
the resulting knowledge graphs after applying the cleaning step described in the previous section.</p>
      <p>Under Constraint 1 refinement criteria, the GNN’s performance improves generally with the
addition of negative samplings, however random sampling evidently has as good a performance as
the proposed one. Under Constraint 2 however, the proposed approach gains an advantage of 0.02
points in Hits@10. The reason the efect of the proposed negative sampling is mitigated in the first
case is deemed to be due to the fact that more entities are involved in the second case which can have
an impact based on the methodology of GNN which is a graph-based algorithm and therefore locality
information and the number of neighbouring nodes (two versus three) can play an important role.
Under the second SKOS constraint with three entities involved, an improved Hits@1 and Hits@10 of
0.71 and 0.82 points are achieved. Among all validated methods, QuatE together with the proposed
sampling method achieves the best performance with a Hits@10 of 1.00 and a remarkable MRR of
0.85 Under either SKOS constraint, using negative sampling improves performance. The addition of
constraint-based sampling, specially in the second case, achieves an improvement of 0.02 points from
0.94 to 0.98 in Hits@10 and from 0.56 to 0.69 in Hits@1. The MRR scores accordingly improve as well.</p>
      <p>Lastly, we evaluate our constraint-aware method using a test set for each KGE method and report the
results in Table 3. The test set is created according to the UMLS hierarchy such that the skos:broader
predicate is valid. Accordingly, this test set evaluates the performance for link prediction in each
case based on the quality of the embeddings and the underlying refined hierarchy. It can be seen
that consistent with the validation results, QuatE and GNN continue to achieve the best performance
under both constraint criteria. The GNN’s high performance with respect to Hits@1 and Hits@10
which are a measure of prediction accuracy, reflect the ability of GNNs in capturing the underlying
structure of the ontology. QuatE benefits from more degrees of freedom and in both constrain settings
remains the best performing KGE model after refinement through our proposed approach. These
results corroborate the strength of using verified triples from a well maintained medical ontology.</p>
      <p>MRR
0.13
0.65
0.68
0.05
0.08
0.08
0.08
0.12
0.07
0.71
0.87
0.85</p>
      <p>Hits@1
0.14
0.53
0.55
0.06
0.03
0.06
0.66
0.68
0.71
0.56
0.56
0.69</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This study assesses the impact of SKOS constraints, informally defined, on the predictive performance
of KGE models. Our experiments reveal that incorporating logically inferred negative samples during
model training enhances learning by leveraging logical formulations derived from SKOS website’s text
definitions. By injecting two types of negative samples into QuatE embedding model, the proposed
method achieves Hits@1 of 0.85, Hits@10 of 1.00, MR of 1.22, and MRR of 0.92 in one scenario,
and Hits@1 of 0.83, Hits@10 of 1.00, MR of 1.26, and MRR of 0.91 in the other scenario depending
on the selected SKOS constraints, which we have verified against the well-known UMLS medical
ontology. We plan to extend this research as a future work by covering the remaining concepts and
their sub-classes in the SemOpenAlex concept hierarchy. We will also explore the incorporation of
ontological axioms into the KGE model learning process.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] W. of Science, Web of science https://www.webofscience.com/,
          <year>2023</year>
          . URL: https://www. webofscience.com/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Bonatti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Presutti</surname>
          </string-name>
          ,
          <article-title>Knowledge graphs: New directions for knowledge representation on the semantic web (dagstuhl seminar 18371), in: Dagstuhl reports</article-title>
          , volume
          <volume>8</volume>
          ,
          <string-name>
            <surname>Schloss</surname>
            <given-names>Dagstuhl</given-names>
          </string-name>
          <source>-Leibniz-Zentrum fuer Informatik</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dalle Lucca Tosi</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. C.</surname>
          </string-name>
          <article-title>dos Reis, Understanding the evolution of a scientific field by clustering and visualizing knowledge graphs</article-title>
          ,
          <source>Journal of Information Science</source>
          <volume>48</volume>
          (
          <year>2022</year>
          )
          <fpage>71</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] SemOpenAlex, Semopenalex ontology https://semopenalex.org/,
          <year>2023</year>
          . URL: https://semopenalex. org/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Caracciolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stellato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Morshed</surname>
          </string-name>
          , G. Johannsen,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rajbhandari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jaques</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Keizer,</surname>
          </string-name>
          <article-title>The agrovoc linked dataset</article-title>
          ,
          <source>Semantic Web</source>
          <volume>4</volume>
          (
          <year>2013</year>
          )
          <fpage>341</fpage>
          -
          <lpage>348</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Baker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Whitehead</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Musker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Keizer</surname>
          </string-name>
          ,
          <article-title>Global agricultural concept space: lightweight semantics for pragmatic interoperability</article-title>
          ,
          <source>npj Science of Food</source>
          <volume>3</volume>
          (
          <year>2019</year>
          )
          <fpage>16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Musa Aliyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ojo</surname>
          </string-name>
          ,
          <article-title>Towards building a knowledge graph with open data-a roadmap, in: e-Infrastructure and e-Services for Developing Countries: 9th International Conference</article-title>
          , AFRICOMM 2017, Lagos, Nigeria,
          <source>December 11-12</source>
          ,
          <year>2017</year>
          , Proceedings 9, Springer,
          <year>2018</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.-K. Tran</surname>
            ,
            <given-names>M. H.</given-names>
          </string-name>
          <string-name>
            <surname>Gad-Elrab</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Stepanova</surname>
          </string-name>
          ,
          <article-title>Improving knowledge graph embeddings with ontological reasoning</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2021</year>
          , pp.
          <fpage>410</fpage>
          -
          <lpage>426</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>M. M. Alam</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Jabeen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ali</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Mohiuddin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <article-title>Afinity dependent negative sampling for knowledge graph embeddings</article-title>
          .,
          <source>in: DL4KG@ ESWC</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <article-title>Entity similarity-based negative sampling for knowledge graph embedding</article-title>
          ,
          <source>in: Pacific Rim International Conference on Artificial Intelligence</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>73</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bordes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Usunier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Garcia-Duran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Yakhnenko</surname>
          </string-name>
          ,
          <article-title>Translating embeddings for modeling multi-relational data</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>26</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehryar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Celebi</surname>
          </string-name>
          ,
          <article-title>Improving transitive embeddings in neural reasoning tasks via knowledgebased policy networks</article-title>
          ,
          <source>in: CEUR Workshop Proceedings</source>
          , volume
          <volume>3337</volume>
          <source>of CEUR Workshop Proceedings</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>Yang</surname>
          </string-name>
          , W.-t. Yih,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <article-title>Embedding entities and relations for learning and inference in knowledge bases</article-title>
          ,
          <source>arXiv preprint arXiv:1412.6575</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Quaternion knowledge graph embeddings</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nickel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Tresp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-P.</given-names>
            <surname>Kriegel</surname>
          </string-name>
          , et al.,
          <article-title>A three-way model for collective learning on multirelational data</article-title>
          .,
          <source>in: Icml</source>
          , volume
          <volume>11</volume>
          ,
          <year>2011</year>
          , pp.
          <fpage>3104482</fpage>
          -
          <lpage>3104584</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Isaac</surname>
          </string-name>
          , E. Summers,
          <article-title>Skos simple knowledge organization system primer</article-title>
          , Working Group Note,
          <volume>W3C</volume>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Choudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Luthra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mittal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <article-title>A survey of knowledge graph embedding and their applications</article-title>
          ,
          <source>arXiv preprint arXiv:2107.07842</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. O.</given-names>
            <surname>Sing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>A comprehensive survey of graph neural networks for knowledge graphs</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>75729</fpage>
          -
          <lpage>75741</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , G. Ye,
          <string-name>
            <given-names>H.</given-names>
            <surname>He</surname>
          </string-name>
          , Not All Negatives Are Worth Attending to:
          <source>Meta-Bootstrapping Negative Sampling Framework for Link Prediction</source>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2312.04815.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kalantidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Sariyildiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Weinzaepfel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Larlus</surname>
          </string-name>
          ,
          <article-title>Hard Negative Mixing for Contrastive Learning</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2020</year>
          , pp.
          <fpage>21798</fpage>
          -
          <lpage>21809</lpage>
          . URL: https://proceedings.neurips.cc/paper/2020/ hash/f7cade80b7cc92b991cf4d2806d6bd78-Abstract.html.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehryar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Celebi</surname>
          </string-name>
          ,
          <article-title>Semantic Annotation of Tabular Data for Machine-to-Machine Interoperability via Neuro-Symbolic Anchoring</article-title>
          ,
          <source>in: CEUR Workshop Proceedings</source>
          , volume
          <volume>3557</volume>
          ,
          <string-name>
            <surname>Rheinisch-Westfaelische Technische Hochschule Aachen* Lehrstuhl Informatik</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <year>2023</year>
          , pp.
          <fpage>61</fpage>
          -
          <lpage>71</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3557</volume>
          /paper5.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Loesch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Celebi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Enhanced</surname>
            <given-names>GAT</given-names>
          </string-name>
          :
          <article-title>Expanding Receptive Field with Meta Path-Guided RDF Rules for Two-Hop Connectivity (</article-title>
          <year>2023</year>
          ). URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3592</volume>
          / paper8.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>