<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Knowledge Graph Embeddings for biomedical Research: Data Integration using Linked Open Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jens Dorpinghaus</string-name>
          <email>jens.doerpinghaus@scai.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marc Jacobs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fraunhofer Institute for Algorithms and Scienti c Computing</institution>
          ,
          <addr-line>Schloss Birlinghoven, Sankt Augustin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Knowledge Graphs are becoming a key instrument for biomedical knowledge discovery and modeling. These approaches rely on structured data, e.g. about related proteins or genes, and form cause-ande ect networks or { if enriched with literature data and other linked data sources { knowledge graphs. A key aspect of analysis on these graphs is the missing context. Here we present a novel semantic approach towards a context enriched Knowledge Graph for biomedical research utilizing data integration with linked data. The result is a general graph concept that can be used for graph embeddings in di erent contexts or layers.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Introduction</p>
      <p>
        Biological and medical researchers considering computational approaches rely
on structured data, e.g. about related proteins or genes, see [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Cause-and-e ect
networks are a special subtype of more general Knowledge Graphs. In principle,
the integration of external data sources and manual curated data is key. Although
several commercial solutions exist, Fakhry et al. state, that the "adoption and
extension of such methods in the academic community has been hampered by the
lack of freely available, e cient algorithms and an accompanying demonstration
of their applicability using current public networks" [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        This and the emerging improvements on large-scale Knowledge Graphs and
machine learning approaches are the motivation for our novel approach on
semantic Knowledge Graph embeddings for biomedical research utilizing data
integration with linked open data. Several similar approaches (often in the context
of drug-repurposing) have been described like Bio2RDF [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], hetionet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], or Open
PHACTS [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Our approach is more focussed on integrating the literature itself
in a FAIR [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and open knowledge graph which is also accessible from public a
public resource: SCAIView3. SCAIView is an information retrieval system that
allows semantic searches in large textual collections by ontological
representations of automatic recognized biological entities [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
3 https://www.scaiview.com/
Copyright ©2019 for this paper by its authors. Use permitted under Creative Commons License
Attribution 4.0 International (CC BY 4.0).
      </p>
      <p>Dorpinghaus, Jacobs et al.</p>
      <p>The basis for generating our large-scale Knowledge Graph representation is
the biomedical literature, e.g. MedLine and PubMed4. These articles or abstracts
are the source for biological relations mentioned above. In addition, meta
information like authors, journals, keywords (so called MeSH-Terms, Medical Subject
Headings), etc. are freely available. Ontologies can be used to contextualize
entities in the Knowledge Graph providing biological or medical relations (cf. 5).
Every ontology will form another knowledge (sub-)graph.</p>
      <p>Using methods of natural language processing (NLP) and text mining, we
can combine and link these knowledge graphs to a giant and very dense new
knowledge graph. This will meet a very general de nition of context. We can see
every knowledge (sub-)graph as context to another. Biological expressions are
context of the corresponding literature, authors are context of a text, named
entities from ontologies found in a text are context to it or to the corresponding
biological expressions.</p>
      <p>
        Our overarching integration schema is based on the Biological Expression
Language6 is widely applied in biomedical domain to convert unstructured
textual knowledge into a computable form. The BEL statements that form
knowledge graphs are semantic triples that consist of concepts, functions and
relationships. Thus they can be easily added to a knowledge graph representing another
layer or context. An example for a large Alzheimer network can be found in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>In the next section we describe the novel concept of semantic graph
embeddings within large-scale Knowledge Graphs. We will present several use-cases
4 See https://www.ncbi.nlm.nih.gov/pubmed/.
5 OLS, https://www.ebi.ac.uk/ols/index
6 BEL, www.openbel.org
and application examples as well as the semantic interoperability layer using
RDF and SPARQL.
2</p>
      <p>Knowledge Graph architecture
A Knowledge Graph is a systematic way to connect information and data to
represent common knowledge. As described above, the context is the most
important topic to generate knowledge or even wisdom.</p>
      <p>
        We de ne knowledge graphs G = (E; R) with entities e 2 E coming from a
formal structure like an ontology O, see [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The relations r 2 R can be
ontology relations, thus in general we can say every ontology O which is part of
the data model is a subgraph of G which means O G. In addition we allow
inter-ontology relations between two nodes e1; e2 with e1 2 O1, e2 2 O2 and
O1 6= O2. More general we de ne R = fR1; :::; Rng as list of either ontologies,
terminologies or any sort of controlled vocabulary containing relations or not.
      </p>
      <p>We de ne contexts C = fc1; :::; cmg as a nite, discrete set. Every node
v 2 G and every edge r 2 R may have one ore more contexts c 2 C denoted by
con(v) or con(r). It is also possible to set con(v) = ;. Thus we have a mapping
con : E [R ! P(C). If we use a quite general approach towards context, we may
set C = E. Thus every inter-ontology relation de nes context of two entities,
but also the relations within an ontology can be seen as context, see gure 1
for an illustration. Here every context is identi ed as a layer (e.g. a document
layer, a molecular layer, a mechanism layer, ...). This allows new connections
between di erent contexts or layers: If two edges e1; e2 2 R1 are connected and
e01; e02 2 R2 with con(e1) = e01 and con(e2) = e02 are not connected, we may add
another edge (e1; e2) with provenance information that this connection comes
from a di erent context, namely R2. Since every layer or context can be seen as
a subgraph forming a surface we can denote the relation between two layers a
knowledge graph embedding.</p>
      <p>It is also possible to get the context of a subgraph Ri G which can be
denominated by con(Ri) or with the notation of graph theory as the extended
induced subgraph by the vertex set Ei from Ri given by Gc[Ei]. This is quite
trivial if context from Ri can only be annotated to vertices in G. Then</p>
      <p>Gc[Ei] = G[Ei] [ f(e; e0) 8e0 2 N (e); e 2 Eig
Here conjEi = Gc[Ei] is the context of Ei restricted to the set of edges (relations)
in the graph. The two edges e0; e00 are implicitly given by this context. It is quite
easy to see that the restriction on context annotated to edges makes the problem
more easy from a computational perspective. Nevertheless, context on edges is
needed from a real-world perspective.</p>
      <p>
        The technical design was done with respect to the microservice architecture
of SCAIView [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We o er both a REST API as well as a Java Message Service
(JMS) interface. As a database backend, we used Neo4j. Here, we used Spring
Data Neo4j to map objects to graphs. Thus our software can be used to perform
Cypher and SPARQL queries. Data can be retrieved in JSON Graph Format or
RDF format.
rpinghaus,
      </p>
      <p>Jacobs
et</p>
      <p>al.
l
e
B
s
a
h
l
e
B
s
a
h</p>
      <p>l
HGNeC
B
s
a
h
hasBel
a
tion
hasBel BELRelation
BELRelation
BELRelation
BELRelation
hashBaeslEvidenc BELRelation
n
tio
ela
R
L
E
B
el
B
s
a
h
el
B
s
haHsSubgraph ha
tion
unc
hasF
ha
sF
un
c
tion
for
some
illustrations
of
di
erent
context
layers.
questions
can
b e
formulated
as
subgraph
structures
of</p>
      <p>For
the
example,</p>
      <p>semaninitial
knowledge
We
may
think
of
complex
examples,
e.g.</p>
      <p>"Give
me
all
pathways
from
A
to</p>
      <p>B
in
the
context
of</p>
      <p>Disease</p>
      <p>C
fo cusing
on
clinical
trials".</p>
      <p>to
of
context,</p>
    </sec>
    <sec id="sec-2">
      <title>PubMed</title>
      <p>networks
Hyp othesis
generation
within
medical
research
and
digital
health
may
lead
to
search
which
for
genomic
or
moleculare
patterns,
diagnosis
or
build
longitudinal
mo dels
build
the
basis
for
a
multitude
of
predictive
and
p ersonalised
medicine
ML
and</p>
      <p>AI
approaches.
This
information
system
can
b e
used
to
retrieve
data
by
context
(cohort
size,
settings,
results,
..)
and
by
content
(imaging
data,
genomic
or
moleculare
measures,
...).</p>
      <p>For
example,
this
system
may
answer
questions
like
Give
me
a
clinical
trial
to
repro duce
my
results
or
to
apply
me
literature
for
phenotyp e
A,
disease
B
age
b etween
C
and
my
D
mo del
or</p>
    </sec>
    <sec id="sec-3">
      <title>Give and a</title>
    </sec>
    <sec id="sec-4">
      <title>CT-scan with characteristic E.</title>
      <p>Here we presented a novel approach that annotates research data with context
information. The result is a knowledge graph representation of data, the context
graph. It contains computable statement representation (e.g. RDF or BEL). This
graph allows to compare research data records from di erent sources as well as
the selection of relevant data sets using graph-theoretical algorithms.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>1. Guidelines for the construction, format, and management of monolingual controlled vocabularies</article-title>
          .
          <source>Standard</source>
          , National Information Standards Organization, Baltimore, Maryland,
          <string-name>
            <surname>U.S.A.</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Belleau</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nolin</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tourigny</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rigault</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morissette</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Bio2rdf: towards a mashup to build bioinformatics knowledge systems</article-title>
          .
          <source>Journal of biomedical informatics 41(5)</source>
          ,
          <volume>706</volume>
          {
          <fpage>716</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Drpinghaus</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darms</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacobs</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Scaiview { a semantic search engine for biomedical research utilizing a microservice architecture</article-title>
          .
          <source>In: Proceedings of the Posters and Demos Track of the 14th International Conference on Semantic Systems - SEMANTiCS2018</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Fakhry</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choudhary</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gutteridge</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sidders</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ziemek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zarringhalam</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Interpreting transcriptional changes using causal graphs: new methods and their practical utility on public networks</article-title>
          .
          <source>BMC bioinformatics 17(1)</source>
          ,
          <volume>318</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Harland</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Open phacts: A semantic knowledge infrastructure for public and commercial drug discovery research</article-title>
          .
          <source>In: International Conference on Knowledge Engineering and Knowledge Management</source>
          . pp.
          <volume>1</volume>
          {
          <issue>7</issue>
          . Springer (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Himmelstein</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lizee</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hessler</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brueggeman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hadley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Green</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khankhanian</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baranzini</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <article-title>Systematic integration of biomedical knowledge prioritizes drugs for repurposing</article-title>
          .
          <source>Elife</source>
          <volume>6</volume>
          ,
          <issue>e26726</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hodapp</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fluck</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zimmermann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Integration of UIMA Text Mining Components into an Event-based Asynchronous Microservice Architecture</article-title>
          .
          <source>In: Proceedings of the LREC 2016 Workshop "Cross-Platform Text Mining and Natural Language Processing Interoperability"</source>
          . pp.
          <volume>19</volume>
          {
          <fpage>23</fpage>
          .
          <string-name>
            <surname>European Language Resources Association</surname>
          </string-name>
          (ELRA), Portoroz, Slovenia (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kodamullil</surname>
            ,
            <given-names>A.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Younesi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bagewadi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hofmann-Apitius</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Computable cause-and-e ect models of healthy and alzheimer's disease states and their mechanistic di erential analysis</article-title>
          .
          <source>Alzheimer's &amp; Dementia</source>
          <volume>11</volume>
          (
          <issue>11</issue>
          ),
          <volume>1329</volume>
          {
          <fpage>1339</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sewer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Talikka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoeng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peitsch</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          :
          <article-title>Quanti - cation of biological network perturbations for mechanistic insight and diagnostics using two-layer causal models</article-title>
          .
          <source>BMC bioinformatics 15(1)</source>
          ,
          <volume>238</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Wilkinson</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumontier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aalbersberg</surname>
            ,
            <given-names>I.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Appleton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Axton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomberg</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boiten</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          , da Silva Santos,
          <string-name>
            <given-names>L.B.</given-names>
            ,
            <surname>Bourne</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.E.</surname>
          </string-name>
          , et al.:
          <article-title>The fair guiding principles for scienti c data management and stewardship</article-title>
          .
          <source>Scienti c data 3</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Knowledge organization systems (kos</article-title>
          )
          <volume>35</volume>
          ,
          <fpage>160</fpage>
          {
          <volume>182</volume>
          (01
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>