<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Knowledge Graph for Ecotoxicological Risk Assessment and E ect Prediction?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Erik B. Myklebust</string-name>
          <email>erik.b.myklebust@niva.no</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics, University of Oslo</institution>
          ,
          <addr-line>Oslo</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Norwegian Institute for Water Research</institution>
          ,
          <addr-line>Oslo</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Exploring the e ects a chemical compound has on a species takes a considerable experimental e ort. Appropriate methods for estimating and suggesting new e ects can dramatically reduce the work needed to be done by a laboratory. In this PhD research we aim at exploring the suitability of using a knowledge graph embedding approach for ecotoxicological e ect prediction. A knowledge graph is being constructed from publicly available data sets, including a species taxonomy and chemical classi cation and similarity. We use ontology alignment techniques to integrate the e ect data into the knowledge graph. Our preliminary experimental results show that the knowledge graph based approach improves the selected baselines.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge graph</kwd>
        <kwd>Semantic embedding</kwd>
        <kwd>Ecotoxicology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>(ii) Using the knowledge graph together with machine learning techniques to
predict e ects. The objectives of this task are twofold:
(a) Limit the search space for the laboratory (binary prediction).
(b) Predict e ects outright with a margin of error (regression).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background and related work</title>
      <p>In this section we introduce some preliminaries and give insights into the current
state of the art e orts applying semantic web technologies within the eld of
toxicology and risk assessment.</p>
      <p>Use case. Ecotoxicology is a multidisciplinary eld that studies the
ecological and toxicological e ects of chemical pollutants on populations, communities
and ecosystems. Risk assessment is the result of the intrinsic hazards of a
substance combined with an estimate of the environmental exposure (i.e., Hazard
+ Exposure = Risk).</p>
      <p>Figure 1 shows a risk assessment pipeline. Exposure is data gathered from
the environment, while e ects are hypothesis that are tested in a laboratory.
These two data sources are used to calculate risk, which is used to nd (further)
susceptible species and the mode of action (MoA) or type of impact a compound
would have over those species. Results from the MoA analysis are used as new
e ect hypothesis.</p>
      <p>
        E ect prediction. Estimating the e ect a compound has on a species is a large
research eld within ecotoxicology. Currently, state-of-the-art solutions such as
Quantitative Structure-Activity Relationship (QSAR) models ( e.g., [
        <xref ref-type="bibr" rid="ref13 ref14 ref7">7, 13, 14</xref>
        ])
exists. However, these are limited in scope. Each QSAR consider small groups of
compounds and a single or a few species. Therefore, a general approach suited
for a larger subset of the domain is favourable.
Knowledge graphs. We follow the RDF-based notion of knowledge graphs [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
which are composed by RDF triples hs; p; oi, where s represents a subject (a class
or an instance), p represents a predicate (a property) and o represents an object
(a class, an instance or a data value e.g., text, date and number). RDF entities
(i.e., classes, properties and instances) are represented by an URI (Uniform
Resource Identi er). A knowledge graph can be split into a TBox (terminology),
often composed by RDF Schema constructors like class subsumption and
property domain and range,3 and an ABox (assertions), which contain relationships
among instances and semantic type de nitions. RDF-based Knowledge Graphs
can be accessed with SPARQL queries, the standard language to query RDF
graphs.
      </p>
      <p>
        There is emerging work in improving the usability of ecotoxicological data
by mapping to knowledge graphs or ontologies, e.g., [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], however, currently this
work is limited. We are unaware of work incorporating the vast array of sources
that is required from beginning to end by a risk assessment system.
Ontology alignment. Ontology alignment is the process of nding mappings or
correspondences between a source and a target ontology or knowledge graph [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
These mappings are typically represented as equivalences among the entities of
the input resources.
      </p>
      <p>
        Currently, mapping ecotoxicological data to di erent sources are under
construction. The ECOTOX web search interface4 now contains mappings to a
external taxonomy source [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] (for a limited number of taxons). Fay et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
indicates a full mapping to external sources exists, however, this is not yet
publicly available.
      </p>
      <p>We are not aware of e orts toward mapping taxonomic classes, e.g., genus,
family, etc. which can reveal inconsistencies in the datasets.</p>
      <p>
        Embedding models. Knowledge graph embedding [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] plays a key role in link
prediction problems where the goal is to learn a scoring function S : E R E !
R. S(s; p; o) is proportional to the probability that a triple hs; p; oi is encoded
as true. Several models has been proposed, e.g., Translating embeddings model
(TransE) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. These models are applied to knowledge graphs to resolve missing
facts in largely connected knowledge graphs, such as DBPedia [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        There is previous work investigating modelling of chemical e ects, e.g., [
        <xref ref-type="bibr" rid="ref12 ref16">16,
12</xref>
        ]. The prediction of ecotoxicological e ects can be seen as a sub-problem. These
works investigate models that use the chemical structures to determine their
e ect on species. Yet, we are not aware of approaches where multiple knowledge
graph embeddings are used to model the interaction between knowledge graphs.
3 The OWL 2 ontology language provides more expressive constructors. Note that the
graph projection of an OWL 2 ontology can be seen as a knowledge graph (e.g., [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]).
4 https://cfpub.epa.gov/ecotox/
      </p>
    </sec>
    <sec id="sec-3">
      <title>Relevance</title>
      <p>The relevance of the research to be conducted during the PhD can be summarized
as follows:
(i) Manually integrating background knowledge into risk assessment systems
is cumbersome since a common vocabulary does not exists. Our approach
will reduce the time spent organizing data, and increase the number of
case studies than can be conducted. A common vocabulary will enhance
the interoperability between several risk assessment systems, increasing the
con dence in the assessments.
(ii) The e ect data used in risk assessment models is the result of time-consuming
laboratory work. By using machine learning techniques with background
knowledge, in the form of a knowledge graph, we aim at being able to limit
the search space for new tests to be analysed in the laboratory. For
example, we aim at recommending the top-ten compounds to test on a speci c
species, rather than conducting experiments using thousands of possible
compounds.
(iii) Design and implementation of a fully- edged recommender system to
predict the level of e ect on a species. For example, DEET (pesticide) has
the potential to kill 50% of the population of the common house y. Such
generalization using the available data and knowledge is the main target
of the research, which aims at reducing to a minimum further laboratory
analysis.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Research questions and hypothesis</title>
      <p>This work aims to address the following questions:
a. Can the disparate data sources used in ecotoxicological risk assessment be
integrated into a knowledge graph to improve accessibility?
b. Can the knowledge graph be used to improve (or diversify) ecotoxicological
e ect prediction over current state-of-the-art models?
The hypothesis associated with the above questions are:
A. It is possible to integrate disparate data sources in a toxicological knowledge
graph using Semantic Web tools.</p>
      <p>B. Extrapolation of e ect data increase the reach of risk assessment systems
while remaining accurate.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Approaches</title>
      <p>This section will describe the approaches used to investigate the hypothesis
above. The evaluation of the hypothesis is described in Section 7.
Hypothesis A. There are multiple sources, varying from tabular, SPARQL
endpoints, REST APIs, and RDF formats, each with its own vocabulary, that
needs to be integrated to enable a uni ed data access. The main sources of data
are:</p>
      <sec id="sec-5-1">
        <title>Species</title>
        <p>T ransform</p>
      </sec>
      <sec id="sec-5-2">
        <title>Alignment</title>
        <p>T ransform
N CBI
Split</p>
      </sec>
      <sec id="sec-5-3">
        <title>ECOT OX</title>
      </sec>
      <sec id="sec-5-4">
        <title>Ef f ects M ap</title>
      </sec>
      <sec id="sec-5-5">
        <title>Compounds</title>
        <p>TERA-KG</p>
      </sec>
      <sec id="sec-5-6">
        <title>Alignment</title>
        <p>SP ARQL
Import</p>
      </sec>
      <sec id="sec-5-7">
        <title>ChEBI</title>
      </sec>
      <sec id="sec-5-8">
        <title>P ubChem</title>
        <p>
          (i) E ect data (ECOTOX [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], example seen in Table 1) in tabular format.
        </p>
        <p>
          Includes limited metadata linked to proprietary identi ers for compounds
and species.
(ii) Compound data from di erent sources. Hierarchies available through
downloadable RDF les and SPARQL endpoints (PubChem [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] and ChEMBL
[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]). Compound features, e.g., Molecular weight, XLogP etc. are available
through the PubChem REST API.
(iii) The tabular NCBI taxonomy [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] is used as the hierarchy for species.
We must map the identi ers used in the e ect data to open standards to take
advantage of the diversity of data sources. The created Toxicological E ects and
Risk Assessment (TERA) knowledge graph with current sources and aggregation
steps is shown in Figure 2. Excerpts of triples from TERA are shown in Table 2.
test id reference number test cas species number
1068553 5390 877430 (2,6-Dimethylquinoline) 5156 (Danio rerio)
2037887 848 79061 (2-Propenamide) 14 (Rasbora heteromorpha)
result id test id endpoint conc1 mean conc1 unit
98004 1068553 LC50 400 mg=kg diet
2063723 2037887 LC10 220 mg=L
# subject predicate object
(i) ecotox:group/Worms owl:disjointWith ecotox:group/Fish
(ii) ncbi:division/2 owl:disjointWith ncbi:division/4
(iii) ecotox:taxon/34010 rdfs:subClassOf ecotox:taxon/hirta
(iv) ncbi:taxon/687295 rdfs:subClassOf ncbi:taxon/513583
(v) compound:CID10198308 rdf:type obo:CHEBI 134899
(vi) compound:CID10198308 pubchem:formula ``C7H6O6S''
(vii) ecotox:chemical/115866 ecotox:affects ecotox:effect/001
(viii) ecotox:effect/001 ecotox:species ecotox:taxon/26812
(ix) ecotox:effect/001 ecotox:endpoint LC50
(x) ecotox:taxon/33155 owl:sameAs ncbi:taxon/311871
        </p>
        <p>
          Improving the knowledge graph can be done with several sources. First, a
dataset containing biological activity, e.g., Chemical ontology (CO) [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Such
datasets would enable ner grained data to be used by the e ect predictor. We
also aim at including an anatomy dataset, such that the biological activity can
be aggregated from proteins to individual level.
        </p>
        <p>
          Another aspect important to e ect prediction is the habitat of the species,
e.g., [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Including the species habitat data will limit the e ect prediction search
space further, e.g., heavy insoluble compounds (sinks in water) would have
little/no e ect on sh.
        </p>
        <p>
          Hypothesis B. The prediction task at hand is depicted in Figure 3. Initially,
we use a naive approach, which is to assume that similar compounds has a
comparable e ect on the same species and vice versa. The state of the art in
risk assessment systems implement akin solutions. However, it is not clear what
constitutes similarity in this context. The similarity between compounds are
quanti able using di erent methods, however, similarity does not imply
similar biological activity [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. For species, the naive solution is to calculate the
taxonomic distance, but again the classi cations of species is not de ned by
the susceptibility to compounds. Consequently, additional sources that describe
these phenomena need to be added to the knowledge graph. When the
knowledge graph is enriched with this data we can explore modelling techniques for
embedding the knowledge graph for the purpose of predicting e ects. We aim
at applying simple embedding methods, TransE [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], DistMult [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], and HolE
[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], until their performance is exhausted. These model may preform adequately
for producing recommendation to the lab, however, as shown in the next section
these models cannot be fully trusted to predicting e ects outright. Therefore, we
intend to include more expressive models, such as Graph Convolution Networks
(GCN) [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Current approaches in knowledge graph embedding do not consider
sparsely connected knowledge graphs, such as the hierarchical structures that
make up TERA. Therefore, we aim at using the classi cation power of GCNs
to embed groups of species or compounds more accurately. This will include the
use of the vast array of chemical properties (experimental or computed) and the
protein classi cation available for most species.
subClassOf
CR
        </p>
        <p>CA
CB
type
type
type
type
c1
c2
c3</p>
        <p>Affects
Not affects</p>
        <p>Affects
Affects
s1
s2
s3
type
type
type</p>
        <p>SA
SB
subClassOf
subClassOf</p>
        <p>SR
We have evaluated three (plus variants) prediction models based on the e ect
data and the TERA knowledge graph. Note that currently the TERA knowledge
graph has been created with the bare minimum of sources required to integrate
the e ect data with external metadata for compounds and species. Selected
results are shown in Table 3. Prediction models used in this preliminary evaluation:
(i) A nearest-neighbour approach. A compound-species pair can inherent an
e ect if another compound or species is close in the knowledge graph.
This method provides a useful baseline. However, the performance of this
method is far from ideal, as it either will over or underestimate e ects
based on the number of neighbours considered.
(ii) A zero-background-knowledge multi-layer perceptron (MLP) model was
applied to the e ects data. This model is able to learn simple relations, e.g.,
(a) Accuracy for the MLP prediction models.</p>
        <p>
          (b) Recall for the MLP prediction models.
s1 and s2 is e ected by c3, therefore, c3 is toxic and will e ect s3. However,
when this model is presented with previously unseen compound-species
pairs, it cannot rely on background knowledge, and hence, the prediction
will be highly awed.
(iii) Using knowledge graph embedding ([
          <xref ref-type="bibr" rid="ref21 ref25 ref5">5, 25, 21</xref>
          ]) on TERA, followed by the
same MLP model architecture as above yields better results for recall
(which is preferred), while accuracy remains similar. In contrast to the
above model, this model is more uncertain when unseen combinations are
presented to the model (in dubio pro reo). As shown in Figures 4a and
4b, lowering the decision threshold (from 0:5 to 0:35) would yield a higher
recall (0:93) for the HolE-based model, while reducing the accuracy (0:75).
        </p>
        <p>The obtained predictions are promising and show the potential usefulness of
the machine learning models in our setting and the bene ts of using the TERA
knowledge graph. As mentioned before, we favour recall with respect to precision.
One the one hand, false positives are not necessarily harmful, while overlooking
the hazard of a chemical may have important consequences. On the other hand,
due to the limited experiments in terms of concentration (i.e., e ect data may
not be complete), some chemicals may look less toxic than others while they may
still be hazardous. At the same time the adoption of a RDF-based knowledge
graph enables the use of an extensive range of Semantic Web infrastructure
that is currently available (e.g., reasoning engines, ontology alignment systems,
SPARQL query engines).
7</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Evaluation plan</title>
      <p>In this section, we introduce the evaluation plan for the success of this project.
We can divide the evaluation of both research questions into qualitative and
quantitative measures.</p>
      <p>The value of the knowledge graph in toxicology research is uncertain at this
stage. The knowledge graph must provide value for the researchers. We can
ensure this by evaluating the quality of the knowledge graph. Our de nition of
quality is that the knowledge graph should have high levels of:
(i) Coverage. The sources included in the knowledge graph must cover the
areas of interest. The coverage also relates to the degree of successful
mappings between the sources. There will be a trade-o between completeness
and correctness of the mappings.
(ii) Integration. The ease of integrating the various sources. This involves
aligning and mapping to attain a consolidated knowledge graph.
(iii) Functionality. The ability of the knowledge graph to be integrated into the
risk assessment systems. This includes keeping the exibility of Semantic
Web technology without a commitment to a schema. We can add new
triples and extend the knowledge graph, without the need of major changes.
(iv) Embedding enrichment. The semantic enrichment the knowledge graph
gives the embeddings compared to embeddings learned from e ects.</p>
      <p>The quantitative evaluation of the knowledge graph generation is tightly
related to that of evaluating the e ect prediction. We can evaluate the ability to
make good e ect predictions quite easily. This can be done with precision, recall,
accuracy, etc. for binary e ects or with mean squared error, R2-score, etc. for
regression. However, the value of these predictions is integrating them into the
risk assessment pipeline. The evaluation metrics must be inline with the ability
our methods have to enhance risk assessment. We will compare environmental
case studies results before and after the use of our modelling results. Since there is
no ground truth data for risk assessment we rely on domain experts to determine
if our contributions adds value to the assessments.</p>
      <p>Risk assessments has currently large margins of errors (experimental errors
etc.), and we may introduce new sources of error with our e ect predictions.
However, we are con dent that errors can also be reduced by greater data
coverage. These are di erent types of errors and part of the evaluation process will
be to nd the optimal trade-o between them.</p>
      <p>The current preliminary results uses random dataset splits for training and
testing the models. We aim at introducing highly selective datasets that can
test predictive performance in di erent scenarios. We will also try a completely
clean test, where we recommend compound-species pairs to be tested in the lab.
This will obviously be limited by the available compounds and test species of
the particular laboratory.</p>
      <p>The methodologies and knowledge graphs will be publicly available such that
feedback from the community can help us evaluate and improve our
contributions.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Re ections</title>
      <p>
        The conducted work falls into one of the main research lines of toxicology
research to enhance the generation of hypothesis to be tested in the laboratory [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
Furthermore, the data integration e orts and the construction of the TERA
knowledge graph is a large contribution to the area of risk assessment. The
availability and accessibility of the best knowledge and data will enable optimal
decision making.
      </p>
      <p>
        Knowledge graph embedding models have been applied in general purpose
link discovery and knowledge graph completion tasks [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. They have also
attracted the attention in the biomedical domain to nd, for example, candidate
genes for a disease, protein-protein interactions or drug-target interactions (e.g.,
[
        <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
        ]). However, we are not aware of the application of knowledge graph
embedding models in the context of toxicological e ect prediction.
      </p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This PhD project is supported by grant 272414 from the Research Council of
Norway. I would like to thank my supervisors, Ernesto Jimenez-Ruiz (The Alan
Turing Institute and University of Oslo), Raoul Wolf (Norwegian Institute for
Water Research), and Knut Erik Tollefsen (Norwegian Institute for Water
Research) for their feedback on this work. In addition, I would also like to thank
Jiaoyan Chen (University of Oxford), Martin Giese (University of Oslo) and Zo a
C. Rudjord (Norwegian Institute for Water Research) for their contribution in
di erent stages of this PhD research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agibetov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimenez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ondresik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solimando</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guerrini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Catalano</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patane</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reis</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spagnuolo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Supporting shared hypothesis testing in the biomedical domain</article-title>
          .
          <source>J. Biomedical Semantics</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ), 9:
          <issue>1</issue>
          {9:
          <issue>22</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Agibetov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Global and local evaluation of link prediction tasks with neural embeddings</article-title>
          .
          <source>In: 4th Workshop on Semantic Deep Learning (ISWC workshop)</source>
          . pp.
          <volume>89</volume>
          {
          <issue>102</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Alshahrani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maddouri</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kinjo</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Queralt-Rosinach</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoehndorf</surname>
          </string-name>
          , R.:
          <article-title>Neuro-symbolic representation learning on biological knowledge graphs</article-title>
          .
          <source>Bioinformatics</source>
          <volume>33</volume>
          (
          <issue>17</issue>
          ),
          <volume>2723</volume>
          {
          <fpage>2730</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Arnaout</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elbassuoni</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>E ective Searching of RDF Knowledge Graphs</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>48</volume>
          (
          <issue>0</issue>
          ) (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Usunier</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Duran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yakhnenko</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Translating embeddings for modeling multi-relational data</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          , pp.
          <volume>2787</volume>
          {
          <fpage>2795</fpage>
          . Curran Associates, Inc. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. ChEBI-ontology:
          <article-title>The european bioinformatics institute (</article-title>
          <year>2019</year>
          ), https://www.ebi.ac.uk/chebi/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Escher</surname>
            ,
            <given-names>B.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bittermann</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henneberger</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knig</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khnert</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klver</surname>
          </string-name>
          , N.:
          <article-title>General baseline toxicity qsar for nonpolar, polar and ionisable chemicals and their mixtures in the bioluminescence inhibition assay with aliivibrio scheri</article-title>
          .
          <source>Environ. Sci.: Processes Impacts</source>
          <volume>19</volume>
          ,
          <issue>414</issue>
          {
          <fpage>428</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>European</given-names>
            <surname>Environment</surname>
          </string-name>
          <article-title>Agency: Linkages of species and habitat types to maes ecosystems (</article-title>
          <year>2015</year>
          ), https://www.eea.europa.eu
          <article-title>/data-and-maps/data/linkages-ofspecies-and-habitat</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shvaiko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Ontology Matching,
          <source>Second Edition</source>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Fay</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elonen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skopinski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pilli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , LaLone, C.:
          <article-title>Enhancing the Utility of the ECOTOX knowledgebase (ECOTOX KB) via ontologybased semantics mapping</article-title>
          .
          <source>In: SETAC Europe</source>
          , Rome, ITALY, May
          <volume>14</volume>
          - 18,
          <year>2018</year>
          . (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Feldman</surname>
            ,
            <given-names>H.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumontier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haider</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hogue</surname>
            ,
            <given-names>C.W.:</given-names>
          </string-name>
          <article-title>Co: A chemical ontology for identi cation of functional groups and semantic comparison of small molecules</article-title>
          .
          <source>FEBS Letters</source>
          <volume>579</volume>
          (
          <issue>21</issue>
          ),
          <volume>4685</volume>
          {
          <fpage>4691</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Forbes</surname>
            ,
            <given-names>V.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calow</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sibly</surname>
            ,
            <given-names>R.M.:</given-names>
          </string-name>
          <article-title>Are current species extrapolation models a good basis for ecological risk assessment?</article-title>
          <source>Environmental Toxicology and Chemistry</source>
          <volume>20</volume>
          (
          <issue>2</issue>
          ),
          <volume>442</volume>
          {
          <fpage>447</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benfenati</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Consensus qsar modeling of toxicity of pharmaceuticals to di erent aquatic organisms: Ranking and prioritization of the drugbank database compounds</article-title>
          .
          <source>Ecotoxicology and Environmental Safety</source>
          <volume>168</volume>
          ,
          <issue>287</issue>
          {
          <fpage>297</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valsecchi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasqualini</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baderna</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marzo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lombardo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benfenati</surname>
          </string-name>
          , E.:
          <article-title>Qsar modeling of daphnia magna and sh toxicities of biocides using 2d descriptors</article-title>
          .
          <source>Chemosphere</source>
          <volume>229</volume>
          ,
          <issue>8</issue>
          {
          <fpage>17</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Kipf</surname>
            ,
            <given-names>T.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Semi-supervised classi cation with graph convolutional networks</article-title>
          .
          <source>CoRR abs/1609</source>
          .02907 (
          <year>2016</year>
          ), http://arxiv.org/abs/1609.02907
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Laender</surname>
            ,
            <given-names>F.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morselli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baveco</surname>
          </string-name>
          , H., den Brink, P.V.,
          <string-name>
            <surname>Guardo</surname>
            ,
            <given-names>A.D.</given-names>
          </string-name>
          :
          <article-title>Theoretically exploring direct and indirect chemical e ects across ecological and exposure scenarios using mechanistic fate and e ects modelling</article-title>
          .
          <source>Environment International</source>
          <volume>74</volume>
          ,
          <volume>181</volume>
          {
          <fpage>190</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isele</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakob</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jentzsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontokostas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morsey</surname>
            , M., van Kleef,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>DBpedia - A largescale, multilingual knowledge base extracted from Wikipedia</article-title>
          .
          <source>Semantic Web</source>
          <volume>6</volume>
          (
          <issue>2</issue>
          ),
          <volume>167</volume>
          {
          <fpage>195</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>Y.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kofron</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Traphagen</surname>
            ,
            <given-names>L.M.:</given-names>
          </string-name>
          <article-title>Do structurally similar molecules have similar biological activity</article-title>
          ?
          <source>Journal of Medicinal Chemistry</source>
          <volume>45</volume>
          (
          <issue>19</issue>
          ),
          <volume>4350</volume>
          {4358 (Sep
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Myklebust</surname>
            ,
            <given-names>E.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimenez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rudjord</surname>
            ,
            <given-names>Z.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wolf</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tollefsen</surname>
            ,
            <given-names>K.E.</given-names>
          </string-name>
          :
          <article-title>Integrating semantic technologies in environmental risk assessment: A vision</article-title>
          .
          <source>In: 29th Annual Meeting of the Society of Environmental Toxicology and Chemistry (SETAC)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>NCBI-Taxonomy</surname>
          </string-name>
          :
          <article-title>The national center for biotechnology information (</article-title>
          <year>2019</year>
          ), https://www.ncbi.nlm.nih.gov/taxonomy
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Nickel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosasco</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poggio</surname>
            ,
            <given-names>T.A.</given-names>
          </string-name>
          :
          <article-title>Holographic embeddings of knowledge graphs</article-title>
          .
          <source>CoRR abs/1510</source>
          .04935 (
          <year>2015</year>
          ), http://arxiv.org/abs/1510.04935
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. PubChem:
          <article-title>National institutes of health (nih) (</article-title>
          <year>2019</year>
          ), https://pubchem.ncbi.nlm.nih.gov/
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. U.S. EPA:
          <article-title>Ecotoxicology knowledgebase (ecotox) (</article-title>
          <year>2019</year>
          ), https://cfpub.epa.gov/ecotox/
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Knowledge graph embedding: A survey of approaches and applications</article-title>
          .
          <source>IEEE Trans. Knowl. Data Eng</source>
          .
          <volume>29</volume>
          (
          <issue>12</issue>
          ),
          <volume>2724</volume>
          {
          <fpage>2743</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>tau Yih</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Embedding entities and relations for learning and inference in knowledge bases</article-title>
          .
          <source>CoRR abs/1412</source>
          .6575 (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>