<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Discovering Information from an Integrated Graph Database</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Erik M. van Mulligen</string-name>
          <email>e.vanmulligen@erasmusmc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wytze J. Vlietstra</string-name>
          <email>w.vlietstra@erasmusmc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rein Vos</string-name>
          <email>rein.vos@maastrichtuniversity.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan A. Kors</string-name>
          <email>j.kors@erasmusmc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Erasmus University Medical Center</institution>
          ,
          <addr-line>Rotterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Maastricht University</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The information explosion in science has become a different problem, not the sheer amount per se, but the multiplicity and heterogeneity of massive sets of data sources. Relations mined from these heterogeneous sources, namely texts, database records, and ontologies have been mapped to Resource Description Framework (RDF) triples in an integrated database. The subject and object resources are expressed as references to concepts in a biomedical ontology consisting of the Unified Medical Language System (UMLS), UniProt and EntrezGene and for the predicate resource to a predicate thesaurus. All RDF triples have been stored in a graph database, including provenance. For evaluation we used an actual formal PRISMA literature study identifying 61 cerebral spinal fluid biomarkers and 200 blood biomarkers for migraine. These biomarkers sets could be retrieved with weighted mean average precision values of 0.32 and 0.59, respectively, and can be used as a first reference for further refinements.</p>
      </abstract>
      <kwd-group>
        <kwd>knowledge based discovery</kwd>
        <kwd>graph databases</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Discovering new information from PubMed and from other biomedical databases is a time consuming and tedious
process [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Retrieving and combining information from these databases has to be performed manually and requires
an understanding of the different information models. In this paper we present a method to harmonize the
information from all these biomedical databases as RDF triples and integrate them within a graph database. To
investigate whether such a harmonized and integrated approach is beneficial we evaluated this against a formal literature
review, performed by a collaborative expert group in neurology research, that identified from (full text) literature
migraine biomarkers in cerebral spinal fluid and blood [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        Many have recognized the potential of computers to support the discovery process of new biomedical information.
The pioneer in this field, Swanson, recognized the potential of relating disconnected fields of knowledge in
biomedicine, in particular by discovering new associations between, as he called it, A and C terms, consisting of single
words or short phrases (2-3 words). He developed a program named ArrowSmith to automatically find B terms that
co-occur with A and C terms in Medline titles3. If the A and C terms were never co-mentioned in a title, a new
potential discovery was identified. Using this approach he was able to discover a connection between Raynaud’s
disease (A) and fish oil (C) through blood coagulation (B), and between migraine (A) and magnesium (C) via blood
clotting (B). These hypotheses were later on proven correct in experimental studies [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The value of this approach has been recognized by many scientists and a series of new research projects were
started to improve on this. One method, explored by Blake and Pratt, was to use concepts as defined by the UMLS
instead of separate terms [
        <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
        ]. In the UMLS thesaurus, different terms that denote the same unit of thought have
been normalized to a single concept. Weeber et al. were the first to mine concepts from both Medline titles as well
as abstracts, by mapping terms to the UMLS thesaurus with the MetaMap concept recognizer [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Weeber et al. were
also succesful in applying their system for a new discovery in drug research, suggesting thalidomide as a treatment
for chronic hepatitis C, among others [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Swanson manually selected the B terms that he considered most relevant for further exploration, but he did put
also much effort in bringing together very different datasets, covering different research fields in medicine. Many
researchers have followed up Swanson’s work and have worked on approaches to algorithmically select the B terms.
The concept-based approaches using UMLS have explored the use of the semantic types of the B concepts. Blake
and Pratt used this approach to discard several semantic types and reported an 81% decrease of the number of B
terms5. Srinivasan et al. applied a similar approach to filter out B terms based on semantic types [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. If the relevant
semantic types were precisely known, the set of terms could be reduced by as much as 91%; if only the obviously
irrelevant semantic types were removed, the number of terms was reduced by an average of 31%. Gordon and
Lindsay evaluated several ranking algorithms borrowed from the information retrieval field when they re-analyzed
Swanson’s fish oil-Raynaud’s Disease discovery, such as Term Frequency-Inverse Document Frequency (TF-IDF)
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. They reported reproduction of 10 of the 12 relevant B-terms for Swanson’s discovery in a list of 35 terms.
Torvik and Smalheiser applied an ensemble algorithm to rank the B-terms that combined eight weighted variables,
such as "B-term occurs in more than one paper within literature sets A and C", "B-term maps to at least one UMLS
semantic category", "B-term first appears recently within Medline as a whole", etc. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. While Swanson originally
used a fixed order approach of first filtering uninformative terms using a stop word list, subsequently term
categorization, and finally manual selection of B-terms, this ensemble algorithm contains all steps of Swanson’s fixed order
approach, but has the advantage of not losing potentially relevant B terms in any of the intermediate steps.
      </p>
      <p>
        Yetsigen-Yildiz et al. compared statistics to rank the B-terms [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Two of them were frequency-based, with the
TF-IDF and the association rules as tested by Hristovski [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] et al., and two were probability based; the Z-score,
which creates literature subsets, and the mutual information score. The association rules were not evaluated against
the Swanson sets, but they were analyzed on their predictions from a subset of Medline’s future published
discoveries.
      </p>
      <p>
        Hristovski et al. were the first to test the added value of incorporating relation predicates into a literature-based
discovery process [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. They applied the UMLS semantic network and the SemRep text mining system to identify
relationships between terms [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Predicates were used to identify discovery patterns: specific combinations of two
predicates between three terms, which when combined would constitute a functional, biologically relevant
association. Although the inclusion of predicates was considered to offer clear advantages, the lack of accuracy of the
relationship extraction hampered practical application.
      </p>
      <p>
        With the ANNI discovery system the co-occurrences between a concept and other concepts in all Medline
abstracts were computed and stored in a so-called concept profile [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The strength of a relationship between two
concepts is expressed as a matching score between their concept profiles. Concepts can be grouped based on their
semantic type and their concept profiles can be matched based on various algorithms: mutual information measure,
log-likelihood, and dot product [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The matching strategy takes into account all the B concepts contained in the
concept profile, filters the resulting C concepts on the required semantic type(s) and ranks the result on matching
score. This approach has been used by Jelier et al. in a study to match the concept profiles for genes from DNA
microarray data with concepts that denote gene functions [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The same approach has been used by Van Haagen et al.
to predict protein-protein interactions by computing the matching score between protein concept profiles at certain
time intervals in Medline [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. An extension of this approach has been developed by using ANNI in mapping
disease-disease relationships for knowledge discovery in multi-morbidity research on somatic and psychiatric diseases
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        Previous work mainly focused on discovering direct relations based on mainly one source of information
(literature). In this paper we describe a novel approach to extend discoveries of associations between a source and target
concept to associations that involve a series of intermediate concepts (paths) combining information mined from
literature, conventional biological databases, semantically enriched information (RDF) sources, and biomedical
ontologies and thesauri. The formalization of this information from different heterogeneous sources into RDF triples
and the integration of these triples in a graph database seems to be logical next step to support information discovery
tasks with multiple intermediate nodes and offers more possibilities to rank the various discovered connections using
graph statistics. We evaluated this approach of using a graph database based on heterogeneous sources for
information discovery by comparing discovered associations with the results of an actual formal literature review.
Our approach semantically integrates triples extracted from Medline abstracts as provided in Semantic Medline [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
with relations obtained from the UMLS 2012AA and databases such as UniProt, EntrezGene, Comparative
Toxicogenomics Database, and RDF triples from the datasets contained in Linked Open Drug Data [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] (DrugBank,
DailyMed, and SIDER) into a graph database. Furthermore, the relations between The subject and object resources
and the predicates in the Semantic Medline triples were already expressed in terms of our ontology and predicate
thesaurus. For UniProt, EntrezGene and the Comparative Toxicogenomics Database the process of making triples
included the mapping of the implicit relations of the database schema to explicit predicates and the mapping of the
subject and object to a RDF resource, i.c. a UMLS, UniProt or EntrezGene identifier. For each UMLS concept in
our ontology we have all the different identifiers and all terms used to refer to the concept. Mapping the information
of a database record to a concept in our ontology was obtained either by matching it to one of its identifiers or by
matching it to one of its terms. The term matching was performed by applying our Peregrine [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] text mining
pipeline.
      </p>
      <p>
        From these different sources we identified 2,669,792 individual concepts, together with about 71 million relations
between them. The relations are based on the relationships defined in the UMLS Semantic Network, the
relationships defined in the UMLS MetaThesaurus (MRREL table), and the predicates defined by Halil and used within
Semantic Medline [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. We harmonized the set by looking at trivial synonyms in this set and mapped these to 171
different predicate identifiers. Each subject and object from a RDF triple are related with an "isa" relation to one or
more semantic types as defined in UMLS. Semantic types on their turn are linked with an "isa" relation to a
semantic group.
      </p>
      <p>The resulting semantically mapped relations, commonly referred to as triples, have been stored in a graph
database. The graph database has been implemented in the Neo4J graph database, version 1.8.324. In order to add the
triples with their provenance we developed an import program that uses the Neo4J using the java API of Neo4J. We
implemented a REST API on top of Neo4J that implements the notion of concepts, labels, semantic types, semantic
groups, and provenance. Each RDF triple is represented in the graph database by making a labeled node using the
preferred term of the ontology concept for both subject and object and the predicate name as labeled edge to
represent the relation between subject and object node. A subject and object node can be linked with multiple semantic
predicate labels and the provenance information implemented as a reference to a text or a database record as the
source of the relation can be added to each edge. Semantic predicates contain a direction and for both directions
labels are provided, typically the active and passive form of a verb. Neo4J has built a path finder algorithm that find
paths between nodes in the graph. We extended this functionality with the use of provenance information in scoring
the various paths. An example of the mapping of the database of UniProt is provided in Table 1.
gene_product_plays_role_in
_biological_process
GO - Molecular
function</p>
      <sec id="sec-2-1">
        <title>GO - Molecular function</title>
      </sec>
      <sec id="sec-2-2">
        <title>GO - Biological process GO - Biological process</title>
        <p>http://www.ncbi.nlm.nih.gov
/gene/7529
https://utsws.nlm.nih.gov/rest/content
/current/CUI/C1149286
https://utsws.nlm.nih.gov/rest/content
/current/CUI/C1323310
https://utsws.nlm.nih.gov/rest/content
/current/CUI/C1155556
https://utsws.nlm.nih.gov/rest/content</p>
      </sec>
      <sec id="sec-2-3">
        <title>Keywords - Coding sequence diversity</title>
      </sec>
      <sec id="sec-2-4">
        <title>Organism</title>
        <p>receptor
signaling pathway
Host-virus
interaction</p>
      </sec>
      <sec id="sec-2-5">
        <title>Cytoplasm perinuclear region of cytoplasm</title>
        <p>Polymorphism
Homo sapiens
(Human)</p>
        <p>
          The challenge of integrating UniProt entries lies in mapping the annotation fields to the corresponding ontology
concepts. We used our concept identification pipeline Peregrine to identify concepts in the free text UniProt
annotation fields [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. The mapping of the implicit relations defined in the UniProt schema to the proper semantic
predicates is a one-time manually effort and requires understanding of the biological meaning of the data. This mapping
process has been performed for all integrated databases. Once created a mapping can be applied to each update of
the database.
        </p>
        <sec id="sec-2-5-1">
          <title>Discovering connections</title>
          <p>Around Neo4J’s basic functionality we provided a web service that implements functionality necessary for our
discovery task. In particular, for inferencing we implemented a path-finding algorithm that extends Neo4J’s
functionality. This simple, path-finding type of inferencing is not following the main, logic-based inferencing approaches such
as implemented with OWL-DL and formal reasoners. The extension of Neo4J’s path-finding function allows one to
specify a set of semantic predicates that restricts the set of triples that can be explored to find a path between the
source and target concepts. The paths lengths are currently limited to a maximum of five triples to avoid
computational explosion. The path function can be modified and can take into account additional information and graph
statistics that may influence the selection of triples, e.g., the amount of provenance information (the sources that
support the relation), the variety of databases that support the triple.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>The graph database has been used in a number of application domains. To evaluate the use of this graph database
with triples obtained from PubMed abstracts and biomedical databases we evaluated the identification of biomarker
compounds marking the imminence of a migraine attack with those reported in a literature review study. We
obtained the set of 61 compounds that have been reported to be measurable in the cerebral spinal fluid, and a set of 200
compounds reported to be measurable in the serum of migraine patients. The latter set was obtained at the same
literature review study but not yet submitted for publication. Both sets were manually constructed by a manual
review process of a corpus of articles retrieved with PubMed, EMBASE, and Web of Science. The objective of our
study was to test whether a graph database could be used to identify a set of linking concepts, similar to the linking
B-terms, between these compounds and migraine. The question was whether this set of linking concepts with their
interconnectivity could be used to identify (1) the original set of compounds, and (2) new compounds of interest.</p>
      <p>Biologically Active Substance
Neuroreactive Substance or Biogenic Amine
Hormone
Enzyme
Vitamin</p>
      <sec id="sec-3-1">
        <title>Receptor Immunologic Factor</title>
      </sec>
      <sec id="sec-3-2">
        <title>Tissue Cell Cell Component Gene or Genome</title>
        <p>Disease or Syndrome
Mental or Behavioral Dysfunction
Body Part, Organ, or Organ Component</p>
      </sec>
      <sec id="sec-3-3">
        <title>Chemical Viewed Structurally</title>
      </sec>
      <sec id="sec-3-4">
        <title>Organic Chemical</title>
      </sec>
      <sec id="sec-3-5">
        <title>Nucleic Acid, Nucleoside, or Nucleotide</title>
      </sec>
      <sec id="sec-3-6">
        <title>Organophosphorus Compound</title>
      </sec>
      <sec id="sec-3-7">
        <title>Amino Acid, Peptide, or Protein</title>
      </sec>
      <sec id="sec-3-8">
        <title>Carbohydrate</title>
      </sec>
      <sec id="sec-3-9">
        <title>Lipid</title>
      </sec>
      <sec id="sec-3-10">
        <title>Steroid</title>
      </sec>
      <sec id="sec-3-11">
        <title>Eicosanoid</title>
      </sec>
      <sec id="sec-3-12">
        <title>Element, Ion, or Isotope</title>
      </sec>
      <sec id="sec-3-13">
        <title>Physiologic Function</title>
      </sec>
      <sec id="sec-3-14">
        <title>Organism Function</title>
      </sec>
      <sec id="sec-3-15">
        <title>Organ or Tissue Function</title>
      </sec>
      <sec id="sec-3-16">
        <title>Cell Function</title>
      </sec>
      <sec id="sec-3-17">
        <title>Molecular Function</title>
        <p>The two sets of compounds were fed to the graph database to obtain the paths between these compounds and
migraine. These paths were analyzed for characteristics (number of publications, range of publication dates, path
length, etc.). Additional compounds that were not part of the initial set have been viewed as potentially new
discovered compounds.</p>
        <p>The final result of this study was a set of concepts found in the paths linking migraine to these sets of compounds.
A selection of this set of linking B-concepts was made on basis of a subset of the semantic types. The Signs and
Symptoms semantic type was excluded based on discussion with the migraine researchers. Furthermore,
Pharmacological Substances and Antibiotics, and concepts which were both a Pharmacological Substance, Organic Chemical,
or Steroids, or Nucleic acids and amino acids, as well as Antibiotics were explicitly excluded, because the migraine
researchers were only interested in endogenous compounds associated with migraine, and not in chemotherapeutic
treatments. The final list of semantic types is shown in Table 2. Using this selected B-concept set we used the
number of different connections between a compound and the B-concept set for reconstructing the initial given set of
compounds and secondly to identify potential new compounds. Several ranking statistics were evaluated and overall
there was only very little difference. From the cerebral spinal fluid set of 61 compounds directly connected to
migraine one could not be identified and from the serum set of 200 compounds directly connected to migraine 23 could
not be identified using this approach. A weighted mean average precision of 0.32 was computed for the cerebral
spinal fluid set and 0.59 for the blood set. Or stated differently, 78% of the unique reference compounds (222
compounds) were found in the top 4% (=1500) results, which means that about one out of ten results was a reference
compound.</p>
        <sec id="sec-3-17-1">
          <title>Error Analysis</title>
          <p>We performed an error analysis on our reference set, by examining compounds from the top 100 of our results that
are not included in our set of reference compounds. The results of this analysis are shown in Table 3. From the 24
compounds that were not in the list, 13 were excluded from the list because they were categorized as inorganic
chemical or as a pharmaceutical preparation and excluded initially from the analysis and 11 were only remotely
connected to migraine and therefore excluded from the result list. From the 49 compounds low on the list 6 were
ranked low because of the ranking mechanism, 28 compounds were connected but due to many connections outside
the migraine cluster ranked low, and 15 compounds were only connected via an ontological relationship (a "isa"
relationship with a class) and where therefore ranked low.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>As mentioned in the introduction, the ranking and filtering of the B-terms determines to a large extent the success of
the knowledge discovery method. A similar issue can be raised about the ranking and relevance of the connecting
paths that our method constructs in a multi-source graph database. With increasing path lengths, at some point each
pair of concepts in the graph database will be connected. It will therefore be important to investigate approaches that
can differentiate between useful and sound discovery paths and those that are noisy and redundant. The platform is
powerful in its potential to implement discovery patterns that combine a rich feature set consisting of semantic
types, semantic groups, semantic predicates, connectivity, and amount of provenance stemming from different
sources.</p>
      <p>From our experiments and our interactions with the expert group thus far it became clear that adding more
ontological grounding to the semantic predicates would be helpful. Similar to semantic types and groups, which denote
the specific properties of concepts, we can imagine that representation of the predicates by concepts with references
to specific types of predicates - e.g., transitive, intransitive, causative, factitive, etc. Predicate types would impose
specific properties to the predicates, such as transitive inference, that could be relevant for the discovery process.</p>
      <p>For this application we did not restrict the discovery connection paths on basis of the combination of a particular
semantic groups or types of concepts with a set of particular predicates. Our first experience is that such a selection
might help in finding more relevant connections. The flexibility of the graph database to support various types of
selections has been used in a different application in the field of adverse drug reactions and in food safety. We will
further investigate in how far these selections are depending on an application and can be formalized in a guideline
on how to use a graph database for discovery.</p>
      <p>The compounds in the top of the result list that were not part of the reference set have not been analyzed yet, but a
quick scan learned that there could be potential interesting compounds that are worth further investigation. This
approach can also potentially "use" the high connectivity of a compound with a reference set to discover new
potential interesting compounds. The verification of this is, however, a tedious and costly process.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The graph database that we constructed combines information extracted from biomedical texts with information
obtained from biological databases. We have shown in this paper that relations from texts and structured databases
can be effectively combined in a single graph database. However, this approach is a first step to use large integrated
datasets to support the discovery process. Research will be required to better understand the importance of graph
statistics for the discovery process. What we present in this paper is a first step that can be used as a reference for
further work.</p>
      <p>The information discovery approach illustrated in this paper shows that relevant compounds as identified in an
actual formal literature review can be retrieved with a fairly high recall. Furthermore, our approach shows that the
connectivity to a set of other concepts has potential. The flexibility of the graph database enables the application of
the approach to other discovery applications and evaluate other approaches to combine graph statistics and filters on
semantic groups and predicates.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Lu Z.:
          <article-title>Pubmed and beyond: a survey of web tools for searching biomedical literature</article-title>
          .
          <source>Database</source>
          , Oxford (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>van Dongen</surname>
            <given-names>R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zielman</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noga</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dekkers</surname>
            <given-names>O.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hankemeier</surname>
            <given-names>T</given-names>
          </string-name>
          .,
          <string-name>
            <surname>van den Maagdenberg</surname>
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Terwindt</surname>
            <given-names>G.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrari M.D.</surname>
          </string-name>
          :
          <article-title>Migraine biomarkers in cerebrospinal fluid: A systematic review and meta-analysis</article-title>
          .
          <source>Cephalalgia</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Swanson</surname>
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smalheiser</surname>
            <given-names>N.R.:</given-names>
          </string-name>
          <article-title>An interactive system for finding complementary literatures: a stimulus to scientific discovery</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>91</volume>
          ,
          <fpage>183</fpage>
          -
          <lpage>203</lpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Swanson</surname>
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Medical literature as potential source of new knowledge</article-title>
          .
          <source>Bull. Med</source>
          . Libr. Assoc.
          <volume>78</volume>
          ,
          <issue>1</issue>
          ,
          <fpage>29</fpage>
          -
          <lpage>37</lpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Blake</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pratt</surname>
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Automatically identifying candidate treatments from existing medical literature</article-title>
          .
          <source>AAAI Spring Symposium on Mining Answers from Texts and Knowledge Bases</source>
          .
          <fpage>9</fpage>
          -
          <lpage>13</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bodenreider</surname>
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The unified medical language system (UMLS): Integrating biomedical terminology</article-title>
          .
          <source>Nucleic Acids Res</source>
          .,
          <volume>32</volume>
          (Database issue),
          <fpage>D267</fpage>
          -
          <lpage>D270</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Weeber</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            <given-names>H</given-names>
          </string-name>
          ., de Jong-van den Berg L.T.W.,
          <string-name>
            <surname>Vos</surname>
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Using concepts in literature-based discovery: Simulating Swansons Raynaud-Fish Oil and Migraine-Magnesium Discoveries</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology 52</source>
          ,
          <issue>7</issue>
          ,
          <fpage>548</fpage>
          -
          <lpage>557</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Weeber</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vos</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronson</surname>
            <given-names>A.R.</given-names>
          </string-name>
          , Molema G.:
          <article-title>Generating hypotheses by discovering implicit associations in the literature: a case report of a search for new potential therapeutic uses for thalidomide</article-title>
          .
          <source>J. Am. Med</source>
          . Inform. Assoc.
          <volume>10</volume>
          ,
          <issue>3</issue>
          ,
          <fpage>252</fpage>
          -
          <lpage>259</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Srinivasan</surname>
            <given-names>P</given-names>
          </string-name>
          .
          <article-title>Text mining: generating hypotheses from Medline</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology 55</source>
          ,
          <issue>5</issue>
          ,
          <fpage>396</fpage>
          -
          <lpage>413</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lindsay</surname>
            <given-names>R.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordon M.D.</surname>
          </string-name>
          :
          <article-title>Literature-based discovery by lexical statistics</article-title>
          . (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Torvik</surname>
            <given-names>V.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smalheiser</surname>
            <given-names>N.R.:</given-names>
          </string-name>
          <article-title>A quantitative model for linking two disparate sets of articles in Medline</article-title>
          .
          <source>Bioinformatics</source>
          <volume>23</volume>
          ,
          <issue>13</issue>
          ,
          <fpage>1658</fpage>
          -
          <lpage>1665</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Yetisgen-Yildiz</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pratt</surname>
            <given-names>W.:</given-names>
          </string-name>
          <article-title>A new evaluation methodology for literature based discovery systems</article-title>
          .
          <source>J. Biomed. Inform</source>
          .
          <volume>42</volume>
          ,
          <issue>4</issue>
          ,
          <fpage>633</fpage>
          -
          <lpage>643</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hristovski</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stae</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peterlin</surname>
            <given-names>B.</given-names>
          </string-name>
          , Dzeroski S.:
          <article-title>Supporting Discovery in Medicine by Association Rule Mining in Medline and UMLS</article-title>
          .
          <source>Medinfo</source>
          <volume>10</volume>
          (
          <issue>Pt2</issue>
          ),
          <fpage>1344</fpage>
          -
          <lpage>1348</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hristovski</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedman</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rindflesch</surname>
            <given-names>T.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peterlin</surname>
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Exploiting semantic relations for literature-based discovery</article-title>
          .
          <source>AMIA annual symposium proceedings. 349</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ahlers</surname>
            <given-names>C.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fiszman</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lang</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rindflesh</surname>
            <given-names>T.C.</given-names>
          </string-name>
          :
          <article-title>Extracting semantic predication from MEDLINE citations for pharmacogenomics</article-title>
          .
          <source>In: Pacific Symposium on Biocomputing 209-220</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Jelier</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuemie</surname>
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veldhoven</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dorssers</surname>
            <given-names>L.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenster</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kors</surname>
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Anni 2.0: a multipurpose text-mining tool for the life sciences</article-title>
          .
          <source>Genome Biol</source>
          .
          <volume>9</volume>
          ,
          <issue>6</issue>
          ,
          <issue>R96</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Jelier</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuemie</surname>
            <given-names>M.J.</given-names>
          </string-name>
          , Roes P.J.,
          <string-name>
            <surname>van Mulligen</surname>
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kors</surname>
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Literature-based concept profiles for gene annotation: The issue of weighting</article-title>
          .
          <source>Int. J. of Med</source>
          . Inform.
          <volume>77</volume>
          ,
          <issue>5</issue>
          ,
          <fpage>354</fpage>
          -
          <lpage>362</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Jelier</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenster</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dorssers</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wouter</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hendriksen</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mons</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delwel</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kors</surname>
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Text-derived concept profiles support assessment of DNA microarray data for acute myeloid leukemia and for androgen receptor stimulation</article-title>
          .
          <source>BMC Bioinformatics 8</source>
          ,
          <issue>14</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>van Haagen</surname>
            <given-names>H.H.</given-names>
          </string-name>
          , 't Hoen P.A., de Morrée A., van
          <string-name>
            <surname>Roon-Mom</surname>
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            <given-names>D.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roos</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mons</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van Ommen</surname>
            <given-names>G.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuemie</surname>
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>In silico discovery and experimental validation of new protein-protein interactions</article-title>
          .
          <source>Proteomics</source>
          <volume>11</volume>
          ,
          <issue>5</issue>
          ,
          <fpage>843</fpage>
          -
          <lpage>853</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Vos</surname>
            <given-names>R.</given-names>
          </string-name>
          , Aarts S.,
          <string-name>
            <surname>van Mulligen</surname>
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Metsemakers</surname>
            <given-names>J</given-names>
          </string-name>
          .,
          <string-name>
            <surname>van Boxtel</surname>
            <given-names>M.P.</given-names>
          </string-name>
          , Verhey F.,
          <string-name>
            <surname>van den Akker M.J.:</surname>
          </string-name>
          <article-title>Finding potentially new multimorbidity patterns of psychiatric and somatic diseases: exploring the use of literature-based discovery in primary care research</article-title>
          .
          <source>J. Am. Med</source>
          . Inform. Assoc.
          <volume>21</volume>
          ,
          <issue>1</issue>
          ,
          <fpage>139</fpage>
          -
          <lpage>145</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Kilicoglu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shin</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fiszman</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosemblat</surname>
            <given-names>G.</given-names>
          </string-name>
          , Rindflesch T.C.:
          <article-title>SemMedDB: a PubMed-scale repository of biomedical semantic predications</article-title>
          .
          <source>Bioinformatics</source>
          <volume>28</volume>
          ,
          <issue>23</issue>
          ,
          <fpage>3158</fpage>
          -
          <lpage>3160</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Linking</surname>
          </string-name>
          <article-title>Open Drug Data (LODD)</article-title>
          , http://www.w3.org/wiki/HCLSIG/LODD
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Peregrine</surname>
          </string-name>
          , https://trac.nbic.nl/data-mining/
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24. Neo4J Developers:
          <fpage>Neo4J</fpage>
          ,
          <source>Graph NoSQL Database</source>
          <year>2012</year>
          , http://neo4j.org/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>