<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The knowledge-driven exploration of integrated biomedical knowledge sources facilitates the generation of new hypotheses</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vinh Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olivier Bodenreider</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Todd Minning</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amit Sheth</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Tropical and Emerging Global Diseases</institution>
          ,
          <addr-line>Univeristy of Georgia</addr-line>
          ,
          <country country="GE">Georgia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Kno.e.sis Center, Wright State University</institution>
          ,
          <addr-line>Dayton, Ohio</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>National Library of Medicine</institution>
          ,
          <addr-line>Bethesda, Maryland</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Knowledge gained from the scienti c literature can complement newly obtained experimental data in helping researchers understand the pathological processes underlying diseases. However, unless the scienti c literature and experimental data are semantically integrated, it is generally di cult for scientists to exploit the two sources e ectively. We argue that, in addition to the semantic integration of heterogeneous knowledge sources, the usability of the integrated resource by scientists is dependent upon the availability of knowledge visualization and exploration tools. Moreover, the integration techniques must be scalable and the exploration interfaces must be easy to use by bench scientists. The end goal of such integrated knowledge sources and exploration tools is to enable scientists to generate novel hypotheses from the knowledge they explore. We tested the feasibility of our approach on a real use case in the domain of human health and parasite biology. On the one hand, we integrated the experimental data generated as part of an ongoing research on Chagas disease with the knowledge extracted from the PubMed articles, using Semantic Web technologies. On the other hand, we developed iExplore, a web tool with a graphical interface for interactive knowledge exploration, that allows non-technical users to explore the integrated knowledge base using a relationship-focused approach. We illustrate the e ectiveness of our approach by describing the knowledgedriven process of using iExplore to generate a new hypothesis for the treatment of Chagas disease.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Translating the knowledge from basic research into practice has been an
important trend in biomedical research in recent years. Translational research aims to
improve health by utilizing a wide range of biomedical resources, using
knowledge from experimental data at the point of care and guiding basic research with
problems encountered in patients. A large amount of biomedical knowledge is
available in the biomedical literature, e.g., in PubMed1 articles, and in structured
1 http://www.ncbi.nlm.nih.gov/pubmed
knowledge sources, including the Uni ed Medical Language System (UMLS) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
and the Entrez Gene database. Recent research [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ] has shown the potential
of using text mining on PubMed articles to advance the biomedical research by
exploiting the associations of genes and diseases from the text to prioritize the
candidate genes. However, we believe that these approaches do not exploit their
potential fully, because their context is limited by design to a speci c subset of
the biomedical literature.
      </p>
      <p>
        We argue that a broad knowledge base is key to enabling the generation of new
hypotheses. Using a large subset of the scienti c literature will bene t
biomedical scientists performing basic or clinical research, and facilitate the
integration of their data with the biomedical literature. We anchored our knowledge
base in comprehensive resources, such as the UMLS [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Entrez Gene. The
UMLS Metathesaurus enables the interoperability among data sources by
providing reference identi ers for biomedical entities. Moreover, the UMLS Semantic
Network provides a consistent categorization of the UMLS Metathesaurus
concepts, together with a set of relations among the categories. The Entrez Gene
database contains the list of genes for various model organisms, as well as their
annotations. The combination of UMLS and Entrez Gene in the schema of the
Biomedical Knowledge Repository (BKR) provides the exibility and
scalability for integrating additional data sources. We illustrate the integration of the
BKR with the experimental data obtained from the research in Chagas disease
in section 2.
      </p>
      <p>We believe that the knowledge exploration process should be supported by a
tool that the scientists nd easy to use and e ective to help them gain new
insights from the integrated knowledge bases of text and experimental data.
The visualization should display su cient contextual information to the users,
and the navigation should be driven by the intuition and background knowledge
of the scientists. We developed iExplore, a web tool that displays the graph
from integrated knowledge bases in an interactive manner using a
relationshipcentric approach. The tool complements to the function of existing tools, e.g.,
RelFinder2. While iExplore supports the exploration process, it is not supposed
to replace the role of the scientists in this process. We explain the knowledge
exploration process enabled by this tool in section 3.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Integration of knowledge sources</title>
      <sec id="sec-2-1">
        <title>The Biomedical Knowledge Repository</title>
        <p>
          The Biomedical Knowledge Repository (BKR) aims to integrate knowledge from
a variety of sources ranging from the scienti c literature to various structured
knowledge bases [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The BKR contains relations extracted from PubMed
documents by SemRep[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] and normalizes biological entities to concepts in the UMLS
and Entrez Genes. It includes approximately 20 million semantic predications
2 http://www.visualdataweb.org/rel nder.php
extracted from 6 million articles in PubMed published from 1999 to 2009. These
semantic predications are transformed into RDF format together with the
provenance information about the article where the predication is extracted.
A semantic predication extracted from the title or abstract of an article is
represented as a set of RDF triples. For example, the title \Trypanosoma cruzi
calreticulin: a possible role in Chagas' disease autoimmunity" of the article with
PubMed ID 19108895 contains one predication, \CALR associated with
Chagas Disease". The set of triples in Table 1 is created to represent this semantic
predication in the BKR.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Experimental Data about Chagas Disease</title>
        <p>The experimental data about Chagas Disease originate from DNA microarray
analysis, proteome analysis, gene knockout and strain creation protocols. DNA
microarray analysis was used to measure the relative transcript abundances for
all of the genes in the T. cruzi genome during the four main life cycle stages of
T. cruzi, namely amastigote, trypomastigote, epimastigote and metacylic
trypomastigote. Whole genome shot-gun proteomic analysis was used to measure
the presence of proteins encoded by T. cruzi genes during the four life cycle
stages. The proteome and transcriptome data have been used to prioritize genes
for the gene knockout and strain creation protocols. To capture the detail of
these experimental protocols, we created two ontolologies: Parasite Experiment
and Parasite Lifecycle, that have been published in the BioPortal3. We use these
ontologies as schema to convert the experimental data into RDF.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Integration</title>
        <p>A uni ed view of experimental data with biomedical literature requires the
mappings between entities of two data sources. Genes are one of the
potential common entities between the BKR and experimental data sources.
However, the experimental data are based on the Trypanosoma cruzi (T. cruzi)
3 http://bioportal.bioontology.org/
genome, while predications extracted from the biomedical literature refer to
human genes. The study of the human orthologs of T. cruzi genes is helpful
because the gene function is usually conserved across species. To bridge the gap
between genes of two organisms, we use the orthologous mapping from T. cruzi
to Homo sapiens (human) from the KEGG Sequence Similarity database4, and
create an RDF triple for each pair of orthologous genes. For example,
orthology between the human gene \CALR" (ID 811 in Entrez Gene) and the T.
cruzi gene Tc00.1047053509011.40 is represented by the following RDF triple:
\EG 811 is orthologous Tc00.1047053509011.40." In practice, such orthology
relations connect entities across the two sources.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Knowledge exploration</title>
      <p>
        Knowledge exploration is the process of establishing a new relationship between
two known concepts and generalizes the notion of literature-based discovery
to sources others than the biomedical literature [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The exploration process
includes two steps, navigation and interpretation.
      </p>
      <p>Navigation requires the visualization of knowledge as a set of relations. Such
relations can be explored by the user based on domain knowledge. We
developed iExplore5 with an intuitive graphical interface to enable navigation. The
tool creates an abstraction layer from the RDF integrated knowledge bases to
visualize the graph of concepts and named relationships. Because knowledge is
represented as a set of relations, users can expand and narrow the graph based
on their interest. Graph expansion is implemented through prede ned SPARQL
queries hidden from users. In practice, biologists drive the exploration guided by
their background knowledge, selecting concepts to expand the graph or
restricting the graph to speci c relations. Interpretation allows biologists to utilize
their background knowledge to generate novel hypotheses from the chains of
concepts identi ed in the navigation phase.</p>
      <p>Example The exploration starts by expanding the concept \Chagas Disease"
and inspecting its related concepts. Of particular interest to us are known
treatments for Chagas disease. We restrict the graph to the \TREATS" relation
using a lter. Among the treatment concepts, we focus on drugs (categorized by
semantic types, such as \Pharmacologic Substance"). We nd the drug
itraconazole, known for treating various parasitic diseases and which has side e ects. We
pursue our exploration by expanding the graph with the relations of the concept
\Itraconazole". Speci cally, we want to explore the genes connected to
\Itraconazole" via the \INHIBITS" relation because they possibly indicate biological
pathways involved in the treatment of Chagas disease. Since only human genes
are present in the graph extracted from the literature, we follow the orthology
relation in order to nd the T. cruzi orthologs of these human genes, i.e., the
4 http://www.genome.jp/kegg/ssdb/
5 Due to limited space, the tool and illustrated examples in section 3 are presented in
the tool's homepage http://knoesis.wright.edu/iExplore/index.html for review.
possible target of the drug in the parasite. In summary, this example establishes
a chain of named relationships from \Chagas Disease" to itraconazole and
human genes to T. cruzi genes. We generate a hypothesis from this chain, that
the T. cruzi orthologs of human genes inhibited by itraconazole may also be
inhibited by itraconazole and thus would be candidates for further studies into
the mechanism(s) of action of the drug itraconazole on T. cruzi.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>We presented the semantic approach that underpins the integration and
complementary use of the two independently generated, heterogeneous knowledge
sources. These integrated knowledge bases together with iExplore allow
biologists to use their knowledge to drive the exploration and generate new biomedical
hypotheses. The validation of such hypotheses is part of our future work. We also
plan to improve iExplore by learning the way biologists use their background
knowledge to generate hypothesis, and then automate the interpretation step
by making recommendations for hypothesis generation. Of note, our approach
can easily be generalized to other diseases and, more generally, to other data
sources integrated and explored together with the BKR following the approach
we demonstrated with Chagas disease. The two-step exploration process can also
be applied to make use of the broader integration of these independently
generated knowledge sources.</p>
      <p>Acknowledgements This research was supported by an appointment to the
Research Participation Program at National Library of Medicine, and the NIH R01
Grant number 1R01HL087795-01A1. We also acknowledge Dr. Thomas
Rindesch, Dr. Cartic Ramakrishnan, Dr. Priti Parikh, Jonathan Mortensen, Joshua
Dotson and Sarasi Lalithsena for help.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>C.</given-names>
            <surname>Ahlers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fiszman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lang</surname>
          </string-name>
          , and
          <string-name>
            <surname>T.</surname>
          </string-name>
          <article-title>Rind esch. Extracting semantic predications from medline citations for pharmacogenomics</article-title>
          .
          <source>In Paci c Symposium on Biocomputing</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>O.</given-names>
            <surname>Bodenreider</surname>
          </string-name>
          .
          <article-title>The uni ed medical language system (umls): integrating biomedical terminology</article-title>
          .
          <source>Nucleic acids research</source>
          ,
          <volume>32</volume>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>O.</given-names>
            <surname>Bodenreider</surname>
          </string-name>
          and
          <string-name>
            <surname>T.</surname>
          </string-name>
          <article-title>Rind esch. Advanced library services: Developing a biomedical knowledge repository to support advanced information management applications</article-title>
          .
          <source>National Library of Medicine</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A.</given-names>
            <surname>Faro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giordano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Spampinato</surname>
          </string-name>
          .
          <article-title>Combining literature text mining with microarray data: advances for system biology modeling</article-title>
          .
          <source>Brie ngs in Bioinformatics</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>D.</given-names>
            <surname>Swanson</surname>
          </string-name>
          .
          <article-title>Fish oil, raynaud's syndrome, and undiscovered public knowledge</article-title>
          .
          <source>Perspectives in biology and medicine</source>
          ,
          <volume>30</volume>
          (
          <issue>1</issue>
          ):
          <fpage>7</fpage>
          ,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>N. Ti n</given-names>
            , J.
            <surname>Kelso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Powell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bajic</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Hide</surname>
          </string-name>
          .
          <article-title>Integration of text-and data-mining using ontologies successfully selects disease gene candidates</article-title>
          .
          <source>Nucleic acids research</source>
          ,
          <volume>33</volume>
          (
          <issue>5</issue>
          ):
          <fpage>1544</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>