<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Smart Searching System for Biomedical Information</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hong-Woo Chun</string-name>
          <email>hw.chun@kisti.re.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chang-Hoo Jeong</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sa-Kwang Song</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yun-Soo Choi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sung-Pil Choi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hanmin Jung</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Korea Institute of Science and Technology Information</institution>
          ,
          <addr-line>245 Daehangno, Yuseong-gu, Daejeon, 305-806</addr-line>
          ,
          <country country="KR">South Korea</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Interactions between Biomedical entities provides meaningful information to detect and invent new drugs for diseases. Natural Language Processing-based Biomedical interaction extraction approach have shown encouraging results in the previous studies. While interaction extraction research has been a popular topic, research about searching and browsing methods for the extracted information has not been an attractive topic relatively. This demonstration presents a smart searching system that provides various analysis tools for Biomedical interactions in whole PubMed. We expect that researchers can discover and develop new research outcomes through the proposed searching system.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 INTRODUCTION</title>
      <p>
        Many automatic Information Extraction (IE) approaches using
Natural Language Processing (NLP) and Text Mining technologies
have been proposed to extract automatically meaningful information
in Biomedical domain. Biomedical entity recognition
        <xref ref-type="bibr" rid="ref1">(Song et al.,
2011)</xref>
        and relation extraction research with respect to
diseasegene association (Chun et al., 2004) and protein-protein interaction
        <xref ref-type="bibr" rid="ref3">(Chun et al., 2011)</xref>
        are examples of the IE research in Biomedical
domain. The IE system recognizes and extracts knowledge from a
massive literature and the extracted knowledge is accumulated in a
knowledge base.
      </p>
      <p>In order to decide research topics, not only IE techniques but
also effective searching methods for the extracted information are
very important. While information extraction research has been one
of the favorite topics, research about searching methods for the
extracted information has not been an attractive topic relatively.
In other words, it is overlooked even though various specialized
searching and browsing methods are necessary to express the
extracted information appropriately.</p>
      <p>The demonstration will show a smart searching system for
Biomedical entities and their interactions. Four types of searching
services are included: Smart slide, Semantic network browsing,
Top5, and Find it.</p>
    </sec>
    <sec id="sec-2">
      <title>BIOMEDICAL INFORMATION IN PUBMED</title>
      <p>Biomedical information in PubMed contains semantic triples
extracted from 21 million PubMed abstracts. A semantic triple
consists of two Biomedical entities and one verb, and two
Biomedical entities are syntacticly a subject and an object for a verb
appeared in a sentence. To extract the semantic triples, various NLP
techniques are applied as the following two steps:</p>
      <p>
        To recognize entities, a machine learning-based named
entity recognizer
        <xref ref-type="bibr" rid="ref4">(Yoshida et al., 2004)</xref>
        is used. The target
concepts in the named entity recognition are the following
six: genes/proteins, diseases, enzymes, drugs, symptoms and
chemical compounds.
      </p>
      <p>
        To extract semantic triples, ENJU full syntactic parser
        <xref ref-type="bibr" rid="ref5">(Miyao
et al., 2009)</xref>
        , and GENIA event ontology
        <xref ref-type="bibr" rid="ref6">(Kim et al., 2006)</xref>
        are
used. GENIA event ontology can cover biomedical relations.
      </p>
      <p>Table 1 describes the statistical information of the extracted
semantic triples from the whole PubMed abstracts.</p>
      <p>Fig. 1. Search result view with auto query generator
3
3.1</p>
      <sec id="sec-2-1">
        <title>Smart slide</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>SMART SEARCHING SYSTEM</title>
      <p>The proposed searching method has a familiar user interface.
Searching process is started with a query, and the auto completion
function recommends candidate queries.</p>
      <p>Search results contain ranked documents as similar as those
of the common searching systems (Figure 1). However, three
differentiated functions are included in the proposed searching
system as follows:</p>
      <p>First, six biomedical entities are highlighted in the results, and
the information about entities are shown if mouse pointer is
positioned over an entity. The information about the entities
contains the corresponding concepts.</p>
      <p>Second, it is easy to use other services by selecting titles and
highlighted entities. Titles are links to websites of the original
PubMed articles. Once an entity is selected, a popup window
shows up another services.</p>
      <p>Third, a query is easily constructed by selecting an entity or
a verb from candidates. The candidate entities and verbs are
all possible entities or event verbs related to the previously
selected entity or verb. For the first input query, candidate verbs
are listed in the next pane. Once a verb is selected, the next
candidate entities are listed in the next pane.</p>
      <p>This service might be helpful for more specific search with more
specific query, and the search results are displayed immediately
when a query is changed like the Google instant searching service.</p>
      <sec id="sec-3-1">
        <title>3.2 Semantic network browsing</title>
        <p>The proposed searching method describes relations among entities.
A vertex and an edge indicate an entity and a verb (relation),
respectively. All entities in the network contain the identifiers of
external public databases such as UMLS, UniProt, BioThesaurus,
KEGG and DrugBank. Thus, network involves not only the
extracted information from texts but also information of other
external databases. If a vertex is selected, synonyms are listed based
on the frequency, and if an edge is selected, all relations between
two entities are shown with the evidence sentences. Moreover, links
to websites for the original documents are also provided (Figure 2).
3.3 Top 5
The proposed searching method shows popular relations for a query.
As for a query, the results show list of entities or relational verbs
based on frequency of co-occurrences (Figure 3). The co-occurrence
indicates a sentence that contains both two entities, or both an entity
and a verb. A Pie type and a bar type are the way to show the results.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.4 Find It</title>
        <p>The proposed searching method can provide latent attributes
for diseases. Erectile dysfunction, an example in Figure 4,
affects cardiovascular disease, associates with diabetes, can cause
blindness, and is popularly treated by PGE. Evidence sentences can
be shown by selecting corresponding entity.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 CONCLUSION</title>
      <p>Researches about information extraction from literature and
construction of a knowledge basehave been actively conducted.
However, researches about various searching and browsing methods
for the extracted information have been relatively neglectful.</p>
      <p>In the proposed approach, a smart searching system is introduced,
and it contains four useful searching methods to utilize a
multifaceted scientific knowledge effectively. We expect that the
proposed searching system provides various opportunities for
researchers to detect and invent new products such as drugs
conveniently.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Song S. K.</given-names>
            ,
            <surname>Choi</surname>
          </string-name>
          <string-name>
            <given-names>Y. S.</given-names>
            ,
            <surname>Chun</surname>
          </string-name>
          <string-name>
            <given-names>H. W.</given-names>
            ,
            <surname>Jeong</surname>
          </string-name>
          <string-name>
            <given-names>C. H.</given-names>
            ,
            <surname>Choi</surname>
          </string-name>
          <string-name>
            <given-names>S. P.</given-names>
            ,
            <surname>Sung</surname>
          </string-name>
          <string-name>
            <surname>W. K.</surname>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Multi-word Terminology Recognition Using Web Search</article-title>
          .
          <source>UNESST</source>
          <year>2011</year>
          , CCIS
          <volume>264</volume>
          ,
          <fpage>233</fpage>
          -
          <lpage>238</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Chun H. W.</given-names>
            ,
            <surname>Tsuruoka</surname>
          </string-name>
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Kim</surname>
          </string-name>
          <string-name>
            <given-names>J. D.</given-names>
            ,
            <surname>Shiba</surname>
          </string-name>
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Nagata</surname>
          </string-name>
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Hishiki</surname>
          </string-name>
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Tsujii</surname>
          </string-name>
          <string-name>
            <surname>J</surname>
          </string-name>
          . (
          <year>2006</year>
          )
          <article-title>Automatic recognition of topic-classified relations between prostate cancer and genes using MEDLINE abstracts</article-title>
          .
          <source>BMC Bioinformatics</source>
          ,
          <volume>7</volume>
          (
          <issue>Suppl 3</issue>
          ):
          <fpage>S4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Chun H. W.</given-names>
            ,
            <surname>Jeong</surname>
          </string-name>
          <string-name>
            <given-names>C. H.</given-names>
            ,
            <surname>Song</surname>
          </string-name>
          <string-name>
            <given-names>S. K.</given-names>
            ,
            <surname>Choi</surname>
          </string-name>
          <string-name>
            <given-names>Y. S.</given-names>
            ,
            <surname>Choi</surname>
          </string-name>
          <string-name>
            <given-names>S. P.</given-names>
            ,
            <surname>Sung</surname>
          </string-name>
          <string-name>
            <surname>W. K.</surname>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Composite Kernel-based Relation Extraction using Predicate-Argument Structure</article-title>
          .
          <source>UNESST</source>
          <year>2011</year>
          , CCIS
          <volume>264</volume>
          ,
          <fpage>269</fpage>
          -
          <lpage>273</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Yoshida</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsujii</surname>
            <given-names>J</given-names>
          </string-name>
          . (
          <year>2004</year>
          ).
          <article-title>Reranking for Biomedical Named-Entity Recognition</article-title>
          .
          <source>BioNLP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Miyao Y.</given-names>
            ,
            <surname>Sagae</surname>
          </string-name>
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Stre</surname>
          </string-name>
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Matsuzaki</surname>
          </string-name>
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Tsujii</surname>
          </string-name>
          <string-name>
            <surname>J</surname>
          </string-name>
          . (
          <year>2009</year>
          ).
          <article-title>Evaluating Contributions of Natural Language Parsers to Protein-Protein Interaction Extraction</article-title>
          .
          <source>Biomedical Informatics</source>
          ,
          <volume>25</volume>
          (
          <issue>3</issue>
          )
          <fpage>394</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Kim J. D.</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ohta</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teteisi</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsujii</surname>
            <given-names>J</given-names>
          </string-name>
          . (
          <year>2004</year>
          ).
          <source>GENIA Ontology</source>
          .
          <source>Technical Report(TRNLP-UT-2006-2)</source>
          . Tsujii Laboratory, University of Tokyo.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>