<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MutDR - A Resource for Protein Mutation-Disease Relations Assembled from Biomedical Literature</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ravikumar Komandur Elayavilli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Majid Rastegar-Mojarad</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hongfang Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Health Sciences Research Mayo Clinic Rochester, MN</institution>
          ,
          <addr-line>55901</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>- Text mining approaches can accelerate the process of assembling knowledge from literature. In this abstract, we present our effort in assembling a resource for protein mutationdisease relations assembled from literature.</p>
      </abstract>
      <kwd-group>
        <kwd>literature mining</kwd>
        <kwd>protein mutation-disease</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>A large amount of information about the role of gene
variants and mutations in diseases is available in curated
databases such as OMIM [1], ClinVar [2], and UniprotKB [3].
However, much of this information remains ‘locked’ in the
unstructured form in the scientific publications. Since manual
curation involves significant human effort and time there is
always a lag in the information between the curated databases
and the literature. The recent findings published in the
literature takes significant time to find its way into the curated
knowledgebase. Text mining approaches can accelerate the
process of assembling this knowledge from the published
literature. However, developing a text-mining system with
semantic understanding capability in the biomedical domain is
very challenging. In an earlier work, we described MutD [4], a
literature mining system that extracts relationship between
protein point mutation and diseases from bio-medical
abstracts. In this abstract, we present access to a PubMed scale
resource through a web interface that allows users to retrieve
protein point mutation-disease relations extracted through
biomedical literature mining.</p>
    </sec>
    <sec id="sec-2">
      <title>II. BACKGROUND</title>
      <p>MutD is a literature mining system that uses an ensemble
of state of the art named entity extraction and normalization
tools and graph based dependency parse representation to
extract relations between protein point mutations and diseases
mentioned in biomedical abstracts. It also extends the scope of
literature mining to across multiple sentences through
discourse processing and heuristics. MutD achieved a
precision of 71% and recall of 58% (F-Measure: 64%) when
compared against the annotations of UniProtKB.</p>
      <p>III.</p>
    </sec>
    <sec id="sec-3">
      <title>METHODS</title>
      <p>In this work, we describe the extension of MutD to create a
PubMed scale resource of literature-mined Protein Point
mutation and disease (PMD) relations, MutDR. Figure 1
outlines the overall workflow in developing MutDR.</p>
      <p>Using MutD, we performed a large-scale mining of PMD
relations on the complete PubMed data set (till May 2016).
The extracted PMD relations were indexed using Elastic
Search and a simple web based search interface was
developed to enable users to retrieve the literature-mined
PMD relations. The user interface has three major
functionalities: 1) Query by a gene/protein, a disease or both
2) Retrieve results ranked according to the relevance of the
query and further by date. 3) Link the normalized entities
genes and diseases to the external knowledge resources
namely UniProtKB and Comparative Toxicogenomics
database (CTD) [5] respectively.</p>
      <p>Fig.1 – Overall architecture workflow</p>
    </sec>
    <sec id="sec-4">
      <title>IV. RESULTS</title>
      <p>MutD extracted 27, 213 protein mutation disease relations
from nearly 81, 048 PubMed abstracts (out of the total 21
million abstracts). Figure 2 shows some of the user interface
features of MutD resource. The PMD relations extracted by
MutD are indexed using Elastic Search [6].</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Amberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Bocchini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Scott</surname>
          </string-name>
          , A. HamoshA, “
          <article-title>McKusick's Online Mendelian Inheritance in Man (OMIM</article-title>
          ),
          <source>” Nucleic Acids Res</source>
          <volume>37</volume>
          :
          <fpage>D793</fpage>
          -
          <lpage>6</lpage>
          ,
          <fpage>2009</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>M. J. Landrum</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>G. R.</given-names>
          </string-name>
          <string-name>
            <surname>Riley</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Jang</surname>
            ,
            <given-names>W. S.</given-names>
          </string-name>
          <string-name>
            <surname>Rubinstein</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Church</surname>
          </string-name>
          , et al.,
          <article-title>“ClinVar: public archive of relationships among sequence variation and human phenotype</article-title>
          .
          <source>” Nucleic Acids Res</source>
          <volume>42</volume>
          :
          <fpage>D980</fpage>
          -
          <lpage>5</lpage>
          ,
          <fpage>2014</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>The UniProt Consortium</surname>
          </string-name>
          , “
          <article-title>Activities at the Universal Protein Resource (UniProt</article-title>
          ),
          <source>” Nucleic Acids Res</source>
          <volume>42</volume>
          :
          <fpage>D191</fpage>
          -
          <lpage>D198</lpage>
          ,
          <year>2014</year>
          K. E. Ravikumar,
          <string-name>
            <surname>K. B. Wagholikar</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          <string-name>
            <surname>Kocher</surname>
          </string-name>
          , and H. Liu, “
          <article-title>Text mining facilitates database curation-extraction of mutation-disease associations from Bio-medical literature</article-title>
          ,
          <source>” BMC Bioinformatics</source>
          ,
          <volume>16</volume>
          (
          <issue>1</issue>
          ),
          <fpage>185</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Grondin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lennon-Hopkins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Saraceni-Richards</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sciaky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. L.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. C.</given-names>
            <surname>Wiegers</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Mattingly</surname>
          </string-name>
          , “
          <source>The Comparative Toxicogenomics Database's 10th year anniversary: update 2015,” Nucleic Acids Res</source>
          <volume>43</volume>
          :
          <fpage>D914</fpage>
          -
          <lpage>D920</lpage>
          , 2015
          <string-name>
            <given-names>R.</given-names>
            <surname>Kuc</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Rogozinski</surname>
          </string-name>
          . Elasticsearch Server.
          <source>Packt Publishing Ltd</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>