<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Use of text mining for Experimental Factor Ontology coverage expansion in the scope of target validation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Şenay Kafkas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ian Dunham</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Helen Parkinson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jo McEntyre</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Resources Used</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>European Bioinformatics Institute - European Molecular Biology Laboratory (EMBL-EBI), and Open Targets Wellcome Genome Campus Hinxton</institution>
          ,
          <addr-line>CB10 1SD</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Understanding the molecular biology and development of disease plays a key role in drug development. Integrating evidence from different experimental approaches with data available from public resources (such as gene expression level changes and reaction pathways affected by pathogenic mutations) can be a powerful approach for evaluating different aspects of target-disease associations. The application of ontologies is of fundamental importance to effective integration. The Target Validation Platform is a user-friendly interface that integrates such evidences from various resources with the aim of assisting scientists to identify and prioritise drug targets. Currently, the EFO is used as the reference ontology for diseases in the platform, importing terms from existing disease ontologies such as the Human Phenotype Ontology as required. In order to generalize the use of EFO from key target-diseases for wider use, we need to compare the target associated disease coverage in EFO with the scope of other available disease terminology resources. In this study, we address this issue by using text mining and present our initial results.</p>
      </abstract>
      <kwd-group>
        <kwd>text mining</kwd>
        <kwd>ontology</kwd>
        <kwd>integration</kwd>
        <kwd>target validation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        Integrating data from de novo experiments with data
available in public data resources in a user friendly interface to
support decision making has been the goal of the Target
Validation Platform (https://targetvalidation.org). This
platform integrates a variety of evidence for a given target
(gene/protein) - disease association, such as reaction pathways
that are affected by pathogenic mutations from Reactome [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
and text mined target-disease associations from the Europe
PubMed Central (Europe PMC) (http://europepmc.org/)
literature database [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The application of disease ontologies is
critical to integrate such different data types.
      </p>
      <p>
        The Experimental Factor Ontology (EFO)
(http://www.ebi.ac.uk/efo/) is the reference resource for
diseases in the platform (“disease” here encompasses both
“disease/phenotype” as the disease/phenotype boundary is
blurred in both the platform’s data sources and ontologically).
Therefore, it is important to understand the disease coverage of
EFO in the scope of target validation, in comparison to the
other available major disease and phenotype resources, in order
to expand its disease coverage. In this study, we address this
issue by using text mining which is a widely used approach in
ontology expansion [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and target-disease association
identification [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], to compare terms available in existing
ontologies and present our initial results.
      </p>
      <p>We used Europe PMC as the literature database, UniProt
for target (gene/protein) names and six major disease
terminologies: EFO (V2.69), the Human Phenotype ontology
(HP) (access date:31-03-2016)
(http://human-phenotypeontology.github.io/), Orphanet Rare Disease Ontology
(ORDO) (V2.1)
(http://www.orphadata.org/cgibin/inc/ordo_orphanet.inc.php), the Human Disease Ontology
(HDO) (06-01-2016 update) (http://disease-ontology.org/), the
Mammalian Phenotype Ontology (MP) (access
date:31-032016)
(http://www.informatics.jax.org/searches/MP_form.shtml),
and Unified Medical Language Systems (UMLS) (2014 AB
Release) (https://www.nlm.nih.gov/research/umls/).</p>
      <p>Europe PMC is one of the largest biomedical literature
databases in the World which provides public access to 31
million abstracts and 3.7 million full text articles, covering
both PubMed and PubMed Central. In our analyses, we used
the latest achieved version of the Open Access full text articles
(~1 Million) (http://europepmc.org/ftp/archive/v.2016.03/)
from the database.</p>
      <p>We generated and refined dictionaries from the human part
of the SwissProt Database (the expert annotated part of
UniProt) (http://www.uniprot.org/) and disease and phenotype
parts of EFO, HP, ORDO, MP, HDO and UMLS before
applying text mining. In the refining process, we filtered out
the terms that would introduce potentially high numbers of
false positives. These are the terms having character length &lt;
3 and the terms that are ambiguous with common English
words (e.g. “Large” is a protein name as well). In addition, we
generated term variations by replacing the widely used Greek
letters in gene/disease names with their symbols (e.g.
replacing “alpha” with α). The final target and disease
dictionaries consisted of a total of 104,434 Uniprot, 26,617
EFO, 18,332 HP, 20,152 ORDO, 29,800 MP, 21,789 HDO
and 75,060 UMLS terms.</p>
      <sec id="sec-1-1">
        <title>B. Target and disease name identification</title>
        <p>
          We used the Europe PMC text-mining pipeline, which is
based on Whatizit [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] to annotate target and disease names in
text with the dictionaries described above. Target and disease
name abbreviations can be ambiguous with some other names
(e.g. ALS which is “Amyotrophic Lateral Sclerosis”, is
ambiguous with “Advanced Life Support”, PMID:26811420).
Therefore, we implemented and used abbreviation filters for
screening out the potential false positive disease/protein
abbreviations introduced during the annotation process. The
abbreviation filters operate based on several heuristic rules.
For example, text sequences within parentheses (i.e. (XYZ)),
appearing in uppercase and having length &lt;6 are identified as
a name abbreviation candidate and are retained as an
annotation only if any of its long forms from the given disease
ontology exists elsewhere in the document.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>C. Target-disease association extraction</title>
        <p>
          The associations are extracted by identifying the
targetdisease co-occurrences at the sentence level and applying
several filtering rules to reduce noise possibly introduced by
the high sensitivity, low specificity co-occurrence approach.
The filtering rules utilise heuristic information from a careful
manual analysis of the text. They include, filtering out all
articles but the “Research” articles (e.g. Reviews, Case
Reports), filtering out target-disease associations appearing in
certain sections such as “Methods” and “References”, and
filtering out target-disease associations that appear only once in
the body of a given article but not in the article's title or
abstract (see [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] for the details).
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>III. RESULTS AND DISCUSSION</title>
      <p>
        Our target-disease extraction system achieves a
MeanAverage Precision value of 81% [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Figure 1 presents a Venn
diagram showing the disease terms found in the corpus that
are associated with targets, after application of the
targetdisease heuristics above, for each of the six different disease
resources. There are 3,859 HDO, 3,384 MP, 1,610 ORDO,
4,277 HP and 17,584 UMLS target associated distinct disease
terms that are not found by EFO. Possible reasons for the
difference in coverage between EFO and the other
terminologies are twofold: nonexistence of a given disease
name in EFO, the coverage of a given disease with different
Fig1. Venn diagram showing overlapping target associated disease terms
synonyms and different classification of a given term in EFO.
For example, “fetal valproate syndrome” and “Chagas
cardiomyopathy” from ORDO are not covered by EFO. “HIV”
is classified as “disease and syndrome” in UMLS, indicating
“HIV infection”, however, in EFO, it is classified as a virus
name. Results suggest that there is some room for
improvement in the EFO and this will be explored for future
releases of EFO.
      </p>
    </sec>
    <sec id="sec-3">
      <title>IV. CONCLUSION AND FUTURE WORK</title>
      <p>In this study, we demonstrate the use of text mining for
analysing and suggesting approaches to expand the
disease/phenotype coverage of EFO within the scope of target
validation. We focused on the target-associated disease terms
from EFO and five other major disease resources, but there is
no reason why this approach could not be applied to other
contexts in efforts to integrate across terminologies and
ontologies. In future, we will extend our analysis to discover
any trends over the resources, to understand the
disease/phenotype target space derived from literature and how
much of the associations that we find in EFO scope is relevant.</p>
    </sec>
    <sec id="sec-4">
      <title>ACKNOWLEDGMENTS</title>
    </sec>
    <sec id="sec-5">
      <title>This work is funded by the Open Targets.</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabregat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sidiropoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Garapati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gillespie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hausmann</surname>
          </string-name>
          et al.,
          <source>“The Reactome pathway Knowledgebase,” Nucleic Acids Res</source>
          .,
          <volume>44</volume>
          (
          <issue>D1</issue>
          ):
          <fpage>D481</fpage>
          -
          <lpage>7</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Europe</surname>
            <given-names>PMC Consortium</given-names>
          </string-name>
          , “
          <article-title>Europe PMC: a full-text literature database for the life sciences and platform for innovation,” Nucleic Acids Res</article-title>
          .,
          <volume>43</volume>
          (Database issue):
          <fpage>D1042</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I.</given-names>
            <surname>Spasic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ananiadou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>McNaught</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          , “
          <article-title>Text Mining and Ontologies in biomedicine:Making sense of raw text</article-title>
          ,” Breefings in Bioinformatics,
          <volume>6</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>239</fpage>
          -
          <lpage>251</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pletscher-Frankild</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pallejà</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tsafou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.X.</given-names>
            <surname>Binder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.J.</given-names>
            <surname>Jensena</surname>
          </string-name>
          , “DISEASES:
          <article-title>Text mining and data integration of disease-gene associations</article-title>
          ,
          <source>” Methods</source>
          ,
          <volume>34</volume>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>89</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Arregui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gaudan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kirsch</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Jimeno, “Text processing through Web services: calling Whatizit,” Bioinformatics, vol.
          <volume>24</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>296</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Ş.</given-names>
            <surname>Kafkas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Dunham</surname>
          </string-name>
          and
          <string-name>
            <surname>J. McEntyre</surname>
          </string-name>
          , “
          <article-title>Literature Evidence in Open Targets- a target validation platform,” Phenotype Day @ISMB 2016, special session of the Bio-</article-title>
          <string-name>
            <surname>Ontologies</surname>
            <given-names>SIG</given-names>
          </string-name>
          ,
          <fpage>8</fpage>
          -
          <issue>12</issue>
          <year>July 2016</year>
          , Orlando, Florida, U.S.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>