<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identifying Chemical Entities based on ChEBI</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tiago Grego</string-name>
          <email>tgrego@fc.ul.pt</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francisco R. Pinto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francisco M. Couto</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Departamento de Química e Bioquímica, Faculdade de Ciências da Universidade de Lisboa</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lasige, Departamento de Informática, Faculdade de Ciências da Universidade de Lisboa</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This software demonstration paper presents Identifying Chemical Entities (ICE), a platform composed by algorithms for chemical entity recognition, entity resolution to a reference database, namely ChEBI, and validation using chemical semantic similarity. It aims to provide the users with an improved display of entity recognition results, exposing outliers which are possible recognition errors and displaying evidence that corroborates consistent chemical entities in the entity recognition and resolution process.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 INTRODUCTION</title>
      <p>Chemical information in the scientific literature is increasing at
a fast pace, making it difficult for researchers to keep up to
date with what is being published. Chemical abstracting services
addressed this issue by having chemistry experts manually extract
the necessary information from the literature. However, with the
exponential growth in publication rate, automatic methods for
chemical entity identification are increasingly needed.</p>
      <p>
        The lack of available chemical terminologies has been an
important aspect for the slow development of chemical text mining
systems, but the recent release of ChEBI
        <xref ref-type="bibr" rid="ref2">(de Matos et al., 2010)</xref>
        allowed their development and some chemical entity recognition
tools are already available. ChEBI is however more than a dictionary
of molecular entities, it is an ontology that provides a structured
classification of the molecular entities. The ontology structure is
essentially a directed acyclic graph (DAG) and comprises three
separate sub-ontologies (Chemical Entity, Role, and Subatomic
Particle). The Chemical Entity sub-ontology provides a structural
relationship between terms while the Role ontology provides a
functional relationship between them, allowing for a thorough
comparison of chemical entities.
      </p>
      <p>ICE (Identifying Chemical Entities) is a software platform that
integrates algorithms for chemical entity recognition in biomedical
literature, resolution of named entities to the ChEBI database, and
validation of annotations using semantic similarity in the ChEBI
ontology to gather annotation evidence. Figure 1 shows an outline
of the ICE architecture.</p>
    </sec>
    <sec id="sec-2">
      <title>RECOGNITION</title>
      <p>
        The software platform uses the algorithms for chemical entity
recognition presented by
        <xref ref-type="bibr" rid="ref4">Grego et al. (2009)</xref>
        . This entity
recognition system follows a machine learning approach using an
implementation of Conditional Random Fields (CRF) to build a
classification model based on a manually annotated patent document
corpus.
      </p>
      <p>The first step in the entity recognition process is the splitting of
the input text into a sequence of tokens, which are then classified
Text
Recognition
Resolution</p>
      <p>Annotated text
mapped to ChEBI</p>
      <p>Validation
Annotated text with
confidence scores
Fig. 1. Architecture of ICE with the three modules for Recognition,</p>
      <p>Resolution and Validation of chemical entities.
according to the previous model. With this method chemical named
entities are located in the input text, however there is no mapping
to a reference database and thus an entity resolution module is
required.
3</p>
    </sec>
    <sec id="sec-3">
      <title>RESOLUTION</title>
      <p>
        For entity resolution to the ChEBI database this software platform
uses the algorithms presented in
        <xref ref-type="bibr" rid="ref5">Grego et al. (2012)</xref>
        . This module
takes as input the string identified as being a chemical compound
name and returns the most relevant ChEBI identifier along with a
confidence score.
      </p>
      <p>A lexical similarity method is used to compare the constituent
words in the input string with the constituent words of each ChEBI
term, to which different weights have been assigned according to
its frequency in the database. A final score is provided with the
mapping and a minimum score threshold can be used to allow for
no mapping to be made in cases where the provided mapping score
is too low, which might be an indication that the term is absent from
ChEBI.
4</p>
    </sec>
    <sec id="sec-4">
      <title>VALIDATION</title>
      <p>A novel algorithm was developed for this software platform to
perform validation of named entities mapped to ChEBI. The
underlying assumption is that most often a text fragment such as
a paragraph has a limited scope, and therefore normally contains
entities that are somehow related to each other. With this in mind,
this algorithm takes as input entities mapped to ChEBI within a text
fragment and searches for relationships between them. The output is
for each input entity the most similar entity within the text fragment,
with the corresponding similarity score.</p>
      <p>
        Using the ontology structure of ChEBI we are able to
compare chemical entities according to both structural and
functional characteristics through several possible semantic
similarity measures
        <xref ref-type="bibr" rid="ref3">(Ferreira and Couto, 2010)</xref>
        . The BOA
framework
        <xref ref-type="bibr" rid="ref6">(Tavares et al., 2011)</xref>
        offers an implementation for
chemical semantic similarity calculation.
      </p>
      <p>Based on the maximum similarity score of each entity we can
filter outliers and corroborate consistent entities. This is performed
using two thresholds. The algorithm has thus four parameters that
can be tuned according to the user requirements: the text fragment
window, which can be the full document, paragraph or sentence; the
semantic similarity measure to be used for the comparison of the
entities; and the two thresholds, which can be tuned to allow for
more precision or recall.</p>
      <p>The final result is advantageous in semi-automated tasks by
providing an improved view over the entity recognition results,
because the user will have an indication of which entities have
increased consistency and are most probably correctly identified, as
well as which entities are outliers in the sense that no similar entities
could be find, which might be an evidence of a recognition error.
5</p>
    </sec>
    <sec id="sec-5">
      <title>EXAMPLE</title>
      <p>As an example lets consider the following sample sentence and
follow the steps of ICE.</p>
      <p>A mixture of ethanol, propanol and acetic acid with a small
amount of sodium chloride.</p>
      <p>In the entity recognition step four entities can be found are now
highlighted.</p>
      <p>A mixture of ethanol, propanol and acetic acid with a small
amount of sodium chloride.</p>
      <p>In entity resolution, ChEBI identifiers are assigned to the entities.</p>
      <p>A mixture of ethanol [CHEBI:16236], propanol [CHEBI:28831]
and acetic acid [CHEBI:15366] with a small amount of sodium
chloride [CHEBI:26710].</p>
      <p>Entities mapped to ChEBI are compared to each other, and those
with high maximum similarity are considered consistent. That is the
case of ethanol and propanol, which have high semantic similarity.
Sodium chloride has a low similarity with the other entities, and is
thus considered an outlier. Acetic acid in the example has reasonable
similarity to both ethanol and propanol, but not high enough to be
considered consistent. The final result would highlight differently
the new three classes of entities.</p>
      <p>In this example, the text fragment window is very small and
thus there are few entities to be compared, which can provide
misleading results. All four entities were in fact correct, but the
indication of consistent and outlier entities still provide interesting
and meaningful explanation that users can use in a semi-automated
fashion.
6</p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSION</title>
      <p>This software demonstration paper presents ICE, a framework that
performs chemical entity recognition and resolution to the ChEBI
database. The entity recognition results are then further processed
using the ChEBI ontology to identify outliers and consistent entities,
providing the user with annotation evidence.</p>
      <p>The goal is to provide the user, potentially a curator performing
semi-automatic annotation tasks, different layers of certainty in the
recognized entities for better analysis of automatic chemical entity
recognition results, as well as providing evidence based on semantic
similarity.</p>
      <p>
        The framework is being extended to other ontologies in addition
to ChEBI using ontology matching techniques that can align shared
concepts between ontologies such as the Gene Ontology and ChEBI
        <xref ref-type="bibr" rid="ref1">(Cruz et al., 2011)</xref>
        .
      </p>
      <p>Also, this software platform will be used in the context of
the project SPNet. This project aims to uncover network motifs
associated with virulence in Streptococcus pneumoniae. This
requires an extensive analysis of the transcriptional and metabolic
networks involved in virulence, and the integration of chemical data
with genomic and proteomic data will be required.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGEMENTS</title>
      <p>The authors want to thank the Portuguese Fundação para a Ciência
e Tecnologia through the financial support of the SPNet project
(PTDC/EIA-EIA/119119/2010), the SOMER project
(PTDC/EIAEIA/119119/2010) and the PhD grant SFRH/BD/36015/2007 and
through the LASIGE multi-annual support.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Cruz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stroe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pesquita</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Couto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Cross</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Biomedical ontology matching using the AgreementMaker system</article-title>
          .
          <source>In Software Demonstration at the International Conference on Biomedical Ontologies (ICBO).</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>de Matos</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alcántara</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dekker</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ennis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hastings</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haug</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spiteri</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Steinbeck</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Chemical Entities of Biological Interest: an update</article-title>
          .
          <source>Nucleic Acids Research</source>
          ,
          <volume>38</volume>
          ,
          <fpage>D249</fpage>
          -
          <lpage>D254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Couto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Semantic similarity for automatic classification of chemical compounds</article-title>
          .
          <source>PLoS computational biology</source>
          ,
          <volume>6</volume>
          (
          <issue>9</issue>
          ),
          <year>e1000937</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Grego</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Pe˛zik,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Couto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            , and
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Identification of chemical entities in patent documents</article-title>
          .
          <source>In Distributed Computing, Artificial Intelligence</source>
          , Bioinformatics, Soft Computing, and Ambient Assisted Living, volume
          <volume>5518</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>942</fpage>
          -
          <lpage>949</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Grego</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pesquita</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bastos</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Couto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Chemical entity recognition and resolution to ChEBI</article-title>
          .
          <source>ISRN Bioinformatics</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Tavares</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bastos</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faria</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grego</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pesquita</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Couto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>The Biomedical Ontology Applications (BOA) framework</article-title>
          .
          <source>In Software Demonstration at the International Conference on Biomedical Ontologies (ICBO).</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>