<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Quality-Assurance Study of ChEBI</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hasan Yumak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ling Chen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CUNY New York</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>NY USA</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>hyumak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>lchen}@bmcc.cuny.edu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Michael Halper</institution>
          ,
          <addr-line>Ling Zheng, Yehoshua Perl, Gai Elhanan</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>New Jersey Institute of Technology Newark</institution>
          ,
          <addr-line>NJ</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Ontologies are important components of many health-information systems. The Chemical Entities of Biological Interest (ChEBI) ontology has become a standard reference for chemicals appearing in biological contexts. As such, assuring the quality of its content is imperative. In fact, ChEBI has a dedicated Web page at which errors and inconsistencies in its concepts can be reported. A study of the correctness of a random sample of ChEBI concepts is carried out. The results show that quite a large number of ChEBI concepts suffer from some kind of problematic modeling. For example, we found that 15.5% of the sample concepts exhibited severe errors of commission, including incorrect hierarchical (is a) and lateral relationships. Errors of omission were also prevalent. The overall results of our quality-assurance (QA) study are presented. Suggestions for enhancing the QA processes in place for ChEBI are discussed.</p>
      </abstract>
      <kwd-group>
        <kwd>ChEBI</kwd>
        <kwd>chemical ontology</kwd>
        <kwd>chemical concept</kwd>
        <kwd>quality assurance</kwd>
        <kwd>modeling error</kwd>
        <kwd>error distribution</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        Ontologies are structures that capture terminological
knowledge for some target domain. Typically large in size and
high in complexity, ontologies have become fundamental
fixtures of health and biological information processing
environments. The Chemical Entities of Biological Interest
(ChEBI) ontology [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is an authoritative reference that models
chemical concepts having biological significance, particularly
from the perspectives of molecular structure and biological role
or application [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Maintained by the European Molecular
Biology Laboratory–European Bioinformatics Institute
(EMBL-EBI), it is an important chemical annotation and
identification standard. As of its February 2016 release, it
comprised a collection of 61,895 concepts (including 47,752
fully annotated compounds), 104,351 is a (hierarchical)
relationships and 65,077 lateral relationships.
      </p>
      <p>
        Due to their scope and complexity, it is nearly impossible
for ontologies such as ChEBI to be free of modeling errors and
inconsistencies. This can hinder their usefulness and adversely
affect the software systems and applications dependent on
them. ChEBI has been used, for example, as a source for
annotations in various bioinformatics databases, including
UniProt which utilized ChEBI for the cofactor comments
related to enzymes [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Metabolites in a human metabolism
model have been annotated with terms from ChEBI [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. ChEBI
is also used to support text mining and chemical analysis. For
example, in a recent study [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], a novel method for computing
semantic similarity between chemical entries based on ChEBI
was introduced to improve the chemical entity identification in
texts. In another study [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a new prediction method based on
information from ChEBI for identifying drugs’ target groups
was proposed. ChEBI’s structural hierarchy has been
integrated with the Gene Ontology (GO) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to allow for data
integration across the biology and chemistry domains. Errors in
ChEBI could be propagated in a deleterious manner into such
applications. Due to this, assuring the quality of the conceptual
content of an ontology is a very important matter. ChEBI, in
fact, employs a GitHub issue tracking system
(https://github.com/ebi-chebi/ChEBI/issues) to enable users to
report various errors and inconsistencies that they encounter
while using the ontology. Those reports are handled by
ChEBI’s curators.
      </p>
      <p>In this paper, we are interested in assessing the percentage
of ChEBI’s concepts that suffer from some type of modeling
issues. We refer to concepts that have errors or inconsistencies
in their modeling as being erroneous concepts. We selected a
random sample of 400 concepts for our study. Two
subjectdomain experts in the field of chemistry were asked to carry
out separate quality-assurance (QA) analyses of the entire
sample and then produce a consensus report. The findings
revealed that a substantial number of ChEBI concepts suffer
from errors of commission and omission, suggesting the need
for a formal initiative to map out an ongoing QA project for
improvement of the quality of the modeling in ChEBI. Further
recommendations for such a project are discussed.</p>
    </sec>
    <sec id="sec-2">
      <title>II. METHODS</title>
      <p>
        A QA study of a random sample of ChEBI concepts was
carried out. ChEBI is updated monthly, and the version we
used in this study was that of February 2016, which contained
61,895 concepts, including 47,752 fully annotated compounds.
The sample was chosen using the basic sampling technique,
simple random sampling without replacement [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The
sampling frame was all 61,895 concepts in the February 2016
release. A random number drawn from a uniform distribution
over the range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] was generated as a key for each concept.
All concepts were sorted using the keys, and the smallest 400
concepts were selected as the random sample [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>It should be noted that ChEBI employs three hierarchies to
classify molecular entities. Its chemical entity hierarchy, the
largest with 60,537 concepts (97.8% of all ChEBI concepts), is
used to classify molecular entities according to their chemical
structure. The subatomic particle hierarchy with 42 concepts
categorizes particles smaller than atoms. ChEBI’s third
hierarchy, the role hierarchy with 1,322 concepts, itself has
three subhierarchies that define the roles in different contexts
for the compounds, namely, the application subhierarchy to
represent the intended use by humans for the compounds (e.g.,
fuel and anti-inflammatory agent), the biological role
subhierarchy to represent the roles of compounds within the
biological context (e.g., growth regulator and inhibitor), and
the chemical role subhierarchy (e.g., acid and base). Note that
there are six concepts belonging to both the chemical entity
and subatomic particle hierarchies, e.g., helion and proton. The
concepts randomly selected for our study came from all three
of the main hierarchies, irrespective of hierarchy.</p>
      <p>The QA analysis was done by a pair of chemistry
subjectdomain experts. In the initial step, each concept from the
sample was inspected by each of the experts separately—
without any communication between them. Their results were
tabulated in two individual error reports. Within a report, the
rationale for the judgment of any error was recorded, and a
suggested correction was proffered. Afterward, a combined
report was prepared listing the respective findings of both
experts for all the concepts. This report was shared with both
experts who were then each asked separately to mark their
agreement or disagreement with the findings of the other
person—and to review their own findings in light of the
other’s. After a review of the other’s report, each expert was
able to change their mind regarding their own judgment of a
modeling error. A concept previously judged to be modeled in
error could instead be deemed to be correct, and vice versa.
After this step of consensus building, a concept was deemed to
be erroneous if both subject-domain experts agree that it was
such.</p>
      <p>
        Most of the QA analysis of ChEBI centered on issues with
concepts’ relationships. The three primary relationships in
ChEBI are the hierarchical is a relationship, capturing
standard subsumption in hierarchies, the relationships has
part, indicating the whole/part association between
compounds, and has role, linking concepts in the chemical
entity hierarchy to concepts in the role hierarchy. There are
seven chemistry-specific lateral relationships, namely, is
conjugate base of, is conjugate acid of, is tautomer of, is
enantiomer of, has functional parent, has parent hydride, and
is substituent group from. In combination, the relationships in
ChEBI can form converses. For example, if concept A is
conjugate base of concept B, then B is conjugate acid of A. A
similar situation exists for the two relationships is tautomer of
and is enantiomer of [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>The types of errors that the experts were looking for
included both errors of commission and errors of omission.
Examples of the former are incorrect hierarchical relationship,
incorrect lateral relationship, and incorrect relationship target.
Examples of the latter are missing hierarchical relationship
and missing lateral relationship.</p>
    </sec>
    <sec id="sec-3">
      <title>III. RESULTS</title>
      <p>A random sample of 400 concepts (0.6%) was selected
from the 61,895 concepts in ChEBI, February 2016 release.
Out of these 400 concepts, 388 (97%) are from the chemical
entity hierarchy, 11 are from the role hierarchy, and one is
from the subatomic particle hierarchy. In the following, we
often refer to a ChEBI concept using its name along with its
unique ChEBI id, written, for example, as “CHEBI: 31900”
(which is the concept with the name Neticonazole
hydrochloride).</p>
      <p>
        Two of the authors (HY and LC), both subject-domain
experts in chemistry and experienced in ontology QA, carried
out the individual QA analyses on the entire sample of
concepts and then produced a consensus report on the errors
discovered. Out of the 400 concepts, the two subject-domain
experts agreed on the errors reported for 167 concepts (41.8%).
The margin of error at the 95% confidence level for a 400
concept sample from a population of 61,895 concepts is 4.9%
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Among the 167 erroneous concepts, 166 of them are
from the chemical entity hierarchy, and the other is from the
role hierarchy.
      </p>
      <p>There were 122 (30.5% = 122 / 400) concepts that
exhibited errors of omission. Of these, 121 concepts were
found to be missing hierarchical relationships, and only one
concept, fatty acid anion 4:0 (CHEBI: 78115), was reported to
be missing the relationship is conjugate base of.</p>
      <p>Table 1 shows the number and percentage of concepts with
errors of commission. For example, 36 concepts (9%) in the
sample were found to have incorrect is a relationships. Note
that some concepts may have multiple kinds of errors. For
example, there are 17 concepts that were reported to have both
errors of commission and omission.</p>
      <p>Table 2 and Table 3 list examples of erroneous concepts
with errors of commission and omission, respectively, along
with their corresponding suggested corrections and the reason
for the error. For example, in Table 2 (Row 3), we see that
Neticonazole hydrochloride (CHEBI: 31900) was originally
modeled as is a hydrochloride; instead, the modeling should
be has part because the mixture contains hydrochloride.</p>
      <sec id="sec-3-1">
        <title>A. Missing chemical classification</title>
        <p>This is the most common error, with 30.25% of the
concepts exhibiting it. For example, there are 23
benzenecontaining compounds in the sample that should be
classified as a benzenoid aromatic compound (CHEBI:
33836). Included among these is
8-hydroxy-3-chlorodibenzofuran (CHEBI: 79743).</p>
      </sec>
      <sec id="sec-3-2">
        <title>B. Incorrect charge difference between conjugate acid and conjugate base</title>
        <p>This is the second most common error (7%). The correct
charge difference between an acid and its conjugate base is
1, and the acid is 1 charge higher. This is because the acid
has one extra proton (H+) compared to its conjugate base.
As indicated in the equation HA  A− + H+, HA is the
conjugate acid of A−, and A− is the conjugate base of HA.
The only difference between HA and A− is one proton; thus,
HA is 1 charge higher than A−. For example, ChEBI
concept 1-(2-carboxyphenylamino)-1-deoxy-D-ribulose
5phosphate (CHEBI: 29112) has as its conjugate base
1-(2carboxylatophenylamino)-1-deoxy-D-ribulose 5-phosphate
(3−) (CHEBI: 58613). However, its conjugate base charge
should be 1−, not 3−, since only one proton is removed from
the acid, not three protons, as shown in its structure in
Table 5.</p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Incorrect chemical classification</title>
        <p>Piperidine (CHEBI: 18049) is classified under Brønsted
acid (CHEBI: 39141), but in fact it is should be classified as
a Brønsted base (CHEBI: 39142). This is because an amine
is a proton acceptor, not a donor, and it acts as a base not as
an acid. Another more common occurrence of this kind of
error is seen for 14 erroneous concepts that are classified as
some class “A,” but, in fact, do not have chemical structure
A. For example, in Fig. 1 showing the chemical structure of</p>
      </sec>
      <sec id="sec-3-4">
        <title>1(3)-O-(alk-1-enyl)-glycerol (CHEBI: 77998), it does not</title>
        <p>contain the following chemical groups carboxylic ester
(CHEBI: 33308), carbonyl compound (CHEBI: 36586), and
ester (CHEBI: 35701). However, it is classified as a
carboxylic ester, carbonyl compound, and ester in ChEBI.</p>
        <p>There are 13 concepts reported with incorrect numbers
of cyclic units. For example, buspirone hydrochloride
(CHEBI: 3224) is classified as an organic heteromonocyclic
compound (CHEBI: 25693) that has one cyclic structure.
However, from the structure seen in Fig. 2, we can see that
this concept contains four cyclic units. Similar errors are
seen in (+)-tephrosone (CHEBI: 66201),
3beta,13-Dihydroxy-16-(hydroxymethylene)-13,17-seco-5alpha-androstan-17-oic acid, delta-lactone (CHEBI: 79677),
c[G(2',5')pA(3',5')p] (CHEBI: 75947), diazoline (CHEBI:
53123), dipyridodiazepine (CHEBI: 63667),
pyrazolopyridazine (CHEBI: 48383), etc.</p>
        <p>Fig. 2. Structure of buspirone hydrochloride (CHEBI: 3224)</p>
      </sec>
      <sec id="sec-3-5">
        <title>E. Incorrect amide classification</title>
        <p>There are seven concepts classified as primary amide
(CHEBI: 33256). However, they should be classified as
secondary amide (CHEBI: 33257). For example, from
Fig. 3, we can clearly see that Arachidonoyl dopamine
(CHEBI: 31231) is a secondary amide (nitrogen group
connected two carbon atoms), while it is denoted as a
primary amide. Similar errors are seen in
beta-D-glucosyl(1&lt;-&gt;1')-N-eico-sanoylsphinganine (CHEBI: 84703),
bistratamide I (CHEBI: 65508), N-(2
hydroxyhexacosanoyl)phyto-sphingosine (CHEBI: 64958),
N-(3oxohexanoyl)homo-serine lactone (CHEBI: 29640),
N-(2hydroxy-docosanoyl)eicosasphinganine (CHEBI: 66983),
and N(4)-acetylcytidine (CHEBI: 70989).</p>
        <p>As an example, the structure of diacylglycerol 38:7
(CHEBI: 86986) does not match with its name. Its name
indicates that there are two esters, while its structure, seen in
Fig. 4, has three R groups, meaning a triacylglycerol, not a
diacylglycerol.</p>
        <p>
          In this study, we discovered an error-rate of 42% in a
random sample of ChEBI concepts. Among these, 15.5%
suffered from severe errors of commission, e.g., incorrect
parents and incorrect relationship targets. The remaining
30.5% exhibited errors of omission. (Some concepts had
both kinds of errors.) One needs to compare this finding to
the reality in other ontologies of a similar caliber. One
ontology for which data of this kind exist is SNOMED CT
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. In several previous studies performed by our SABOC
team for evaluating various QA methodologies for
SNOMED CT, we were also measuring the percentage of
erroneous concepts in random control samples [
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16">13-16</xref>
          ]. In
those studies, we encountered error-rate percentages of
8.3%, 29%, 13%, 8.8%, and 9%, respectively, for the
control samples. Hence, the average control sample error
rate was 13.62%. In this light, the results of the present
study can be taken to be troubling.
        </p>
        <p>Let us note that ChEBI employs a star status (rating)
system to indicate the level of annotation applied to a
concept by the ChEBI curatorial team. A concept manually
annotated by the team has a “3-star” status. A concept
manually annotated by a third party has a “2-star” status. A
preliminary concept loaded automatically from a data source
but not yet manually annotated is designated with a “1-star”
status. We looked at the distribution of erroneous concepts
according to their star status (see Table 6). Out of the 400
concepts that we analyzed, 292 concepts had a 3-star status,
and 116 of those concepts (39.7%) were deemed to be
erroneous. Among the 1-star concepts, 3.3% were erroneous
(Table 6). And 64.1% of the 2-star concepts were erroneous.
Collectively, the non-3-star concepts exhibited an error-rate
of 47.2% (51 out of 108). So, while it was not surprising
that the 3-star concepts showed an overall lower error rate,
they still contributed significantly to the error findings with
a rate of nearly 40%.</p>
        <p>
          In the present study, we assessed the frequency of errors
of commission and omission in a random sample of
concepts from ChEBI. During our original design and
analysis, we postulated that concepts with higher numbers
of parents would exhibit higher error rates. This postulation
was based on a recurring, substantiated theme in our
previous ontological QA research that more complex
concepts are prone to exhibit higher error rates than
concepts in random control samples. Concepts with multiple
parents are indeed more complex than concepts with a
single parent due to multiple inheritance of properties and
convergence of definitional paths. For example, in our QA
research on the CORE problem list of SNOMED CT, we
indeed have shown that the expected error rate increases
with the number of parents [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
        <p>However, as seen in Fig. 5—and contrary to our
expectations—our postulation was wrong. In fact, our
findings show that the error rate is inversely proportional
with respect to the number of parents. For example, 12,128
concepts in ChEBI have two parents. From these, 97 were
chosen for our random sample; 33 of them (34.0%) were
found to be erroneous. Of the 38 three-parent concepts in
the sample, 12 (31.6%) were erroneous. Further reductions
in the error rates were seen for four-parent (22.2%) and
fiveparent concepts (14.3%).</p>
        <p>
          ChEBI is user driven, and user requests can be made via
the ChEBI submission tool [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. For a new concept request,
users need to provide minimal unique information, including
the classifications for the new entity. Users can also report
issues or bugs using ChEBI’s GitHub issue tracking system
(https://github.com/ebi-chebi/ChEBI/issues). As of June
2016, there were 2,933 closed issues and 234 open issues in
the tracking system. After ChEBI’s curators have verified
requests, new concepts and properties are made available in
subsequent releases. For example, a user reported on
December 11, 2015 that an is a relationship should be added
between the concepts endocannabinoid (CHEBI: 67197) and
lipid (CHEBI: 18059). On January 28, 2016, a ChEBI
curator responded that the change was done. From an
inspection of the ontology in its January 2016 version, we
can see that endocannabinoid (CHEBI: 67197) has only one
is a relationship to cannabinoid (CHEBI: 67194), while there
is a new is a to lipid (CHEBI: 18059) in the latest (June 2016)
version. To date, the errors we found in this study have been
submitted to ChEBI via GitHub.
        </p>
        <p>The role of ChEBI in chemistry applications is
significant. Therefore, findings of problems at this high
level are of major concern. During the past 20 years,
SNOMED CT (in which we have seen lower error rates in
random concept samples) has been managed by a variety of
large professional organizations, such as the College of
American Pathologists (CAP), the National Library of
Medicine (NLM), and the IHTSDO. Unfortunately, ChEBI
does not have the same level of resources that have been
available for the maintenance of SNOMED CT. Hence a
creative solution to handle the QA chores of ChEBI is
needed utilizing ChEBI’s curatorial board. It certainly
makes sense, as a start, for the curatorial board of ChEBI to
conduct a follow-up study of an even larger sample of
concepts than the one used in the present study to further
assess the error rates in ChEBI for errors of omission and
commission.</p>
        <p>
          In future work, ChEBI’s curatorial board may want to
identify criteria that can be used in locating subsets of
concepts that are more likely to be erroneous. As noted, the
number of parents as such a criterion did not prove useful,
but maybe there are others that will. Employing useful
methodologies to help automate aspects of QA efforts
should increase the yield of corrections with respect to the
curators’ expended time. Another potential way to measure
the complexity of concepts is by their number of
relationships. In a future study, we will test to see if ChEBI
concepts with larger numbers of relationships have higher
error rates than concepts with fewer relationships. In a
recent study, this was shown to be true with statistical
significance for the concepts in the Biological Process
hierarchy of the National Cancer Institute thesaurus (NCIt)
[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>V. CONCLUSIONS</title>
      <p>In this paper, we reported on a quality-assurance (QA)
study that was carried out on a sample of ChEBI concepts
by two chemistry subject-domain experts. The results
revealed that quite a few ChEBI concepts suffer from some
kinds of modeling problems. Our consensus report found
that 15.5% of the concepts from our sample exhibited severe
errors of commission. Particularly prevalent were errors of
the type “incorrect and missing chemical classification” and
“incorrect charge differences between conjugate acids and
conjugate bases.” These findings are particularly troubling
taking into account the importance of ChEBI and the many
applications dependent on it. In general, it appears that the
QA processes in place for ChEBI could use further
refinement. For example, a targeted effort to review all
charge differences between conjugate acids and conjugate
bases in ChEBI seems warranted.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENT</title>
      <p>Research reported in this publication was supported by
the National Cancer Institute of the National Institutes of
Health under Award Number R01CA190779. The content is
solely the responsibility of the authors and does not
necessarily represent the views of the National Institutes of
Health.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hastings</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Owen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dekker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ennis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Muthukrishnan</surname>
          </string-name>
          , et al. ChEBI in 2016:
          <article-title>Improved services and an expanding collection of metabolites</article-title>
          .
          <source>Nucleic Acids Res</source>
          .
          <year>2016</year>
          ;
          <volume>44</volume>
          (
          <issue>D1</issue>
          ):
          <fpage>D1214</fpage>
          -
          <lpage>1219</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Degtyarenko</surname>
          </string-name>
          , P. de Matos,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ennis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hastings</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zbinden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McNaught</surname>
          </string-name>
          , et al.
          <article-title>ChEBI: a database and ontology for chemical entities of biological interest</article-title>
          .
          <source>Nucleic Acids Res</source>
          .
          <year>2008</year>
          ;
          <volume>36</volume>
          (Database issue):
          <fpage>D344</fpage>
          -
          <lpage>350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>UniProt. UniProt:</surname>
          </string-name>
          <article-title>a hub for protein information</article-title>
          .
          <source>Nucleic Acids Res</source>
          .
          <year>2015</year>
          ;
          <volume>43</volume>
          (Database issue):
          <fpage>D204</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>I.</given-names>
            <surname>Thiele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Swainston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Fleming</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hoppe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sahoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. K.</given-names>
            <surname>Aurich</surname>
          </string-name>
          , et al.
          <article-title>A community-driven global reconstruction of human metabolism</article-title>
          .
          <source>Nat Biotechnol</source>
          .
          <year>2013</year>
          ;
          <volume>31</volume>
          (
          <issue>5</issue>
          ):
          <fpage>419</fpage>
          -
          <lpage>425</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lamurias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Ferreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Couto</surname>
          </string-name>
          .
          <article-title>Improving chemical entity recognition through h-index based semantic similarity</article-title>
          .
          <source>J Cheminform</source>
          .
          <year>2015</year>
          ;
          <article-title>7(Suppl 1 Text mining for chemistry and the CHEMDNER track):S13.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y. F.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. H.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , K. Y. Feng,
          <string-name>
            <given-names>H. P.</given-names>
            <surname>Li</surname>
          </string-name>
          , et al.
          <article-title>Prediction of drugs target groups based on ChEBI ontology</article-title>
          .
          <source>Biomed Res Int</source>
          .
          <year>2013</year>
          ;
          <year>2013</year>
          :
          <fpage>132724</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Adams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Batchelor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Z.</given-names>
            <surname>Berardini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dietze</surname>
          </string-name>
          , et al.
          <article-title>Dovetailing biology and chemistry: integrating the Gene Ontology with the ChEBI chemical ontology</article-title>
          .
          <source>BMC Genomics</source>
          .
          <year>2013</year>
          ;
          <volume>14</volume>
          :
          <fpage>513</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>V. J.</given-names>
            <surname>Easton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>McCol</surname>
          </string-name>
          .
          <source>Statistics Glossary v1.1</source>
          ,
          <string-name>
            <given-names>chapter</given-names>
            <surname>Sampling</surname>
          </string-name>
          .
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Sunter</surname>
          </string-name>
          .
          <article-title>List Sequential Sampling with Equal or Unequal Probabilities without Replacement</article-title>
          . Applied Statistics.
          <year>1977</year>
          ;
          <volume>26</volume>
          (
          <issue>3</issue>
          ):
          <fpage>261</fpage>
          -
          <lpage>268</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>ChEBI user manual</article-title>
          . Available from: http://www.ebi.ac.uk/chebi/userManualForward.do [accessed 4/22/2016].
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Tanur</surname>
          </string-name>
          .
          <article-title>Margin of Error</article-title>
          . In: International Encyclopedia of Statistical Science. Springer Berlin Heidelberg;
          <year>2011</year>
          :
          <fpage>765</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>SNOMED</given-names>
            <surname>CT</surname>
          </string-name>
          . Available from: https://www.nlm.nih.gov/healthit/snomedct/ [accessed 4/22/2016].
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Halper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Hripcsak,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Perl</surname>
          </string-name>
          , et al.
          <article-title>Analysis of error concentrations in SNOMED</article-title>
          .
          <source>Proceedings of the AMIA 2007 Annual Symposium</source>
          .
          <year>2007</year>
          :
          <fpage>314</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Halper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Perl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          , et al.
          <article-title>Auditing complex concepts of SNOMED using a refined hierarchical abstraction network</article-title>
          .
          <source>J Biomed Inform</source>
          .
          <year>2012</year>
          ;
          <volume>45</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ochs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Perl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Halper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          , et al.
          <article-title>Scalability of abstraction-network-based quality assurance to large SNOMED hierarchies</article-title>
          .
          <source>Proceedings of the AMIA 2013 Annual Symposium</source>
          .
          <year>2013</year>
          :
          <fpage>1071</fpage>
          -
          <lpage>1080</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ochs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Perl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Min</surname>
          </string-name>
          , et al.
          <article-title>Scalable quality assurance for large SNOMED CT hierarchies using subjectbased subtaxonomies</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          .
          <year>2015</year>
          ;
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <fpage>507</fpage>
          -
          <lpage>518</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Perl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Halper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Elhanan</surname>
          </string-name>
          , et al.
          <article-title>The readiness of SNOMED problem list concepts for meaningful use of electronic health records</article-title>
          .
          <source>Artif Intell Med</source>
          .
          <year>2013</year>
          ;
          <volume>58</volume>
          (
          <issue>2</issue>
          ):
          <fpage>73</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hastings</surname>
          </string-name>
          , P. de Matos,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dekker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ennis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Harsha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kale</surname>
          </string-name>
          , et al.
          <article-title>The ChEBI reference database and ontology for biologically relevant chemistry: enhancements for 2013</article-title>
          .
          <source>Nucleic Acids Res</source>
          .
          <year>2013</year>
          ;
          <volume>41</volume>
          (Database issue):
          <fpage>D456</fpage>
          -
          <lpage>463</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19] S. de Coronado,
          <string-name>
            <given-names>M. W.</given-names>
            <surname>Haber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sioutos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Tuttle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. W.</given-names>
            <surname>Wright. NCI Thesaurus</surname>
          </string-name>
          <article-title>: using science-based terminology to integrate cancer research results</article-title>
          .
          <source>Stud Health Technol Inform</source>
          .
          <year>2004</year>
          ;
          <volume>107</volume>
          (Pt 1):
          <fpage>33</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>