<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards standardized evidence descriptors for metabolite an- notations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Schober</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Reza M Salek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Steffen Neumann</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>European Bioinformatics Institute (EMBL-EBI), European Molecular Biology Laboratory, Wellcome Trust Genome Campus</institution>
          ,
          <addr-line>Hinxton, Cambridge, CB10 1SD</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Leibniz Institute of Plant Biochemistry, Dept. of Stress and Developmental Biology</institution>
          ,
          <addr-line>Weinberg 3, 06120 Halle</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Motivation: Data on measured abundances of small molecules from biomaterial is currently accumulating in the literature and in online repositories. Unless formal machine-readable evidence assertions for such metabolite identifications are provided, quality assessment based re-use will be sparse. Existing annotation schemes are not universally adopted, nor granular enough to be of practical use in evidence-based quality assessment. Results: We review existing evidence schemes for metabolite identifications of variant semantic expressivity and derive requirements for a 'compliance-optimized' yet traceable annotation model. We present a pattern-based, yet simple taxonomy of intuitive and self-explaining descriptors that allow to annotate metabolomics assay results both in literature and data bases with evidence information on small molecule analytics gained via technologies such as mass spectrometry or NMR. We present example annotations for typical mass spectrometry molecule assignments and outline next steps for integration with existing ontologies and metabolomics data exchange formats. Availability: An initial draft and documentation of the metabolite identification evidence code ontology is available at https://github.com/DSchober/MIECO. Supplementary material can be found at goo.gl/NCsA7w * Contact: dschober@ipb-halle.de</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <sec id="sec-2-1">
        <title>Background</title>
        <p>
          Metabolomics investigates the distribution and abundance of small
molecules in organisms, mainly applying assay methods like Gas
Chromatography/ Liquid Chromatography (GC/LC), Mass
Spectrometry (MS), Nuclear Magnetic Resonance Spectroscopy (NMR),
Ultraviolet (UV) and Infra Red (IR) spectroscopy. To convert this
analytical data into usable systemic knowledge, identification and
annotation of metabolites in biomaterials is essential
          <xref ref-type="bibr" rid="ref4 ref5">(Creek et al.
2014)</xref>
          , e.g. to indicate that study X provides evidence by assay Y for
occurrence of metabolite Z (in sample Q under condition R).
However, the degrees of confidence in identification statements can
vary greatly between researchers and studies and are difficult to
communicate among users in a crisp, yet concise and accurate
fashion
          <xref ref-type="bibr" rid="ref19 ref5">(Schymanski et al. 2014)</xref>
          . An author's method for reporting the
identification evidence in free text may be dependent on the context
and is usually hard to follow by an external recipient, be it another
scientist or a computational agent like a search engine. Attempts
were made to set up traceable annotation systems to indicate the
quality levels of evidence assignments. Different proposals were put
forward to allow biologists to indicate their identifications with
evidence information in a standardized manner, ranging from
domainspecific simple four level schemes
          <xref ref-type="bibr" rid="ref21">(Sumner et al. 2007)</xref>
          to complex
domain-independent description logic (DL) based ontologies for
automatic evidence reasoning (
          <xref ref-type="bibr" rid="ref2">Bölling et al. 2014</xref>
          ). Yet, none of these
efforts has gained greater momentum so far.
        </p>
        <p>
          The PhenoMeNal data standards workpackage
          <xref ref-type="bibr" rid="ref14 ref8">(PhenoMeNal
website 2016)</xref>
          and the Metabolite Identification Task Group of the
Meta
          <xref ref-type="bibr" rid="ref2">bolomics Society (Creek et al. 2014</xref>
          ) both aim to foster the
development and harmonization of metabolite evidence reporting. As
part of this endeavor, we briefly review existing schemes, identify
their compliance problems and present a domain specific, simple
and compliance-maximized ontology to assist metabolite evidence
assignments.
1.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Overview on existing evidence schemes</title>
        <p>
          Among the minimum reporting standards put forth by the
Metabolomics Standards Initiative
          <xref ref-type="bibr" rid="ref10">(Fiehn et al. 2007)</xref>
          , a four level evidence
scheme was proposed to enable researchers to specify the degree of
confidence in metabolite annotations
          <xref ref-type="bibr" rid="ref21">(Sumner et al. 2007)</xref>
          :




        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Level 1: Confident Identification based on two orthogo</title>
        <p>nal evidences using defined reference standards measured
under identical analytical conditions.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Level 2: Putative Identification based on similar physi</title>
        <p>cochemical properties or library spectra similarities (no
authentic reference standard).</p>
      </sec>
      <sec id="sec-2-5">
        <title>Level 3: Putative Identification of Compound-Class i.e.</title>
        <p>classification based on similar physicochemical properties
or spectral similarity with a compound class.</p>
        <p>
          Level 4: Known Unknowns that are unidentified, yet can
be differentiated and quantified based on spectral data.
This broad numeric classification is proposed for wider usage, e.g.
as recommended assay annotation for the metabolomics journal
          <xref ref-type="bibr" rid="ref13">(Metabolomics Journal, 2016)</xref>
          . Although easy to use, its drawback
is a lack of granularity in the detail of what evidences can be
expressed in a formal and search engine-friendly manner. This lack of
utility might be the reason why it has been sparsely adopted by the
metabolomics community
          <xref ref-type="bibr" rid="ref9">(Everett 2016)</xref>
          . Realizing these
drawbacks, and based on an earlier suggestion of adding computable
numeric indicators
          <xref ref-type="bibr" rid="ref4 ref5">(Creek et al. 2014)</xref>
          , Sumner et al (2014) proposed
a more granular assay-technology centered scheme, allowing to
assert numeric weights to its granular evidence components, resulting
in an additive quantitative identification score. At the same time, the
earlier four level scheme was expanded
          <xref ref-type="bibr" rid="ref2">by Schymanski et al. (2014</xref>
          ),
providing an enriched scheme with five identification levels that are
accompanied by minimal assay data requirements. Level 1 to 3 map
to the Sumner levels, but provide a better granularity, with 2a and
2b distinguishing different sets of evidence in absence of a reference
standard, whereas level 4 is an addition describing formula based
annotations and Level 5 describes the known unknowns via an exact
mass.
        </p>
        <p>
          Besides the aforementioned domain-specific approaches, advanced
ontological proposals of more general, domain-independent nature
have been introduced. These leverage on automated evidence
reasoning, based on description logic (DL) semantics, with
axiomatizations by means of fine-grained DL patterns: The Evidence Code
Ontology (ECO), (Chi
          <xref ref-type="bibr" rid="ref2">bucos et al. 2014</xref>
          ) applies a basic pattern around
the notions ‘Assertion method’ and ‘Evidence’. At a recent
workshop
          <xref ref-type="bibr" rid="ref7 ref8">(ECO-OBI Workshop 2016)</xref>
          it became evident that besides
domain-specific coverage gaps, restructuring is required to allow for
OBI
          <xref ref-type="bibr" rid="ref1">(Bandrowski et al. 2016)</xref>
          usage. The transition from a simple
enumerated evidence list towards axiomatised reasoning for
autoclassification is reflected in the recent name change: ‘Evidence and
Conclusion Ontology’
          <xref ref-type="bibr" rid="ref14 ref7 ref8">(ECO website 2016)</xref>
          . Diverging from its
initial Gene Ontology inspired simple set-up, this effort now competes
with complex DL-heavy approaches like the Semantic Evidence
(SEE) reasoning ontology (
          <xref ref-type="bibr" rid="ref2">Bölling et al. 2014</xref>
          ) and there is danger
of development slow-down due to the added complexity. Coverage
is currently sparse on metabolomics technologies, i.e. important top
level terms like
'mass spectrometry evidence used in automatic
assertion' EquivalentTo:
'mass spectrometry evidence' AND
        </p>
        <p>(used_in SOME 'automatic assertion')
are missing. An analysis of other existing domain ontologies
revealed sparse coverage for terms used in metabolite identification
and distribution across multiple namespaces (see Supplementary
Material), making it necessary to build a coherent artefact1. As ECO
still is expected to gain a greater user base, we decided to re-use this
ontology by importing and expanding it and leveraging on its basic
evidence pattern.
2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>METHODS</title>
      <p>
        To allow for granular coverage, we apply an ontology driven
approach to generate descriptors for small molecule evidence
assignments, covering all major assay methods in metabolomics. We
imported ECO in Protégé 4.3 and started our additions in an OWL
ontology called Metabolite Identification Evidence Code Ontology
(MIECO). OBI-core2 was imported to gain BFO 2, the OBI upper
level and core RO object properties. These are utilized to increase
the granularity of assertion modes, i.e. providing ontology design
patterns for composite assertion specifications. As we also aim for
backwards compatibility to the
        <xref ref-type="bibr" rid="ref21">Sumner (2007)</xref>
        scheme, we intend to
employ DL reasoning to automatically map and classify MIECO
evidences onto earlier evidence schemes.
      </p>
      <p>
        In a first round of ontology design, we derived terms in a data
driven bottom-up approach from an in-house LC-MS use case
(UPLC/ESI-QTOF MS) based on an untargeted analysis of
semipolar root exudate metabolites3 of Arabidopsis thaliana
        <xref ref-type="bibr" rid="ref20 ref5">(Strehmel et
al. 2014)</xref>
        . Based on the natural language identification statements in
1 Term imports and cross-referencing methodologies like MIREOT in
OntoFox are still fragile and often provide little help in modularisation and
domain border decisions.
      </p>
      <p>
        2 http://obi-ontology.org/page/Core
this paper, we generated an enriched list of around 20 evidence
descriptors. (see Supplemental material). Additional input were the
alphanumeric terms from Tab.2 in
        <xref ref-type="bibr" rid="ref22">Sumner et al. (2014)</xref>
        , which lists
the most common descriptors proposed as identification indicators,
distributed among the five most prominent assay methods. From
that, we started expansion by first identifying an intuitive lexical
pattern that is easily understandable by our end users and that covers
most of the classes generated from our initial paper text. This will
form the basis for later reconfiguration into DL patterns and
automatic pattern-based term creation via TermGenie
        <xref ref-type="bibr" rid="ref5">(Dietze et al.
2014)</xref>
        , which assists pragmatic pre-coordination of only those terms
really required in practice. We exemplarily annotate assignments
from Table 2 in the LCMS use case paper
        <xref ref-type="bibr" rid="ref20 ref5">(Strehmel et al. 2014)</xref>
        over
a range of evidence levels.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>RESULTS</title>
      <p>Our analysis indicates that the simple four level scheme in use today
is not granular enough to provide reliable quality indicators and
foster quality and trust analysis for metabolite annotation. Formal DL
approaches, on the other hand, tend to develop slowly and shielding
their complexity from users has been a known problem, preventing
end user compliance and community growth. Here we strive for a
pragmatic middle way, starting with an intuitive taxonomy of data
driven and frequently used descriptors, and later exploiting DLs
combinatorial semantics to generate axioms and additional evidence
terms from patterns where required via TermGenie, as described
above.</p>
      <p>We provide a first draft of a taxonomy of pre-coordinated
descriptors for metabolite identification. In contrast to emerging
DLheavy approaches, our scheme is less complex, as we focus on
simple and easy to use terms for maximum end user compliance.
MIECO currently consists of ~ 90 terms, axiomatization being
sparse at this early stage.
3.1</p>
      <sec id="sec-4-1">
        <title>Pattern proposals</title>
        <p>Metabolite identification refers to assertions that support that a
compound under investigation is either of a certain totally defined
molecular structure (Identification on the leaf node universal level) or
is of a certain backbone structure (annotation on the superclass
level). We therefore structure the main evidence axis according to
the following assertion taxonomy and compositional pattern:
Assertion e.g. comprising a taxonomy of
Annotation/Characterisation/Classification/Identification
of
Molecular structural element e.g. Molecule, class, functional
group, element, Isotope
by (
Assay Outcome e.g. Assay outcome with sub-assay details e.g.
MS, MS2, LC/RT, Isotope data, adduct data, precursor
(quantifier) ion type
used in
Assertion method e.g. Run against reference standard,
Comparison to reference database, Author inference)</p>
        <sec id="sec-4-1-1">
          <title>3 The data is available as MTBLS160 in</title>
          <p>http://www.ebi.ac.uk/metabolights/reviewerLgTnoHUrFb
Metabolights at
We currently try to render this linguistic pattern compatible with
ECO, but we hope the patterns will still be intuitive to understand
by peer users. The part in round brackets is of 1:n cardinality.</p>
          <p>
            We strive for an end-user compliant naming pattern for future
MIECO class names. MIECO labels are currently generated in a
bottom up manner as extracted from the use case. The relation of the
naming pattern and the MIECO term labels will be, that in a future
release we can add the formally consistent pattern-generated
TermGenie labels as alternative labels to existing user-preferred
            <xref ref-type="bibr" rid="ref18">(Schober et al. 2009)</xref>
            MIECO labels. The difference to the above
lexical pattern is that the Assertion taxonomy has been transformed
into an assertion relation (object property) hierarchy:
asserted by (annotated by)
characterised by
classified by
          </p>
          <p>identified by</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>Overall evidence naming pattern: MolecularStructureElement [annotation relation] AssayOutcome used_in AssertionMethod</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>Example Evidence Class: Guanosine identified_by ‘LCMS fragmentation pattern’ used_in ‘similarity to authentic reference standard’</title>
          <p>3.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Example annotations and mappings</title>
        <p>
          Table 1 exemplarily shows evidence annotations via MIECO terms
for different feature assignments from the use case paper
          <xref ref-type="bibr" rid="ref20 ref5">(Strehmel
et al. 2014, modified after supplementary Table 2 therein)</xref>
          and
compares these to earlier evidence schemes.
        </p>
        <p>
          Aside these encompassing overall assignment annotations,
MIECO terms can also be used to annotate on a highly granular
level, i.e. annotating each evidence contributor for a whole set of
experimental LCMS characteristics of a molecule. E.g. for feature
#2 (Guanosine), we can assign






‘MIECO_0000094: Characterisation by sum formula’,
annotating the Elem. Comp. of C10H13N5O5.
‘MIECO_0000005: Characterisation by LC RT similarity’,
annotating the Retention time 46s
‘MIECO_0000012: Characterisation by online HD exchange
experiment identified substructure revealing exchangeable
protons’, annotating the evidence of the 6 exchange protons
matching the H-functional groups.
‘MIECO_0000016: Characterisation by collision induced
dissociation (CID) MS2 with mass and isotope pattern of
quasi-molecular fragment ion in negative ESI mode’, annotating the
[M-H]Quantifier Ion.
‘MIECO_0000009: Characterisation by m/z value in MS1’,
annotating the 282.08 m/z of the parent ion.
‘MIECO_0000010: Characterisation by fragmentation pattern in
MS2’, annotating the MS2 fragment ions m/z 150,133.
Reliable metabolite identification is not easy to achieve and
communicate in a traceable, reproducible manner. Among the reasons is
that identification assertions made by humans often embrace
implicitness, e.q. as a consequence of the complexity of the underlying
state of affairs. The reason for this is the combinatorial cross
dependencies of the steps in a metabolite identification process. This
can consist of multiple singular assertions that act together in a
nonlinear synergistic fashion to ultimately produce an evidence. This is
why highly granular models are required and ontologies are
currently the best formalisms to capture such detail. If the granular
subprocesses of an identification process are not explicitly named, the
danger of subjectiveness arises and unreliable identification
measures can decrease not only scientific resolution and
interpretation, but may also blur further processing, re-use and knowledge
generation. However, the reliability is often not additive/linear, but
rather an emerging property of inference chains. Multiple per se
anecdotal evidences can result in a strong evidence, because the parts
are reinforcing each other. Emergence is unfortunately not easy to
be captured in Boolean symbolic formalisms like ontologies
          <xref ref-type="bibr" rid="ref11">(Haken
H., Schober D. 2008)</xref>
          , but as a proxy for such knowledge could be
introduced by assigning weights to single MIECO evidences and
then let a rule-based system judge/compute on the overall evidence
of the combined evidence contributors.
        </p>
        <p>Also, as molecular structures are an important part of our
annotation pattern, we have to consider at what level in a
generalization/specialization taxonomy a distinction between class and
individual is made4.</p>
        <p>
          As next steps we need to increase the amount of pre-coordinated
terms beyond use case coverage and we plan to foster interaction
and cross play with ECO. The addition of weighted numeric
information as proposed in
          <xref ref-type="bibr" rid="ref22">Sumner et al. (2014)</xref>
          will later lead to a
quantified scoring scheme. We currently investigate usage within
community resources like Metabolights
          <xref ref-type="bibr" rid="ref15">(Haug et al. 2013)</xref>
          and
integration into Galaxy Workflows. We currently test curator compliance
by annotating experiment entries within Metabolights5 with MIECO
terms, i.e. test if users apply MIECO correctly. Suitability for use in
existing metadata representation systems like ISA
          <xref ref-type="bibr" rid="ref16">(Sansone et al.
2012)</xref>
          and format exchange data standards like mzTab or
mzIdentML
          <xref ref-type="bibr" rid="ref12">(Jones et al. 2012)</xref>
          are analysed. E.g. mzTab expanded
the
          <xref ref-type="bibr" rid="ref21">Sumner et al. (2007)</xref>
          scheme with added identification
reliabilities: Small molecule identifications reported in an mzTab file can be
assigned a reliability, reported as an integer between 1-3 for
proteomics results (1: high reliability, 2: medium reliability, 3: poor
reliability). This is easy to use but still untraceable insofar as these
indicators cannot be derived from provided granular metadata.
        </p>
        <p>As a next step we need to investigate curator tools for assisted
semi-manual annotation, as well as the best storage places for
metabolite evidence metadata. We will also look into text mining and
embedding into annotation tools for computer assisted high
throughput annotation, rendering an ever increasing amount of data
accessible to evidence-based threshold filtering for quality data re-use.
Here the transition to a quantitative background model for numeric
evaluation is necessary.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSION</title>
      <p>
        Our initial MIECO draft is domain-specific use case (bottom up)
driven vocabulary for annotating metabolite assignments with assay
specific evidence terms. It was designed in a pragmatic manner,
attempting to shield the end-users from DL axiomatizations and
subscribing to complexity reduction strategies descri
        <xref ref-type="bibr" rid="ref2">bed earlier
(Schober et al. 2014</xref>
        ). Although in an early stage, we hope this domain
restriction and its simple design, reflecting a lexical design pattern
that is intuitive to a domain specialist, will contribute to reducing
development times in the future. Our ontology-based metabolite
identification and evidence scheme could become a handy asset in
judging data provenance and reliability of identification assertions,
i.e. allowing to set confidence thresholds for search and retrieval
4 i.e. is a metabolite fully identified as Leucine? Or is there a need
to further disambiguate the analyte from Isoleucine ? Are trans fatty
acids the same as cis fatty acids? ChEBI does not always separate
class from instance level.
tasks. When more mature, this annotation scheme could gain
momentum in the larger metabolomics community and is envisioned to
contribute to a more traceable catalog of descriptors for small
molecule assignment evidence.
      </p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGEMENT</title>
      <p>We like to thank the Baltimore ECO workshop participants, Nadine
Strehmel, Christoph Ruttkies and the Metabolite Identification task
group of the Metabolomics Society for inspiration and feedback.
This work has been financed by the EC Horizon 2020 project
PhenoMeNal, grant agreement number 654241.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Bandrowski A.</given-names>
            ,
            <surname>Brinkman</surname>
          </string-name>
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Brochhausen</surname>
          </string-name>
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Brush</surname>
          </string-name>
          <string-name>
            <given-names>M.H.</given-names>
            ,
            <surname>Bug</surname>
          </string-name>
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Chibucos</surname>
          </string-name>
          <string-name>
            <surname>M.C.</surname>
          </string-name>
          , et al. (
          <year>2016</year>
          )
          <article-title>The Ontology for Biomedical Investigations</article-title>
          .
          <source>PLoS ONE</source>
          <volume>11</volume>
          (
          <article-title>4): e0154556</article-title>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.0154556
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Bölling C.</given-names>
            ,
            <surname>Weidlich</surname>
          </string-name>
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Holzhütter</surname>
          </string-name>
          <string-name>
            <surname>H.G.</surname>
          </string-name>
          (
          <year>2014</year>
          ),
          <article-title>SEE: structured representation of scientific evidence in the biomedical domain using Semantic Web techniques</article-title>
          .
          <source>J Biomed Semantics; 5(Suppl</source>
          <volume>1</volume>
          ): S1. doi:
          <volume>10</volume>
          .1186/2041-1480-5-S1-S1 http://www.jbiomedsem.com/content/5/S1/S1
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Chibucos</surname>
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mungall</surname>
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balakrishnan</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christie</surname>
            <given-names>K.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huntley</surname>
            <given-names>R.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blake</surname>
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giglio</surname>
            <given-names>M.</given-names>
          </string-name>
          , (
          <year>2014</year>
          ),
          <article-title>Standardized description of scientific evidence using the Evidence Ontology (ECO)</article-title>
          .
          <source>Database</source>
          .
          <year>2014</year>
          ,
          <year>2014</year>
          :
          <fpage>bau075</fpage>
          -
          <lpage>10</lpage>
          .1093/database/bau075, http://www.ncbi.nlm.nih.gov/pmc/articles/PMC4105709
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Creek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunn</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fiehn</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Griffin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Hall,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Mistrik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Schymanski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            ,
            <surname>Sumner</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          , et al. (
          <year>2014</year>
          ),
          <article-title>Metabolite identification: are you sure? And how do your peers gauge your confidence?</article-title>
          <source>Metabolomics</source>
          ,
          <volume>10</volume>
          , pp.
          <fpage>350</fpage>
          -
          <lpage>353</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Dietze H</surname>
          </string-name>
          . et al. (
          <year>2014</year>
          ),
          <article-title>Termgenie-a web-application for patternbased ontology class generation</article-title>
          .
          <source>J. Biomed. Semant.</source>
          ,
          <volume>5</volume>
          , 48, http://bioinformatics.oxfordjournals.org/external-ref?access_num=
          <volume>10</volume>
          .1186/2041-1480-5-48&amp;link_type=DOI
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Dunn</surname>
            <given-names>WB</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erban</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            <given-names>RJM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Creek</surname>
            <given-names>DJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brown</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Breitling</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hankemeier</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goodacre</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kopka</surname>
            <given-names>J</given-names>
          </string-name>
          . et al. (
          <year>2013</year>
          ),
          <article-title>Mass appeal: metabolite identification in mass spectrometry-focused untargeted metabolomics</article-title>
          .
          <source>Metabolomics</source>
          .
          <year>2013</year>
          ;
          <volume>9</volume>
          :
          <fpage>p44</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>ECO-OBI workshop</surname>
          </string-name>
          (
          <year>2016</year>
          ); https://docs.google.com/document/d/1Y-gxKHpHAyS_ngiAS1NkBLT4lttPAOx8ECHkCf9kSw/edit#heading=h.r8a1r3t5hcrz,
          <source>accessed 21.6</source>
          .2016
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>ECO website</surname>
          </string-name>
          (
          <year>2016</year>
          ) http://www.evidenceontology.
          <source>org/; accessed 3.9</source>
          .2016
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Everett</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , (
          <year>2016</year>
          ),
          <article-title>New NMR-Based Methods for Assessing Confidence in Known Metabolite Identification, Metabolomics Spotlight Article</article-title>
          .
          <source>In MetaboNews website May</source>
          <year>2016</year>
          , http://www.metabonews.ca/May2016/MetaboNews_May2016.htm#spotlight ,
          <source>last accessed 17.6</source>
          .2016
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Fiehn O.</given-names>
            ,
            <surname>Robertson</surname>
          </string-name>
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Griffin</surname>
          </string-name>
          <string-name>
            <given-names>J</given-names>
            .,
            <surname>van der Nikolau</surname>
          </string-name>
          <string-name>
            <surname>B</surname>
          </string-name>
          , et al. (
          <year>2007</year>
          ),
          <article-title>The metabolomics standards initiative (MSI) Metabolomics</article-title>
          .
          <year>2007</year>
          ;
          <volume>3</volume>
          :
          <fpage>175</fpage>
          -
          <lpage>178</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11306-007
          <source>-0070-6</source>
          .
          <article-title>5 Metabolights is a dedicated open data repository for metabolomics ex-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Haken H.</given-names>
            ,
            <surname>Schober</surname>
          </string-name>
          <string-name>
            <surname>D.</surname>
          </string-name>
          (
          <year>2008</year>
          ),
          <source>personal email communication. 25.8</source>
          .2008
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Jones</surname>
            <given-names>AR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisenacher</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mayer</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kohlbacher O</surname>
          </string-name>
          . et al. (
          <year>2012</year>
          ),
          <article-title>The mzIdentML data standard for mass spectrometry-based proteomics results</article-title>
          .
          <source>Molecular and Cellular Proteomics: MCP</source>
          .
          <year>2012</year>
          ;
          <volume>11</volume>
          (
          <issue>7</issue>
          ):
          <fpage>M111</fpage>
          014381. doi:
          <volume>10</volume>
          .1074/mcp.
          <source>M111</source>
          .014381.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Metabolomics</given-names>
            <surname>Journal</surname>
          </string-name>
          , Instructions for authors (
          <year>2016</year>
          ), http://www.springer.com/life+sciences/biochemistry+%26+biophysics/journal/11306?detailsPage=pltci_1709154#,
          <source>accessed 19.6</source>
          .2016
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>PhenoMeNal website</surname>
          </string-name>
          (
          <year>2016</year>
          ) http://phenomenal-h2020.eu/home, accessed
          <volume>17</volume>
          .6.2016
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Haug</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salek</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Conesa</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hastings</surname>
            <given-names>J.</given-names>
          </string-name>
          , et al. (
          <year>2013</year>
          ),
          <article-title>MetaboLights-an open-access general-purpose repository for metabolomics studies and associated meta-data</article-title>
          .
          <source>Nucl. Acids Res</source>
          .
          <volume>41</volume>
          (
          <issue>D1</issue>
          ):
          <fpage>D781</fpage>
          -
          <lpage>D786</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gks1004
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Sansone S.A.</given-names>
            ,
            <surname>Rocca-Serra</surname>
          </string-name>
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Field</surname>
          </string-name>
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Maguire</surname>
          </string-name>
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Taylor</surname>
          </string-name>
          <string-name>
            <surname>C.</surname>
          </string-name>
          , et al. (
          <year>2012</year>
          ),
          <article-title>Toward interoperable bioscience data</article-title>
          .
          <source>Nature Genetics</source>
          ,
          <volume>44</volume>
          ,
          <fpage>121</fpage>
          -
          <lpage>126</lpage>
          (
          <year>2012</year>
          ), doi:10.1038/ng.1054
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Schober D.</given-names>
            ,
            <surname>Boeker</surname>
          </string-name>
          <string-name>
            <surname>M.</surname>
          </string-name>
          (
          <year>2010</year>
          ),
          <article-title>Ontology Simplification: new buzzword or real need? OBML 2010 Workshop Proceedings</article-title>
          . Edited by:
          <string-name>
            <surname>Herre</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoehndorf</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelso</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <article-title>Institut fuer Medizinische Informatik, Statistik und Epidemiologie (IMISE)</article-title>
          ,
          <source>Markus Loeffler</source>
          .
          <year>2010</year>
          ,
          <fpage>M1</fpage>
          -
          <lpage>5</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Schober D.</given-names>
            ,
            <surname>Smith</surname>
          </string-name>
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Lewis</surname>
          </string-name>
          <string-name>
            <given-names>S.E.</given-names>
            ,
            <surname>Kusnierczyk</surname>
          </string-name>
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Lomax</surname>
          </string-name>
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Mungall</surname>
          </string-name>
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Taylor C</surname>
          </string-name>
          .F.,
          <string-name>
            <surname>Rocca-Serra</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sansone</surname>
            <given-names>S-A.</given-names>
          </string-name>
          (
          <year>2009</year>
          ),
          <article-title>Survey-based naming conventions for use in OBO Foundry ontology development</article-title>
          .
          <source>BMC bioinformatics</source>
          <year>2009</year>
          ,
          <volume>10</volume>
          :
          <fpage>125</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Schymanski</surname>
            ,
            <given-names>E. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jeon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gulde</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fenner</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruff</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>H. P.</given-names>
          </string-name>
          , et al. (
          <year>2014</year>
          ),
          <article-title>Identifying small molecules via high resolution mass spectrometry: communicating confidence</article-title>
          .
          <source>Environmental Science and Technology</source>
          ,
          <volume>48</volume>
          (
          <issue>4</issue>
          ),
          <fpage>2097</fpage>
          -
          <lpage>2098</lpage>
          . doi:
          <volume>10</volume>
          .1021/es5002105
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Strehmel N.</given-names>
            ,
            <surname>Bottcher</surname>
          </string-name>
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Schmidt</surname>
          </string-name>
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Scheel</surname>
          </string-name>
          <string-name>
            <surname>D.</surname>
          </string-name>
          (
          <year>2014</year>
          ),
          <article-title>Profiling of secondary metabolites in root exudates of Arabidopsis thaliana</article-title>
          .
          <source>Phytochemistry</source>
          <volume>108</volume>
          :
          <fpage>35</fpage>
          -
          <lpage>46</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Sumner L.W.</given-names>
            ,
            <surname>Amberg</surname>
          </string-name>
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Barrett</surname>
          </string-name>
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Beale</surname>
          </string-name>
          <string-name>
            <surname>M.</surname>
          </string-name>
          et al. (
          <year>2007</year>
          ),
          <article-title>Proposed minimum reporting standards for chemical analysis</article-title>
          .
          <source>Metabolomics</source>
          .
          <year>2007</year>
          ;
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <fpage>211</fpage>
          -
          <lpage>221</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11306-007- 0082-2.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Sumner L. W.</surname>
          </string-name>
          , Lei
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Nikolau</surname>
          </string-name>
          <string-name>
            <given-names>B.J.</given-names>
            ,
            <surname>Saito</surname>
          </string-name>
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Roessner</surname>
          </string-name>
          <string-name>
            <given-names>U.</given-names>
            ,
            <surname>Trengove</surname>
          </string-name>
          <string-name>
            <surname>R.</surname>
          </string-name>
          (
          <year>2014</year>
          ):
          <article-title>Proposed quantitative and alphanumeric metabolite identification metrics</article-title>
          .
          <source>Metabolomics</source>
          <volume>10</volume>
          :
          <fpage>1047</fpage>
          -
          <lpage>1049</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11306-014-0739-6.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>