<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Powered by Vertabelo, Design Your Database Online, http://vertabelo.com
gcm_lkb_expert</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Ontology-Driven Metadata Enrichment for Genomic Datasets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Colombo</string-name>
          <email>andrea55.colombo@mail.polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Ceri[</string-name>
          <email>stefano.cerig@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano</institution>
          ,
          <addr-line>Via Ponzio 34/5, 20133, Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>201</volume>
      <fpage>8</fpage>
      <lpage>09</lpage>
      <abstract>
        <p>Data-driven genomic research requires accessing several repositories of genomic datasets, produced by international consortia, which provide open access to extremely valuable and well curated biological content. The associated metadata, describing experimental and biological conditions, are highly heterogeneous; consequently, dataset collection and integration is di cult { it requires data conversions and term matching which needs to be done by humans, with biological expertise. In this paper, we present a method and tools for ontology-driven metadata enrichment. We select few relevant features which are provided by most repositories, and then we comparatively evaluate several search services providing ontological access, eventually associating each feature with the speci c ontologies which are most suited to describe them. We also provide an expert validation of the approach. The method and tools are deployed in a large repository of open data, which will be soon available to the research community.</p>
      </abstract>
      <kwd-group>
        <kwd>Data Integration Genomic Datasets tion Open Data Bioinformatics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>With the growth of diversity and complexity of scienti c databases, the role of
metadata { describing their content and data production process { is
becoming more relevant. In particular, genomic computing often requires collecting
datasets from multiple heterogeneous sources; unfortunately, metadata
describing datasets across such sources are structured di erently, they are often
incompatible or incomplete. This raises a huge problem of data integration, which
can be solved through ontological mediation, bridging the sources and enabling
metadata interoperability.</p>
      <p>In this paper, we describe metadata enrichment, which is the process of
annotating existing structured metadata with ontological terms, their de nitions,
synonyms, ancestors, and descendants, to instrument a semantically enriched
search of datasets linked to such metadata. Metadata enrichment is performed
? These two authors contributed equally to this work.
at the end of a data integration procedure for data loading, cleaning and mapping
that is outside of the scope of this paper.</p>
      <p>
        Metadata are converted to t the format of a Genomic Conceptual Model
(GCM, [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]), gathering the most important properties shared between
heterogeneous sources. GCM is centered on the item entity (representing an
experimental unit stored as a le of genomic regions) and organized as a four-pointed-star
whose parts describe connected aspects about biology, technology, extraction,
and management of the item.
      </p>
      <p>Among all GCM attributes, we call \ontological" the ones that are present
in all sources and require ontological agreement, thus are worthy of enrichment.
These are: the Platform of items, i.e., the NGS platform used for sequencing,
the Ethnicity and Species of donors, i.e., the individual of the organism from
which the biological material is derived; the Disease, storing information about
the pathology investigated with the sample; the Tissue and CellLine of
samples, which distinguish the kind of biological material used for the experiment;
the Technique or assay used to produce the genomic experiment (e.g.,
\Chipseq", \miRNA-seq", \Genotyping Array"); the speci c Feature or aspect
described by the experiment (e.g., \Copy Number Variation",\Histone Modi
cation",\Transcription Factor"); and the Target gene or protein of experiments
(e.g., \CTCF", \MYC").</p>
      <p>
        In this paper, we propose a metadata enrichment system, speci c for genomic
datasets, with a four-fold contribution: 1. description of the existing services to
search ontologies related to biomedical content; 2. scoring and selection of the
service and ontologies most relevant for our data; 3. organization of ontological
knowledge in a well-structured taxonomy; 4. production of a tool for ontological
annotations extraction. As a rst integration e ort, we include three important
data sources used in the genomic research community, namely: Genomic Data
Commons (GDC, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]), with over 310,000 les covering aspects of cancer genomics;
the Encyclopedia of DNA Elements (ENCODE, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]), with almost 420,000 les
related to functional DNA sequences and regulatory elements controlling gene
expression; Roadmap Epigenomics Project (REP, [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]) containing around 2,000
datasets related to genetic variation.
      </p>
      <p>The paper is structured as follows. Section 2 presents our solution to the
problem of selecting appropriate search services and ontologies to annotate metadata.
Section 3 describes how the enrichment procedure works and how we validated
the process. Section 4 overviews related work, and nally Section 5 concludes
the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Search Service and Ontology Selection</title>
      <p>First, we present the four most used and well-known ontology search services in
literature (see Section 2.1), and how we score them (see Section 2.2) in order
to select the most appropriate search service for our project (see Section 2.3).
Next, we compare the ontologies provided by that search service, and select the
speci c ontology that is most suitable to annotate values for each ontological
attribute (see Section 2.4).</p>
      <sec id="sec-2-1">
        <title>Ontology Search Services</title>
        <p>Ontological access to genomic data is well supported by several search services,
which are capable in turn to integrate a high number of ontologies. Therefore,
we are initially concerned in choosing the best search service, that will then be
used within our system as broker to the underlying ontologies. We consider four
di erent search services, which appear suitable for our purpose.</p>
        <p>
          BioPortal [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] is a repository of biomedical ontologies and terminologies
whose access is provided through a Web portal and Web services. We exploit its
term search service, an endpoint which takes a free text input and provides a
result in json format, listing a (con gurable) number of annotations to
ontological terms, showing di erent degrees of matching with the free text. These can
be considered as possible annotations for the input text. Each term is
identied by the pair hontology; idi, describing the code which references the ontology
inside the BioPortal system and an identi cation number which references the
term inside the ontology. A term also contains a single preferred label and its
synonyms. An annotation is composed by a term and a match type: \PREF" if
the match with the term is established with the preferred label or \SYN" if the
match is with one of the term synonyms.
        </p>
        <p>
          Ontology Recommender [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] is a BioPortal service that receives a free
text or a list of keywords and suggests a set of ontologies appropriate for
annotating the indicated terms, considered all together. The structure of annotations
is identical to BioPortal's. Additionally, Recommender provides four scores that
re ect how well the ontology (set) annotates the input data: Coverage, measures
with which extent the ontology represents the input; Acceptance, indicates how
well-known and trusted the ontology is by the biomedical community; Detail,
shows the level of speci cation provided by the ontology for the input data;
Specialization, indicates how specialized the ontology is w.r.t the input data
domain.
        </p>
        <p>
          Ontology Lookup Service (OLS, [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]) provides ontology search,
visualization, and ontology-based services. The accepted input is a keyword, the
provided result is a list of annotations, similar to the other services but not
including a match type. In the API request, a eldList parameter can be used
to specify the speci c elements to be included in the output along with other
formatting preferences.
        </p>
        <p>Zooma1 is a service from OLS which provides mappings between textual
input and a manually curated repository of text-to-ontology-term mappings. If
no mappings are found, it uses the basic OLS search. In addition to the usual
annotation information, Zooma also returns a con dence label associated to the
annotation, ranging from HIGH to LOW.</p>
        <p>
          We exclude other important ontology search portals such as HeTOP [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and
UMLS [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], as they are more focused on multilingual support and medical
terminologies, therefore do not include many ontologies that are important to
annotate our values. Also the NCBO Annotator [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is not considered since its
functionalities are completely covered by the Ontology Recommender.
        </p>
        <sec id="sec-2-1-1">
          <title>1 https://www.ebi.ac.uk/spot/zooma/</title>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Scoring</title>
        <p>Every search service provides a search API, which is repeatedly used for the
score evaluation. For each API call we store: the used service; the attribute from
GCM characterizing the values (the \type" of the values); the original raw value
deriving from the GCM, imported through the mapping phase; possible parsed
values deriving from a simple syntactic pre-processing of raw values (e.g., removal
of punctuation, split of long expressions. . . ); the hontology,ontology idi pair,
uniquely identifying an ontological term in a service; pref label and synonym,
respectively the textual expression primarily used for the term and its alternative
versions; score, textual information regarding the goodness of a match, directly
retrieved from the services, if available.</p>
        <p>In total, we performed 1,783 API calls to each of the four services,
corresponding to 1,299 original values to be enriched; some of these were splitted during
a pre-processing phase. As a result, we retrieved 1,783 interesting matches from
BioPortal, 885 from Recommender, 1,782 from OLS, and 1,779 from ZOOMA,
all of which were used for the following processing after calculating our scores.</p>
        <p>Starting from the retrieved information, we calculate the match score as a
measure of how well a term matches a value, by using a scoring system that is
speci cally designed for the task, which is next described. The general formula
returning the match score value, shown in Eq. 1, subtracts from an initial
maximum number (10, when there is a perfect match with a pref label, 9 with a
synonym) a penalty measuring how the raw value di ers from the label retrieved
from the services:
match score(raw; label) = f10; 9g
distance(raw; label)
(1)</p>
        <p>
          To compute the distance, we use a modi ed version of Needleman-Wunsch
algorithm [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], a protein and nucleotide sequence alignment algorithm which is
widely used in bioinformatics. In the original algorithm, the input is represented
by two strings whose letters need to be aligned. The letters may have a \match",
a \mismatch" or an \indel" (i.e., adding a gap in one of the strings). In our
modi ed version, we de ne each word as a distinct letter of the original algorithm
and we add another type of mismatch, i.e., the swap. All in all, the total distance
is calculated as a sum of distances between words:
{ Match: Two words are the same, then their distance is 0
{ Mismatch: Two words are di erent, then their distance is 2.5
{ Swap: Two consecutive words traded places, then their distance is 0.5
{ Delete: One word is deleted from the raw, then their distance is 2
{ Insert: A new word is added to the raw then their distance distance is 1
The indicated distance values are chosen in such a way that the number of
deletions is minimized (i.e., we penalize a label which does not include a word
present in raw ) and the swap is preferred to indel and mismatch. For example,
for the raw \breast invasive carcinoma", the label \invasive breast carcinoma"
(i.e., one swap) is considered better than \breast carcinoma" (i.e., one deletion).
        </p>
        <p>
          Additional calculated scores are: onto suitability, a measure of how
much an ontology is adequate for a given attribute, calculated as the average
match score over all raw values for that attribute; onto acceptance, a measure
of how well-known and trusted the ontology is by the biomedical community,
computed through Recommender Web Services [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]2; the overall score,
obtained by multiplying each raw value match score by a weighted average of the
two measures typical of the ontology.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Service Evaluation</title>
        <p>Table 1 describes the obtained results. The \Service Properties" part contains an
overview of service properties. BioPortal and Recommender provide a match type
(MT) in their APIs response, which means that they specify if the input text is
more similar to the preferred label rather than to one of the synonyms associated
to a term. Recommender o ers the additional function of searching for multiple
key-words at the same time (MK) and consequently suggests a minimal set of
ontologies suitable for annotating the maximum possible number of key-words.
This function is also o ered by ZOOMA which, however, in practice just
performs multiple single key-word requests and lists all results at the same time.
Only Recommender executes a good attempt of annotating free texts (FT).
BioPortal's set of ontologies is much broader than OLS' since minor e orts are also
included. ZOOMA exploits search results from OLS but also provides results
coming from previous manual curation works as an additional service to the
user.
Aggregated scores COocvcuerraregnece</p>
        <p>
          The \Example scoring" part contains an example of how services are
rewarded based on the matching terms they nd. To evaluate the match, we use
the overall score described in Section 2.2. When the disease-related text
\cervical adenocarcinoma" is searched, BioPortal suggests, on top of others, the three
terms \ncit c4029", \efo 0001416", and \doid 3702", while Recommender just
2 It is derived from the number of visits to the ontology page in BioPortal and the
presence or absence of the ontology in UMLS [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
provides one result, \ncit c4029". Our algorithm for Occurrence computes the
set of terms which occur the highest amount of times in the top three matches
of the services (in this case [\ncit c4029",\efo 0001416"]) and assigns a weighted
reward (1 if the set only contains one entry, 0.5 if it contains 2, and so on) to the
services which include that term in the top results. Indeed BioPortal scores 1
since it contains both top results, while Recommender scores 0.5 since it contains
just one. Coverage is 1 when the service provides at least one result, 0 otherwise.
        </p>
        <p>We use as scores for service selection the average Occurrence and Coverage
over all the searched raw values. On this basis, OLS is selected as the best suited
search service to pursue the enrichment annotations in our system.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Ontology Selection</title>
        <p>Based on the overall score described in Section 2.2, we also aggregate results
over speci c attributes and ontologies. This calculation produces, as a result,
one top ontology for each attribute. Since most of the times only one ontology
does not provide an acceptable coverage for all the values belonging to that
attribute, we use an algorithm to compute a small set of ontologies to annotate
values from an attribute. Such algorithm rst tries to match values only with
the rst ontology, then tries to match only the ones left unmatched with the
following ontologies, until a xed point for coverage is found. If the computational
costs become too high, the algorithm can be stopped at a prede ned threshold
coverage, considered acceptable. In our case we set the threshold equal to 95%.</p>
        <p>The resulting choice of ontolgies sets is: OBI for Platform, NCIT for Ethnicity,
NCBITaxon for Species, NCIT for Disease, UBERON for Tissue, fEFO,CLg for
CellLine, NCIT for Feature, and OGG for Target. All the above choices meet the
set threshold. Our best choice for the attribute Technique is the set fOBI,EFOg,
but for this attribute we are not able to achieve the coverage threshold, as we
reach a best coverage of 85.7%.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Metadata Enrichment</title>
      <p>After selecting such sets, we proceed with the enrichment of the values contained
in the ontological attributes of the GCM. Section 3.1 presents the relational
schema which supports this phase. Section 3.2 describes the enrichment
process and Section 3.3 shows how the automatic annotation is aided by curators
intervention. Finally, Section 3.4 overviews the expert validation.
3.1</p>
      <sec id="sec-3-1">
        <title>Relational Schema</title>
        <p>
          Figure 1 describes the logic schema of the relational database. The Genomic
Conceptual Model frame contains the tables from the GCM (of which we only
show in detail the ones which have ontological attributes). The Local
Knowledge Base (LKB) frame stores all the information retrieved from OLS services
and relevant to annotate our values. The main tables are: vocabulary (storing
the reference term ids), synonym (containing synonyms of the preferred label in
the vocabulary), reference (identi ers of equivalent terms in alternative
ontologies), ontology (dimension table for used ontologies), and relationship
(representing links between terms in the ontology). The Expert Support frame contains
the tables used to contain information for expert users. Each GCM ontological
attribute X is equipped with a companion-attribute X tid, which references the
ontological term in the vocabulary table (e.g., Platform with value \Illumina
Human Methylation 450" is associated to Platform tid = 10, representing the
vocabulary object OBI 0001870, taken from the Ontology of Biomedical
Investigations [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]). The Vocabulary table is the central entity of the LKB schema. The
tid column is the primary key which is referenced by all other tables in LKB
and from the tables in GCM. Also tables from the LKB and from the Expert
with the ontologies sets indicated in Section 2.4. When a best match score,
calculated as in Eq. 1, is found and is above the threshold 5.0, we select the
corresponding term and proceed with the annotation, otherwise the decision is
delegated to data curators (see Section 3.3).
        </p>
        <p>Once the term has been selected, we populate the tables of the LKB with all
the information derived from OLS regarding the term: description, iri, synonyms,
xrefs, hypernyms and hyponyms (both of is a and part of kinds). The depths
of ancestors and descendants retrieved from the ontology are con gurable by
constant speci cation. The automatic enrichment process currently annotates
about 83% of the total raw values, while the remaining are handled using a
manual curation procedure.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Biologists Support</title>
        <p>We propose two procedures which allow experts curators to support the
annotation algorithm; we assume them to be knowledgeable about biological data
management and to be expert in genomic data curation.</p>
        <p>In the rst procedure, a curator can examine all cases in which the algorithm
is not able to provide a high quality match (i.e., the service provides either partial
matches with low score or no result). The low scores matches are proposed as
suggestions so that the curator may select one of them. In any case, a manual
annotation can always be provided. The procedure can be con gured so that it
also shows the cases with the same score.</p>
        <p>The second procedure is started when a pre-existing annotation is not
adequate (i.e., a tid column has been lled with a wrong vocabulary term). In this
case, the curator can invalidate the annotation and provide an alternative.
3.4</p>
      </sec>
      <sec id="sec-3-3">
        <title>Expert Validation</title>
        <p>We conducted a validation by engaging six experts with good biological
knowledge. For each considered attribute, we presented to them a set of annotations
(i.e., matches between a raw and an ontological term, equipped with its
descriptions) automatically produced by the enrichment procedure. We asked them to
rate the associations according to how accurate they are w.r.t their knowledge.</p>
        <p>The questionnaire contains up to 20 matches for each attribute (or less in
the case of Platform, Species, Technique, and Feature, which contain less found
matches), selected randomly from their value pools, therefore considered
representative of the sets. The test allows ve choices: 1. exact, 2. good, 3.
acceptable, 4. wrong, 5. do not know.</p>
        <p>In Table 2, in the rst row we indicate, for each attribute, the ratio between
the number of automatically annotated values and the number of their total
distinct values. Then, we show in detail the results from the attributes presented
to experts. The averaged results highlight that in 83.06% of cases the experts
marked as exact or good the examined matches, in 8.81% they rated them as
acceptable, and only the remaining 7.05% were marked as wrong. In the 1.08%
of cases the experts declared they were not able to evaluate the match.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Related</title>
    </sec>
    <sec id="sec-5">
      <title>Works</title>
      <p>
        Many works in the literature consider the problem of recognizing ontological
concepts to perform semantic annotation of data. For example: Bodenreider [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
proposes a (dated) survey on the use of ontologies in biomedical data
management and integration; the works [
        <xref ref-type="bibr" rid="ref11 ref18">11, 18</xref>
        ] debate solutions devoted to data
integration; Giles et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] focus on concept extraction from datasets of a speci c
source; the works [
        <xref ref-type="bibr" rid="ref14 ref5">14, 5</xref>
        ] consider the problem of metadata authoring by using
BioPortal ontology-based recommendations, with a focus on metadata manual
creation and preparation. A number of articles have addressed the problem of
choosing ontologies for semantic enrichment. Among these: Wilkinson et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
present the FAIR principles, which de ne a set of characteristics that data
resources and infrastructures should exhibit; [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] identify key search factors for
biomedical ontologies to help biomedical experts in selecting the best-suited
ones in their search cases. In Section 2.1 we presented BioPortal [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], Ontology
Recommender [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], Ontology Lookup Service [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and Zooma, since we believe
UMLS [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], HeTop [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and Annotator [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] were not suited for our purpose.
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>Annotating metadata with terms from ontologies and providing an expansion
to hypernyms and hyponyms allows for easier and semantically exible dataset
search. We provide selection criteria for choosing among search services and
ontologies, and a user-friendly process for assisting biologists in checking that
suggested terms are indeed acceptable. We also provided an internal validation
of annotations produced by our process. As future work, we intend to improve
the matching algorithm by exploiting the ontology structures and information.
We plan to integrate more sources and test our method on a comprehensive
database.</p>
      <p>The implementation of the metadata enrichment system described in this
paper is is available at: https://github.com/DEIB-GECO/Metadata-Enricher.
It is used in the broader context of a genomic repository, developed within the
GeCo Project3, which will be available for use in the near future. Enriched
metadata help users in locating datasets for genomic data extraction and analysis,
either on their original sources or within our repository.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgment</title>
      <p>This research is funded by the ERC Advanced Grant 693174 GeCo (data-driven
Genomic Computing).</p>
      <sec id="sec-7-1">
        <title>3 http://www.bioinformatics.deib.polimi.it/geco/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bandrowski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.:
          <article-title>The ontology for biomedical investigations</article-title>
          .
          <source>PloS one 11(4)</source>
          ,
          <year>e0154556</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bernasconi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.:
          <article-title>Conceptual modeling for genomics: Building an integrated repository of open data</article-title>
          . In: Mayr,
          <string-name>
            <given-names>H.C.</given-names>
            ,
            <surname>Guizzardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            , Ma, H.,
            <surname>Pastor</surname>
          </string-name>
          ,
          <string-name>
            <surname>O</surname>
          </string-name>
          . (eds.) Conceptual Modeling. pp.
          <volume>325</volume>
          {
          <fpage>339</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The uni ed medical language system (UMLS): integrating biomedical terminology</article-title>
          .
          <source>Nucleic acids research 32(suppl 1)</source>
          ,
          <source>D267{D270</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Biomedical ontologies in action: role in knowledge management, data integration and decision support</article-title>
          .
          <source>Yearbook of Medical</source>
          Informatics p.
          <volume>67</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Egyedi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          , et al.:
          <article-title>Embracing semantic technology for better metadata authoring in biomedicine</article-title>
          .
          <source>In: Proceedings of SWAT4LS International Conference</source>
          <year>2017</year>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Consortium</surname>
            <given-names>ENCODE</given-names>
          </string-name>
          :
          <article-title>An integrated encyclopedia of DNA elements in the human genome</article-title>
          .
          <source>Nature</source>
          <volume>489</volume>
          (
          <issue>7414</issue>
          ),
          <volume>57</volume>
          {
          <fpage>74</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.B.</given-names>
          </string-name>
          , et al.:
          <article-title>Ale: automated label extraction from GEO metadata</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>18</volume>
          (
          <issue>14</issue>
          ),
          <volume>509</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Health multi-terminology portal: a semantic added-value for patient safety</article-title>
          .
          <source>Studies in health technology and informatics 166</source>
          ,
          <issue>129</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          , et al.:
          <article-title>The NCI Genomic Data Commons as an engine for precision medicine</article-title>
          .
          <source>Blood</source>
          <volume>130</volume>
          (
          <issue>4</issue>
          ),
          <volume>453</volume>
          {
          <fpage>459</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The open biomedical annotator</article-title>
          .
          <source>In: AMIA Summit on Translational Bioinformatics</source>
          . pp.
          <volume>56</volume>
          {
          <issue>60</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , et al.:
          <article-title>A system for ontology-based annotation of biomedical data</article-title>
          .
          <source>In: International Workshop on Data Integration in The Life Sciences</source>
          . pp.
          <volume>144</volume>
          {
          <fpage>152</fpage>
          . Springer (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Jupp</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>A new Ontology Lookup Service at EMBL-EBI</article-title>
          . In: Malone,
          <string-name>
            <surname>J.</surname>
          </string-name>
          , et al. (eds.)
          <source>Proceedings of SWAT4LS International Conference</source>
          <year>2015</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kundaje</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.:
          <article-title>Integrative analysis of 111 reference human epigenomes</article-title>
          .
          <source>Nature</source>
          <volume>518</volume>
          (
          <issue>7539</issue>
          ),
          <volume>317</volume>
          {
          <fpage>330</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mart</surname>
            nez-Romero,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Fast and accurate metadata authoring using ontologybased recommendations</article-title>
          .
          <source>In: AMIA Annual Symposium Proceedings</source>
          . vol.
          <year>2017</year>
          , p.
          <fpage>1272</fpage>
          . American Medical Informatics Association (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mart</surname>
            nez-Romero,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <source>NCBO Ontology Recommender</source>
          <volume>2</volume>
          .
          <article-title>0: an enhanced approach for biomedical ontology recommendation</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <volume>21</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Needleman</surname>
            ,
            <given-names>S.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wunsch</surname>
          </string-name>
          , C.D.:
          <article-title>A general method applicable to the search for similarities in the amino acid sequence of two proteins</article-title>
          .
          <source>Journal of molecular biology 48(3)</source>
          ,
          <volume>443</volume>
          {
          <fpage>453</fpage>
          (
          <year>1970</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al.:
          <article-title>Where to search top-k biomedical ontologies? Brie ngs in Bioinformatics p</article-title>
          .
          <year>bby015</year>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.H.</given-names>
          </string-name>
          , et al.:
          <article-title>Ontology-driven indexing of public datasets for translational bioinformatics</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>10</volume>
          (
          <issue>2</issue>
          ),
          <source>S1</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Whetzel</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          , et al.:
          <article-title>Bioportal: enhanced functionality via new web services from the national center for biomedical ontology to access and use ontologies in software applications</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>39</volume>
          (
          <issue>suppl 2</issue>
          ),
          <source>W541{W545</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Wilkinson</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.D.</surname>
          </string-name>
          , et al.:
          <article-title>The FAIR guiding principles for scienti c data management and stewardship</article-title>
          .
          <source>Scienti c Data</source>
          <volume>3</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>