<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Annotation process, guidelines and text corpus of small non-coding RNA molecules: the MiNCor for microRNA annotations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jose´ Camilla Sammartino</string-name>
          <email>j.sammartino.88@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Krallinger</string-name>
          <email>mkrallinger@cnio.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alfonso Valencia</string-name>
          <email>avalencia@cnio.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro Nacional de</institution>
          ,
          <addr-line>Investigaciones Oncolo ́gicas., Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centro Nacional de</institution>
          ,
          <addr-line>Investigaciones Oncolo ́gicas., Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Molecular Medicine, and Medical Biotechnology., University of Naples Federico II</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>MicroRNA are small non-coding
molecules that act as post-transcriptional
regulators of gene expression in a wide
spectrum of biological states. Mostly, the
information about microRNA is
embedded in unstructured data (text files) which
needs specific text mining techniques
for its retrieval and analysis. These
are generally based on supervised (or
semi-supervised) learning methods, which
require collections of neatly annotated
and categorised training data. In this study
we propose a comprehensive granular
annotation protocol for the annotation
of non-coding RNA molecules, focusing
primarily on microRNA mentions. This
annotation protocol was used to construct
a manually annotated corpus (MiNCor
Gold) for microRNA mentions as well as
a large semi-automatically generated
microRNA mentions silver standard corpus
(MiNCor Silver) and a large microRNA
name dictionary. Therefore, the efficiency
of these standards was evaluated using a
named entity recognition (NER) system in
comparison with another microRNA
mentions standard freely available online. The
NER system trained with our silver corpus
showed a better performance, with higher
precision (96,67% vs. 94,00%) and recall
(97,57% vs. 95,00%) on their test data and
on our (precision 89,26% vs. 88,97% and
recall 90,03% vs. 86,74%). The corpora
and guidelines are freely downloadable at
http://zope.bioinfo.cnio.es/
mincor/minacor.tar.gz.</p>
      <p>
        MicroRNAs are small non-coding RNA molecules
involved in the post-transcriptional regulation of
gene expression. In the last decade they have
been linked to a wide spectrum of
biological/developmental processes and diseases
including cancer, metabolic disorders or infectious
diseases
        <xref ref-type="bibr" rid="ref24 ref26 ref27 ref29 ref3">(Bayoumi et al., 2016; Smith et al., 2015;
Pogue et al., 2014; Ohtsuka et al., 2015; Pileczki
et al., 2016)</xref>
        . MicroRNAs are post-transcriptional
regulators of gene expression acting on the
messenger RNA target. The maturation of
microRNAs is a double step-and-area process, starting
in the Nucleus of the cell, where is cleaved then
exported in the Cytoplasm where is subjected to
another cleavage which produce a double-strands
microRNA of 22 nucleotides. This dsmicroRNA
is recognised by the RNA-Induced Silencing
Complex (RISC)
        <xref ref-type="bibr" rid="ref30">(Stroynowska-Czerwinska et al.,
2014)</xref>
        . Even though the exact mechanism of
action of RISC is yet fully understood, there are
evidences that RISC is lead to the messenger RNA
(mRNA) target by the microRNA, which has a
homologous sequence to the 3’ - UnTranslated
Region (3’ - UTR) of the target. The binding to this
region allows the regulation process, that can
happen before or during the translation in protein of
the messenger, which means that is possible to
have or don?t have protein products
        <xref ref-type="bibr" rid="ref21">(Morozova
et al., 2012)</xref>
        . Temporal and spatial expression
of these molecules is important as much as their
expression levels, a modification in one of these
can lead to a dysregulation of the biological
processes in which they are involved, with effects that
can expand to entire biological pathways.
Different studies show the importance of a correct
microRNA post-transcriptional regulation to prevent
the development of pathological states and
development defects
        <xref ref-type="bibr" rid="ref27 ref30 ref5">(Bhaskaran and Mohan, 2014)</xref>
        , but
their importance is also enlightened by their
possible application as fast, specific and non-invasive
biomarkers in a large spectrum of harmful states
        <xref ref-type="bibr" rid="ref16 ref28 ref4">(Rubio et al., 2016; Benz et al., 2016; Larrea et al.,
2016)</xref>
        . Furthermore, these molecules can be used
as target in pharmacological therapies and
clinical application (for the diagnosis and the
followup)
        <xref ref-type="bibr" rid="ref18 ref19 ref7">(Lin et al., 2014; Du et al., 2014; Mao et
al., 2013)</xref>
        . This promoted the publication of an
increasing number of publications especially
devoted to the study of microRNA biology as well
as predictive bioinformatics analysis
methodologies tailored to the characterisation of miRNA
expression and target prediction.
      </p>
      <p>
        Biomedical Natural Language Processing
(BioNLP) techniques and text mining strategies
can be applied for the retrieval, filtering and
analysis of knowledge from unstructured data such
the scientific literature. One of the main hurdles
for the implementation of text mining building
block components is the construction of manually
annotated text-bound corpora, as they require
usually a considerable human workload together
with annotators with deep domain knowledge and
basic linguistic expertise. The development of
corpora is a time-consuming, tedious and very so
much needed process for Text Mining and BLP
methods
        <xref ref-type="bibr" rid="ref23">(Neves, 2014)</xref>
        . To promote advances
in BLP, different competitive evaluations have
been held
        <xref ref-type="bibr" rid="ref10 ref11 ref14">(Hersh et al., 2004; Kim et al., 2004;
Hirschman et al., 2005)</xref>
        , in which distinct groups
participated in different tasks, ranking from
document retrieval, NER to complex relation/event
extraction tasks
        <xref ref-type="bibr" rid="ref12">(Hunter and Cohen, 2006)</xref>
        . Those
challenges resulted in valuable text corpora that
have been re-used by the biomedical text mining
community.
      </p>
      <p>Despite the release of several manually
annotated text corpora devoted to biological entities,
there isn’t one and only manual, work or
reference that can be considered as a general guide to
build specific guidelines which are usually
written based on the background knowledge of the
authors or don’t include all the possible
characteristics. Furthermore, the annotation process can be
very variable and complex, due to the
interconnection of different disciplines (medical/biological
and linguistic) and the different aims of the
annotation (chemical compounds, disease,
connection between mutated proteins and disease, case
reports).</p>
      <p>The assembling of a corpus requires specific
documents that describe the annotation process
and define its guidelines.</p>
      <p>
        As for microRNAs, several attempts have been
made to facilitate the extraction of information
directly from the literature
        <xref ref-type="bibr" rid="ref17 ref2 ref22 ref8">(Bagewadi et al., 2014;
Griffiths-Jones et al., 2006; Li et al., 2015; Naeem
et al., 2010; Xie et al., 2013)</xref>
        . To our
knowledge there are three freely-available corpora for
microRNA (miRNA) mentions, two of them,
MirBase and MirTex
        <xref ref-type="bibr" rid="ref17 ref8">(Griffiths-Jones et al., 2006; Li et
al., 2015)</xref>
        , do provide very short annotation
guidelines
        <xref ref-type="bibr" rid="ref1 ref20 ref9">(Ambros et al., 2003; Meyers et al., 2008;
Griffiths-Jones, 2004)</xref>
        for the annotation of
microRNA mentions which mostly focus on the
identification of single mentions, without considering
more granular annotation types. The third
corpus (SCAI corpus)
        <xref ref-type="bibr" rid="ref2">(Bagewadi et al., 2014)</xref>
        , does
provide additional details and a set of annotation
rules, but we believe that it underspecified some
of the relevant annotation criteria and it primarily
focuses only on human microRNA Mentions. For
instance it covers the annotation of species
specific prefixes, e.g. hsa for human miRNAs, but
does not annotated terms such as human
preceding miRNA mentions. Moreover, general prefixes
(anti-, onco-, pre-, pri-), specific miRNA class
names (angiomir, antagomir, isomir), as well as
non-coding RNA names are not included in the
annotation process.
      </p>
      <p>Here we propose a comprehensive annotation
protocol for labelling microRNA mentions in
biomedical literature. It encompasses all
microRNA mentions regardless of the species or
origin, the maturing step or the classification and
includes also a class of non-coding RNA names
and miRNA clusters. This annotation protocol has
been iteratively refined and was then used for the
annotation of the MiNCor corpus, which as used
for the evaluation of several microRNA mentions
recognition approaches. We believe that the
release of this MiNCor corpus guidelines might be
useful as an annotation template for the corpus
construction of other biomedical entities.</p>
      <p>We tested our corpus in comparison with the
SCAI corpus, which is to our knowledge, the one
whose guidelines are the most comprehensive so
far. Therefore, to test the efficiency of our corpus
we trained and tested a named entity recognition
(NER) system with it and evaluated the results in
comparison with SCAI, whose trainer and tester
for the NER system were the only ones with
characteristics that could be compared to ours.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Annotation protocol and guidelines</title>
      <p>
        The guidelines for the MiNCor annotation
protocol is composed of a 14 pages written manual
defined by a biotechnologist with extensive
biological knowledge, integrating information from
previous miRNA corpora, revision of multiple
different resources (NCBI, MeSH terms, miRNA review
articles) and the model of the Manual for
annotation of chemical entities of the CHEMDNER
corpus
        <xref ref-type="bibr" rid="ref15">(Krallinger et al., 2015)</xref>
        . The annotation
protocol is structured into rule types together with
example cases, which we call the GPNCE annotation
system, standing for general rules, positive rules,
negative rules, class rules and examples. We
believe that structuring the annotation protocol into
such rules, makes it easier to follow the
annotation criteria by the human annotators during the
labelling of the mentions.
2.1
      </p>
      <sec id="sec-2-1">
        <title>The GPCNE annotation protocol</title>
        <p>We based our guidelines on a three phase
annotation protocol that we called GPNCE (General,
Positive, Negative, Class and Examples).
We firstly describe the different classes that can
be identified in literature. We cover in detail
six different classes of microRNA mentions: (1)
general microRNA names, (2) specific microRNA
names, (3) multiple microRNA mentions, (4)
nested microRNA mentions, (5) microRNA
cluster mentions and (6) other/non-coding RNA
mentions. Figure 1 provides different examples for the
classes.</p>
        <p>
          In the second phase we propose three types of
rules for the annotation: General, Positive and
Negative. The General Rules describe the
decisions the annotator should take into account
during the annotation process (what constitutes at a
general level a miRNA mention and how to deal
with cases of uncertainty). The Positive Rules
describe how to annotate correct miRNA
mentions, what to include in the mentions (positive
word-boundaries, prefixes, suffixes, symbols) and
illustrate the criteria through different examples
of correctly annotated mentions. The Negative
Rules, cover what needs to be excluded from the
mentions during the annotation
(negative-wordboundaries, prefixes, part-of-speech entities, other
entities, wrong mentions). All the rules described
included examples with positive cases (specified
by the check mark symbol) and negative ones
(specified by the cross mark symbol). The
definition of these rules was based on the Manual for
annotation of chemical entities of the CHEMDNER
corpus
          <xref ref-type="bibr" rid="ref15">(Krallinger et al., 2015)</xref>
          .
        </p>
        <p>The last phase, Examples, consist in two appendix
at the end of the manual in which are represented
different examples of annotated mentions in
sentences extracted from abstracts and possible errors
to avoid. To help visualising the correct labels,
those were highlighted with a specific colour
coding system in regard of the class to which them
belong.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>MiNCor Corpora</title>
      <p>
        We decided to build two different corpus for
microRNA mention following the GPCNE
annotation protocol. The first one is a manually
labeled corpus of 102 abstracts called MiNCor
Gold, the second is a semi-automatically
generated corpus of 302K sentences called
MiNCor Silver. The corpora and the guidelines can
be downloaded at http://zope.bioinfo.
cnio.es/mincor/minacor.tar.gz. The
directory contains six files with a ’README.txt’
file which contains all the descriptions of the other
files.
3.1
Using the MiNCor annotation guidelines a domain
expert annotator performed a manual labelling
of 102 abstracts with at least one microRNA
mention. The abstracts were randomly selected
from 3869 abstracts retrieved using the MeSH
query ”mirna” on Pubmed but restricting the
search only to papers published in 2016. The
labelling was performed manually using the
customised AnnotateIt web-interface http:
//ubio.bioinfo.cnio.es/people/
fleitner/mirnaner_test_250.html,
similar to the one used for the annotation of
the CHEMDNER-Patents Corpus
        <xref ref-type="bibr" rid="ref15">(Krallinger
et al., 2015)</xref>
        , but adjusting the different classes
and labels to the microRNA Mention Classes, a
schematic overviw of the annotation protocol is
summarised in figure 2. Out of the 102 abstracts
we extracted a total of 1154 mentions. Table 1
provides an overview of the distribution of the
microRNA class types.
      </p>
      <sec id="sec-3-1">
        <title>Type of Mention</title>
        <p>General microRNA
Specific microRNA
Multiple microRNA
Nested microRNA
Cluster microRNA
ncRNA microRNA</p>
        <p>To validate the annotation process 20 of these
abstracts were randomly selected and de novo
annotated using the same annotation guidelines by
a second annotator. The results obtained were
then compared with the first annotation
considering only perfect mention matches. The
annotator agreement scores resulted in: Precision:
99,00, Recall: 99,45 and F1: 99,22. All the
nonoverlapping labeled entities were analysed and
modified only when the annotators could agree, if
there still was uncertainty the mentions were left
unlabelled. The errors in the annotation mostly
concerned the non-coring RNA class mentions :
- [...the long intergenic non-coding RNAs
(lincRNAs) expressed in...]
- Should we consider ’non-coding RNAs
(lincRNAs)’ as a unique mention or as a two separate
mention ’non-coding RNAs’, ’lincRNAs’?
- Should ’long intergenic’ be included in the
mention?
In this example, if we follow the guidelines, the
correct labelling should be ’long intergenic
noncoding RNAs’ and ’lincRNAs’, but in this type of
cases the annotators labelled differently (one did
the label as a single mention and the other did
as two separate mentions) showing uncertainty.
Therefore, we suggest to not label the mention
when there are doubts.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>MiNCor Silver</title>
        <p>
          The MiNCor Silver was obtained from sentences
that were derived from PubMed results
containing the MeSH term microRNA as well as
additional manually defined microRNA search query
terms (”mirna”; ”microrna”; ”non-coding RNA”;
”lin-4”; ”let-7”, ”antagomir”; ”oncomir”). The
research was limited to all the abstracts ad full
text papers published starting from 2016. All
the resulting files were segmented into sentences
and then tagged using a large dictionary of
microRNA names (MiNCor lexicon). This
dictionary contained names derived from multiple
microRNA databases as well as microRNA
mentions detected by GNormplus. We carried out a
dictionary expansion step taking into account the
nomenclature guidelines of microRNAs by
considering core terms (e.g. miRNA, microRNA),
prefixes (e.g. hsa, mmu) and suffixes (e.g. -101;
-23a/b). A dictionary pruning step was carried out
to remove highly ambiguous mentions (e.g. ’MIR’
was found to be referred to other entities as for
example the ’Space Station Mir’
          <xref ref-type="bibr" rid="ref13">(Johannes et al.,
2016)</xref>
          ). After applying the dictionary look-up we
additionally used a cascade of rules to adjust the
mention boundaries, to cover for instance
mentions of co-ordinated microRNAs or lists of
microRNAs. In the end our training silver corpus had
a total of 302’560 sentences, over 3’000’000 of
tokens and over 175K labeled microRNA mentions,
the sum of the entities used in the different
dictionaries for the post-processing amount to 788’784
different terms in total.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Comparison with other corpus</title>
      <p>
        We decided to use both the MiNCor corpus and
the SCAI miRNA corpus
        <xref ref-type="bibr" rid="ref2">(Bagewadi et al., 2014)</xref>
        as a test set to evaluate the performance of a
CRF-based miRNA entity tagger. We chose
the SCAI corpus, because it did provide
annotation criteria, and thus allowed some
interpretation of differences in the annotation process.
We constructed a miRNA entity tagger using
the NERsuite toolkit
        <xref ref-type="bibr" rid="ref6">(Cho et al., 2010)</xref>
        . The
NERsuite toolkit is freely available at http:
//nersuite.nlplab.org. It is based on
the CRFsuite (http://www.chokkan.org/
software/crfsuite/ an implementation of
the Conditional Random Field)
        <xref ref-type="bibr" rid="ref25">(Okazaki, 2007)</xref>
        and includes different feature types commonly
used for biomedical NER tasks, including aspects
covering lemmatization, Part-Of-Speech and word
morphology.
      </p>
      <p>We trained two different NER models, one using
the miRNA SCAI training set (201 manually
labeled abstracts) and another based on a large
silver standard MiNCor training dataset comprising
302K sentences. We generated a CRF model both
using the SCAI training set and the MiNCor silver
standard training set.</p>
      <p>Both the two corpora used for the train of the
models were segmented in sentences and tokenised.
At token level the two corpora were lemmatised,
labelled with Part-Of-Speech and chunking tags,
and labelled following the I.O.B. format. The
results were defined in terms of Precision, Recall
and F1. The obtained results with the two
models are shown in table 2, for the SCAI test set, and
in table 3, using the MiNCor gold standard test set.
Table 4 shows the overall statistics of the two test
corpora (MiNCor Gold and SCAI).</p>
      <sec id="sec-4-1">
        <title>Score</title>
        <p>Precision
Recall
F-score
Using our guidelines we manually annotated 102
abstracts retrieved form Pubmed using the MeSH
query ”mirna” and filtering the results
including only the recent ones (2016). At the same
time, with a more refined search on Pubmed
(including different MeSH queries) we extracted
302K sentences that were semi-automatically
labeled following our guidelines and subsequently
pruned with dictionary look-up and a cascade
of rules to adjust the mention boundaries. We
then tested our corpora in comparison with the
SCAI manually labelled corpus using the
NERsuite toolkits to perform the named entity
recognition task for microRNA mention in
literature. We used the MiNCor Gold as our test
and the MiNCor Silver as the trainer to build
our model. To obtain the SCAI model we
trained the NERsuite with their trainer
downloadable at http://www.scai.fraunhofer.
de/mirna-corpora.html. As shown in
Table 2 and Table 3, the microRNA tagger models,
trained using the SCAI training set (Table 2) and
our dictionary/rule-based Silver Standard training
set (Table 3), report lower scores when using our
corpus as gold standard (second column of the two
tables). This is due to the more granular
definition of the microRNA mentions and by including
for instance also other ncRNA types that were not
labelled in the used training collections. On the
other end, our model had a better performance in
comparison with SCAI on both test sets, this is
due to our model, even though not being manually
curated, covers more possible mentions, including
microRNA mentions for all different species and
biosynthesis steps, furthermore, it includes more
classes of mentions, leading to a more
comprehensive identification.</p>
        <p>Even if our model had a better performance, the
resulting score wasn’t perfect. Some of the main
sources of errors related to the microRNA mention
recognition was due to mention of lists of
microRNAs, where microRNA mentions are expressed as
multiple overlapping entity mentions (mir-1, -23,
-33 and -101). Other errors occurred in the
labelling of non-coding RNAs.</p>
        <p>Non-coding RNA mentions are hard to define
because there isn’t a specific nomenclature to which
the researcher can refer. Nevertheless , there are
resources online (NCBI, MeSH terms, miRNA
review articles, books) that can help in the definition
of this class. What we tried to do was to give rules
for the identification of non-coding RNA
mentions, where the most important was that in case
of uncertainty the mention shouldn’t be labelled,
which results in a lower accuracy for the model.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and future works</title>
      <p>Here we have presented the MiNCor corpora and
the Guidelines for the Annotation for microRNA
and non-coding RNA mentions in scientific
literature. The aim of this work was to provide
annotation guidelines that are comprehensive and
explicative, using different examples for the
annotation and rules to help the annotator during
the process. The availability of exhaustive
guidelines for the annotation of biomedical entities is a
very important contribution for Biomedical
Natural Language Processing tasks, because gives the
researcher the possibility to have a standardised
tool that can help in the definition of a line of
research even without extensive knowledge of the
field. Furthermore, the possibility to use
predefined guidelines for the construction of corpora
can reduce the time needed for the process.
We also constructed two corpora (gold and
silver) using our guidelines and tested them with
a named entity recognition task using the
NERsuite toolkit and comparing the results with
another microRNA tagger already available.
Manually curated corpora are considered a gold standard
in Natural Language Processing because they can
generally reach higher level of accuracy. In our
case that is not true, which provide an example
of a good surrogate for manually annotated gold
standard corpora. At the moment there aren’t very
large gold standard for microRNA mention that
encompass all the possible characteristic and types
of mention, which is why our MiNCor Silver can
be considered a better option, even though not
being manually curated, as shown by the results we
obtained.</p>
      <p>In the future, our intent is to enlarge our guidelines
with other types of non-coding RNAs (e.g.
ribosomial RNAs, transfer RNAs) that are not included
at the moment, provide a larger corpus of
microRNAs derived from full text and patent abstract
sentences and describe additional rules to help
defying the relations of these molecules with other
biological entities (e.g. chemical compounds, genes,
proteins).</p>
      <p>Boya Xie, Qin Ding, Hongjin Han, and Di Wu. 2013.
mircancer: a microrna–cancer association database
constructed by text mining on literature.
Bioinformatics, page btt014.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Victor</given-names>
            <surname>Ambros</surname>
          </string-name>
          , Bonnie Bartel, David P Bartel, Christopher B Burge, James C Carrington, Xuemei Chen, Gideon Dreyfuss, Sean R Eddy, SAM GriffithsJones,
          <string-name>
            <surname>Mhairi Marshall</surname>
          </string-name>
          , et al.
          <year>2003</year>
          .
          <article-title>A uniform system for microrna annotation</article-title>
          .
          <source>Rna</source>
          ,
          <volume>9</volume>
          (
          <issue>3</issue>
          ):
          <fpage>277</fpage>
          -
          <lpage>279</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Shweta</given-names>
            <surname>Bagewadi</surname>
          </string-name>
          , Tamara Bobic´,
          <string-name>
            <surname>Martin</surname>
            <given-names>HofmannApitius</given-names>
          </string-name>
          , Juliane Fluck, and
          <string-name>
            <given-names>Roman</given-names>
            <surname>Klinger</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Detecting mirna mentions and relations in biomedical literature</article-title>
          .
          <source>F1000Research</source>
          , 3.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Ahmed S Bayoumi</surname>
          </string-name>
          , Amer Sayed, Zuzana Broskova,
          <string-name>
            <surname>Jian-Peng</surname>
            <given-names>Teoh</given-names>
          </string-name>
          , James Wilson, Huabo Su, YaoLiang Tang, and Il-man
          <string-name>
            <surname>Kim</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Crosstalk between long noncoding rnas and micrornas in health and disease</article-title>
          .
          <source>International journal of molecular sciences</source>
          ,
          <volume>17</volume>
          (
          <issue>3</issue>
          ):
          <fpage>356</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Fabian</given-names>
            <surname>Benz</surname>
          </string-name>
          , Sanchari Roy, Christian Trautwein, Christoph Roderburg, and Tom Luedde.
          <year>2016</year>
          .
          <article-title>Circulating micrornas as biomarkers for sepsis</article-title>
          .
          <source>International journal of molecular sciences</source>
          ,
          <volume>17</volume>
          (
          <issue>1</issue>
          ):
          <fpage>78</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>M</given-names>
            <surname>Bhaskaran</surname>
          </string-name>
          and
          <string-name>
            <given-names>M</given-names>
            <surname>Mohan</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Micrornas history, biogenesis, and their evolving role in animal development and disease</article-title>
          .
          <source>Veterinary Pathology Online</source>
          ,
          <volume>51</volume>
          (
          <issue>4</issue>
          ):
          <fpage>759</fpage>
          -
          <lpage>774</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>HC</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N</given-names>
            <surname>Okazaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Miwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J</given-names>
            <surname>Tsujii</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Nersuite: a named entity recognition toolkit</article-title>
          .
          <source>Tsujii Laboratory</source>
          , Department of Information Science, University of Tokyo, Tokyo, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Bowen</given-names>
            <surname>Du</surname>
          </string-name>
          , Zhe Wang, Xin Zhang, Shipeng Feng, Guoxin Wang,
          <string-name>
            <surname>Jianxing He</surname>
            ,
            <given-names>and Biliang</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Microrna-545 suppresses cell proliferation by targeting cyclin d1 and cdk4 in lung cancer cells</article-title>
          .
          <source>PloS one</source>
          ,
          <volume>9</volume>
          (
          <issue>2</issue>
          ):
          <fpage>e88022</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Sam</given-names>
            <surname>Griffiths-Jones</surname>
          </string-name>
          , Russell J Grocock, Stijn Van Dongen,
          <string-name>
            <surname>Alex Bateman</surname>
          </string-name>
          , and
          <string-name>
            <surname>Anton</surname>
          </string-name>
          J Enright.
          <year>2006</year>
          .
          <article-title>mirbase: microrna sequences, targets and gene nomenclature</article-title>
          .
          <source>Nucleic acids research</source>
          ,
          <volume>34</volume>
          (
          <issue>suppl 1</issue>
          ):
          <fpage>D140</fpage>
          -
          <lpage>D144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Sam</given-names>
            <surname>Griffiths-Jones</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>The microrna registry</article-title>
          .
          <source>Nucleic acids research</source>
          ,
          <volume>32</volume>
          (
          <issue>suppl 1</issue>
          ):
          <fpage>D109</fpage>
          -
          <lpage>D111</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>William</given-names>
            <surname>Hersh</surname>
          </string-name>
          , Ravi Teja Bhupatiraju, and
          <string-name>
            <given-names>Sarah</given-names>
            <surname>Corley</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Enhancing access to the bibliome: the trec genomics track</article-title>
          .
          <source>Medinfo</source>
          ,
          <volume>11</volume>
          (Pt 2):
          <fpage>773</fpage>
          -
          <lpage>777</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Lynette</given-names>
            <surname>Hirschman</surname>
          </string-name>
          , Alexander Yeh,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Blaschke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Alfonso</given-names>
            <surname>Valencia</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Overview of biocreative: critical assessment of information extraction for biology</article-title>
          .
          <source>BMC bioinformatics</source>
          ,
          <volume>6</volume>
          (
          <issue>Suppl 1</issue>
          ):
          <fpage>S1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Lawrence</given-names>
            <surname>Hunter</surname>
          </string-name>
          and
          <string-name>
            <given-names>K Bretonnel</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Biomedical language processing: what's beyond pubmed? Molecular cell</article-title>
          ,
          <volume>21</volume>
          (
          <issue>5</issue>
          ):
          <fpage>589</fpage>
          -
          <lpage>594</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Bernd</given-names>
            <surname>Johannes</surname>
          </string-name>
          , Vyacheslav Salnitski, Alexander Dudukin, Lev Shevchenko, and
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Bronnikov</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Performance assessment in the pilot experiment on board space stations mir and iss</article-title>
          .
          <source>Aerospace medicine and human performance</source>
          ,
          <volume>87</volume>
          (
          <issue>6</issue>
          ):
          <fpage>534</fpage>
          -
          <lpage>544</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Jin-Dong</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Tomoko Ohta, Yoshimasa Tsuruoka, Yuka Tateisi, and
          <string-name>
            <given-names>Nigel</given-names>
            <surname>Collier</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Introduction to the bio-entity recognition task at jnlpba</article-title>
          .
          <source>In Proceedings of the international joint workshop on natural language processing in biomedicine and its applications</source>
          , pages
          <fpage>70</fpage>
          -
          <lpage>75</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Krallinger</surname>
          </string-name>
          , Obdulia Rabal, Florian Leitner, Miguel Vazquez, David Salgado, Zhiyong Lu, Robert Leaman, Yanan Lu, Donghong Ji, Daniel M Lowe, et al.
          <year>2015</year>
          .
          <article-title>The chemdner corpus of chemicals and drugs and its annotation principles</article-title>
          .
          <source>Journal of cheminformatics</source>
          ,
          <volume>7</volume>
          (
          <issue>S1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Erika</given-names>
            <surname>Larrea</surname>
          </string-name>
          , Carla Sole, Lorea Manterola, Ibai Goicoechea, Mar´ıa Armesto, Mar´ıa Arestin, Mar´ıa M Caffarel,
          <string-name>
            <given-names>Angela M Araujo</given-names>
            , Mar´ıa Araiz, Marta
            <surname>Fernandez-Mercado</surname>
          </string-name>
          , et al.
          <year>2016</year>
          .
          <article-title>New concepts in cancer biomarkers: Circulating mirnas in liquid biopsies</article-title>
          .
          <source>International journal of molecular sciences</source>
          ,
          <volume>17</volume>
          (
          <issue>5</issue>
          ):
          <fpage>627</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Gang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Karen E Ross</given-names>
            ,
            <surname>Cecilia N Arighi</surname>
          </string-name>
          , Yifan Peng,
          <string-name>
            <surname>Cathy H Wu</surname>
            , and
            <given-names>K</given-names>
          </string-name>
          <string-name>
            <surname>Vijay-Shanker</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>mirtex: A text mining system for mirna-gene relation extraction</article-title>
          .
          <source>PLoS Comput Biol</source>
          ,
          <volume>11</volume>
          (
          <issue>9</issue>
          ):
          <fpage>e1004391</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Regina</given-names>
            <surname>Lin</surname>
          </string-name>
          , Ling Chen, Gang Chen, Chunyan Hu, Shan Jiang, Jose Sevilla, Ying Wan, John H Sampson,
          <article-title>Bo Zhu, and</article-title>
          <string-name>
            <given-names>Qi-Jing</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Targeting mir23a in cd8+ cytotoxic t lymphocytes prevents tumordependent immunosuppression</article-title>
          .
          <source>The Journal of clinical investigation</source>
          ,
          <volume>124</volume>
          (
          <issue>12</issue>
          ):
          <fpage>5352</fpage>
          -
          <lpage>5367</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Yiping</given-names>
            <surname>Mao</surname>
          </string-name>
          , Ramkumar Mohan, Shungang Zhang, and
          <string-name>
            <given-names>Xiaoqing</given-names>
            <surname>Tang</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Micrornas as pharmacological targets in diabetes</article-title>
          .
          <source>Pharmacological research</source>
          ,
          <volume>75</volume>
          :
          <fpage>37</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Blake C Meyers</surname>
          </string-name>
          ,
          <string-name>
            <surname>Michael J Axtell</surname>
            , Bonnie Bartel,
            <given-names>David P Bartel</given-names>
          </string-name>
          , David Baulcombe, John L Bowman, Xiaofeng Cao, James C Carrington, Xuemei Chen,
          <string-name>
            <surname>Pamela J Green</surname>
          </string-name>
          , et al.
          <year>2008</year>
          .
          <article-title>Criteria for annotation of plant micrornas</article-title>
          .
          <source>The Plant Cell</source>
          ,
          <volume>20</volume>
          (
          <issue>12</issue>
          ):
          <fpage>3186</fpage>
          -
          <lpage>3190</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Nadya</given-names>
            <surname>Morozova</surname>
          </string-name>
          , Andrei Zinovyev, Nora Nonne,
          <string-name>
            <surname>Linda-Louise</surname>
            <given-names>Pritchard</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexander N Gorban</surname>
          </string-name>
          , and
          <string-name>
            <surname>Annick</surname>
          </string-name>
          Harel-Bellan.
          <year>2012</year>
          .
          <article-title>Kinetic signatures of microrna modes of action</article-title>
          .
          <source>Rna</source>
          ,
          <volume>18</volume>
          (
          <issue>9</issue>
          ):
          <fpage>1635</fpage>
          -
          <lpage>1655</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Haroon</given-names>
            <surname>Naeem</surname>
          </string-name>
          , Robert Ku¨ffner, Gergely Csaba, and
          <string-name>
            <given-names>Ralf</given-names>
            <surname>Zimmer</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>mirsel: automated extraction of associations between micrornas and genes from the biomedical literature</article-title>
          .
          <source>BMC bioinformatics</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <fpage>135</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Mariana</given-names>
            <surname>Neves</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>An analysis on the entity annotations in biological corpora</article-title>
          .
          <source>F1000Research</source>
          , 3.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Masahisa</given-names>
            <surname>Ohtsuka</surname>
          </string-name>
          , Hui Ling, Yuichiro Doki, Masaki Mori, and George Adrian Calin.
          <year>2015</year>
          .
          <article-title>Microrna processing and human cancer</article-title>
          .
          <source>Journal of clinical medicine</source>
          ,
          <volume>4</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1651</fpage>
          -
          <lpage>1667</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Naoaki</given-names>
            <surname>Okazaki</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Crfsuite: a fast implementation of conditional random fields (crfs).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Valentina</given-names>
            <surname>Pileczki</surname>
          </string-name>
          , Roxana Cojocneanu-Petric, Mahafarin Maralani, Ioana Berindan Neagoe, and
          <string-name>
            <given-names>Robert</given-names>
            <surname>Sandulescu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Micrornas as regulators of apoptosis mechanisms in cancer</article-title>
          .
          <source>Clujul Medical</source>
          ,
          <volume>89</volume>
          (
          <issue>1</issue>
          ):
          <fpage>50</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Aileen</surname>
            <given-names>I Pogue</given-names>
          </string-name>
          ,
          <article-title>James M Hill,</article-title>
          and
          <string-name>
            <surname>Walter</surname>
          </string-name>
          J Lukiw.
          <year>2014</year>
          .
          <article-title>Microrna (mirna): sequence and stability, viroid-like properties, and disease association in the cns</article-title>
          .
          <source>Brain research</source>
          ,
          <volume>1584</volume>
          :
          <fpage>73</fpage>
          -
          <lpage>79</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Mercedes</given-names>
            <surname>Rubio</surname>
          </string-name>
          , Quique Bassat, Xavier Estivill, and
          <string-name>
            <given-names>Alfredo</given-names>
            <surname>Mayor</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Tying malaria and micrornas: from the biology to future diagnostic perspectives</article-title>
          .
          <source>Malaria journal</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Tanya</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Cha</given-names>
            <surname>Rajakaruna</surname>
          </string-name>
          , Massimo Caputo, and
          <string-name>
            <given-names>Costanza</given-names>
            <surname>Emanueli</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Micrornas in congenital heart disease</article-title>
          .
          <source>Annals of translational medicine</source>
          ,
          <volume>3</volume>
          (
          <issue>21</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Anna</given-names>
            <surname>Stroynowska-Czerwinska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Agnieszka</given-names>
            <surname>Fiszer</surname>
          </string-name>
          , and
          <string-name>
            <surname>Wlodzimierz</surname>
          </string-name>
          J Krzyzosiak.
          <year>2014</year>
          .
          <article-title>The panorama of mirna-mediated mechanisms in mammalian cells</article-title>
          .
          <source>Cellular and Molecular Life Sciences</source>
          ,
          <volume>71</volume>
          (
          <issue>12</issue>
          ):
          <fpage>2253</fpage>
          -
          <lpage>2270</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>