<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Correcting Crowdsourced annotations to Improve Detection of Outcome types in Evidence Based Medicine</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>University of Liverpool</string-name>
          <email>shindsg@liverpool.ac.uk</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MRC Hub for Trials Methodology Research Network</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>The validity and authenticity of annotations in datasets massively influences the performance of Natural Language Processing (NLP) systems. In other words, poorly annotated datasets are likely to produce fatal results in at-least most NLP problems hence misinforming consumers of these models, systems or applications. This is a bottleneck in most domains, especially in healthcare where crowdsourcing is a popular strategy in obtaining annotations. In this paper, we present a framework that automatically corrects incorrectly captured annotations of outcomes, thereby improving the quality of the crowdsourced annotations. We investigate a publicly available dataset called EBM-NLP, built to power NLP tasks in support of Evidence based Medicine (EBM) primarily focusing on health outcomes. Contact Author 1http://biocreative.sourceforge.net/bionlp tools links.html 2https://metamap.nlm.nih.gov/</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Evidence Based Medicine (EBM) is a popular health research
paradigm that enforces healthcare decision making through
the explicit and judicious use of current best evidence
[Sackett et al., 1996]. In practice, researchers in EBM widely use
a framework entitled PICO, representing a collection of
elements that form the basis of clinical questions, i.e. Patients,
Interventions, Comparators and Outcomes [Huang et al.,
2006]. This framework has significantly contributed towards
various key health-care delivery indicators such as
identification of evidence of the effectiveness of a certain treatment or
diagnosis, strategies to evaluate quality of studies and
mechanisms implemented in healthcare [Santos et al., 2007].</p>
      <p>EBM supported by NLP involves extraction of evidence
from biomedical literature powered by several opensource
tools such as BioNLP1MetaMap tools2. This extraction
largely entails extraction of PICO framework elements. This
paper focuses on the extraction of outcomes, emphasizing the
flaws/faults discovered in crowdsourced annotations of health
outcomes within medical research abstracts.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Outcome detection in EBM</title>
      <p>An outcome is a measurement or an observation used to
capture and assess the effect of treatment such as assessment of
side effects (risk) or effectiveness (benefits) [Williamson et
al., 2017]. Examples of outcomes are blood pressure,
anxiety, stress, fatigue and quality of life. In this paper, we
attempt to review and rectify flaws in outcome annotations
utilizing NLP methods. Flaws examined here are ideally
perceived as errors made while manually annotating outcomes
within medical research abstracts. These may vary from,
capturing non-outcomes as outcomes to capturing unnecessary
text such as context as part of an outcome to compressing
multiple outcomes into a single outcome and many others,
as Section 3 discusses. Flaws such as these constrain the
ability to build systems resilient enough to detect outcomes.
From an NLP perspective, the more fragile the quality of
annotations is, the less accurate the prediction models would
be. Ultimately this hampers the overall objective of building
systems that enhance the effective search for evidence within
published literature hence impeding the aims of EBM [Nye et
al., 2018].
1.2</p>
    </sec>
    <sec id="sec-3">
      <title>How and what did we do to achieve the goals in the study?</title>
      <p>We investigate a recently published corpus EBM-NLP [Nye
et al., 2018], comprising abstracts annotated with outcome
types using a highly scrutinised crowdsource labelling
strategy. An outcome type is a classification or category that
collectively embodies a group of outcomes measured during
Randomised Control Trials (RCTs). This investigation begins
with an assessment of whether the annotations retain the true
identity of an outcome as defined in previous paragraph, and
if not, what flaws recur across these annotations. Flaws are
carefully identified with the supervision of domain experts
inorder to eliminate any non-medical judgment or analysis that
would bias our approach. This is followed by formulating
constraints to examine the syntactic and semantic structure of
the annotations to correct the identified flaws.</p>
      <p>In summary, this paper reveals the noise (flaws) discovered
in crowdsourced outcome annotations, It proposes an
approach to mitigate the downsides of noisy data. It concludes
with NLP tasks performed using SOTA approaches (biLSTM,
CNN and SVM) [Young et al., 2018] to evaluate the impact
made by the adopted corrective techniques.
With the increasing trend in applying NLP and machine
learning to healthcare, a tremendous amount of effort from
researchers in this space has been directed towards building
tools to create high-quality datasets. This has birthed a host
of web-based text annotation tools such as APLenty [Nghiem
and Ananiadou, 2018], BRAT3 and Prodigy4. Whilst such
tools enhance text annotation through their rich features, they
still require people to manually annotate data, which often
is a costly and tedious task. Yang et al. (2019)
empirically prove that annotation tasks can be difficult, however
they discover, that despite the noise in crowdsourced
annotations, reasonable models can possibly achieve similar
performance on these annotations as they would when using less
expert data. Vittayakorn and Hays (2011) embarked on work
closely related ours, however focusing on a computer vision
dataset. They assessed the quality of user-annotations by
defining annotation quality functions that calculated scores
representative of the ground-truth of an annotation. Our
proposed approach explores the syntactic and semantic structure
of annotation spans to automatically filter out errors in a
preannotated medical dataset, hence improve the quality of the
annotations.
3</p>
      <sec id="sec-3-1">
        <title>Flaws discovered in crowdsourced annotations of outcomes in healthcare</title>
        <p>Faulting during manual annotation is almost inevitable,
because of the complexity, ambiguity and variation in how
heathcare terms are described across different studies [Xu
et al., 2016] and [Dodd et al., 2018]. The huge wage
expectations from domain expert annotators makes it even
worse [Bruno, 2018]. This section breaks down the
different flaws observed in outcome annotations,
Flaw 1: Inclusion of unnecessary text that is either
supportive of the actual outcome or an elaborated context of an
outcome. Two kinds of unnecessary text identified and
presented in Table 1 are,</p>
        <sec id="sec-3-1-1">
          <title>1. Statistical metrics.</title>
          <p>Statistical terms such as mean, median, standard
deviation are relevant in reporting results but are
not considered as outcomes themselves.
2. Modification or descriptive Part-Of-Speech (POS).</p>
          <p>Comparative POS such as adjectives, conjunctions
and adverbs were captured as part of the sequence
of words in outcome-spans. E.g. Lower in Lower
maternal attachment can also be higher which are
comparative adjectives describing the change as
applied to an outcome maternal attachment.</p>
          <p>Flaw 2: Failure to identify independent or rather granular
out- comes. This was observed across the following,
1. Multiple outcomes annotated as a single outcome.</p>
          <p>Some outcome-spans were captured as a sequence</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3https://brat.nlplab.org/installation.html 4https://prodi.gy/</title>
          <p>Incorrectly captured Outcome
1. mean arterial blood pressure
2. median Survival
1. Improved ADHD symptoms
2. Lower maternal attachment
arterial blood pressure
Survival
ADHD symptoms
maternal attachment
of distinct outcomes syntactically separated by
either logical conjunctions (and/or) or punctuation
characters such as commas, full and semi-colons.
2. Outcomes co-joined by a dependency term.</p>
          <p>These included outcome-spans that depicted two or
more distinct but related outcomes. e.g. Systolic
and Diastolic blood pressure represents two
different but related outcomes and POSas Table 2 below
indicates,
Incorrectly captured Outcome
Correct Outcome
cardiovascular
events(myocardial infarction, stroke
andcardiovascular death)
1. myocardial infarction
2. stroke
3. cardiovascular death
Systolic and Diastolic
bloodpressure
1. Systolic blood pressure
2. Diastolic blood pressure</p>
          <p>Flaw 3: Capturing Measurement tools, metrics and results
as outcomes.
scores within Work-related stress scores is metric result
reported during RCTs, but the outcome itself is
Workrelated stress. Other examples may include tools such
as questionnaires and tests used in RCTs. Examples are
shown in Table 3,
Incorrectly captured Outcome
Correct Outcome
1. Quality of life Questionnaire
2. Work-related stress scores
3. Weight-test
Quality of life
Work-related stress
Weight
Flaw 4: Imprecise outcome annotations resulting from
inadequate domain knowledge of annotators, Examples in
Table 4.</p>
          <p>1. Non-outcomes incorrectly captured e.g. Severity,</p>
          <p>Effect sizes, significant improvement.
2. Misrepresented outcome types, especially in the</p>
          <p>Mortality outcome type.</p>
          <p>Outcome-span</p>
          <p>Incorrect Type</p>
          <p>Correct Type
Nauseas and Vomiting
suicidal ideations</p>
          <p>Mortality
Mortality</p>
          <p>Physical
Mental
Flaw 5: Combining annotations of outcomes in non-human
studies together with those in human studies.</p>
          <p>Despite the validity of outcomes in non-human species,
they ought to be separately annotated. Example, time
needed to treat commercial beef cattle is an outcome
extracted from non-human medical abstracts included in
outcome annotations for human medical abstracts.
4
4.1</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Proposed Hybrid approach to correct outcome annotations</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Part-Of-Speech Tagging</title>
      <p>Biomedical NLP is supported by a number of POS taggers.
These include, MedPost/SKR Tagger, which was trained on
5,700 manually tagged Medline sentences achieving 97.43%
accuracy on a test set of 1000 sentences [Smith et al., 2004].
GENIA tagger also reported accuracies higher than 97% in
POS tagging sentences from various combination of Genia
corpus5, Wall Street Journal and PennBioIE datasets
[Tsuruoka et al., 2005]. In our approach, we use a POS
tagger available in spaCY6, a SOTA NLP industry-scale
library [Honnibal and Montani, 2017] for advanced NLP. We
train the tagger on the Medpost corpus, publicly available
corpus containing 6,700 Medline sentences annotated with
60 POS tags [Smith et al., 2004]. The trained tagger
subsequently assigned POS tags to every individual word in our
dataset. The model conforms to Penn Treebank POS tagging
guidelines, with a few adjustments that include,</p>
      <p>All words that ended with ‘+’ such as CIN2+ were
assigned noun tags, ‘NN’. This catered for some medical
compounds/substances with a similar syntax that could
have not appeared in the training set.</p>
      <p>Punctuation symbols such as period (.), single
quotation (’) and semi-colon (;) were eliminated because the
EBM-NLP dataset had several of these as redundant
punctuation tokens.</p>
      <p>Square brackets retained their syntax as a POS tag i.e.
‘[’ and ‘]’ were tagged as ‘[’ and ‘]’ respectively.
4.2</p>
    </sec>
    <sec id="sec-5">
      <title>Dealing with Statistical Terms</title>
      <p>Statistical terms within outcomes were eliminated
irrespective of their position in the outcome-spans. These terms were
referenced from a couple of sources including the
international institute of statistics glossary7 and another in a book
for medical device clinical trials [Abdel-Aleem and
Abdelaleem, 2009].
4.3</p>
    </sec>
    <sec id="sec-6">
      <title>Rule-based Chunking</title>
      <p>The chunking algorithm (chunker) relies on a set of rules
to determine where the chunk of interest (correct
outcomespan) begins and ends. These rules are handcrafted
linguistic constraints created to influence the capturing of sequences
of words relevant to an outcome within the incorrect
crowdsourced outcome-spans. Exposed to outcome-spans from the
5https://www.nlm.nih.gov/bsd/medline.html
6https://spacy.io/usage/training
7http://isi.cbs.nl/glossary/bloken00.htm
previous step, this chunker uses underlying syntactical
patterns known as regular expressions to programmatically
extract one or more sub text-spans that constitute of the
actual outcome-span of interest. For example, given an
incorrect outcome-span such as “lower JJR maternal JJ
attachment NN”. Based on one of the predefined constraints
below that suggests removal of comparative POS such as
comparative adjectives tagged ’JJR’, the chunker uses the
positional information of word tagged with the unwanted POS i.e.
“lower JJR” to strip it off and retain “maternal attachment”
as the outcome. Below is a list of chunking constraints used,
Penalizing POS tags including TO (infinitive marker),
II (Preposition), CC (coordinating conjunction) and DD
(determiner).</p>
      <p>Words tagged with these POS tags were deemed
irrelevant and therefore removed when they were located at,
– Start or end of outcome-spans. e.g the DD
memory NN loss NN.
– First or last within a two-worded outcome-span.e.g.</p>
      <p>and CC fatigue NN.
– Every position in an outcome-span, i.e. all words
tagged with a mixture of only these.</p>
      <p>Eliminating contextually comparative or quantification
terms from start or ending positions of an outcome-span
sequence. Comparative terms included comparative
adjectives and adverbs with tags JJR and RRR such as
longer and better resp, then superlative adjectives and
adverbs with tags JJT and RRT such as highest and
most. We additionally considered a set of terms
depicting quantity and their synonyms extracted from
WordNet [Miller, 1995]. These included total, average,
increase and decrease.</p>
      <p>Removing unnecessary word sequences at the start of
outcome-spans. Unwanted starting sequence included
(NNS II) or (NNS DD) or (NNS TO) e.g. predictors of
is unnecessary in predictors NNS of II sex NN risk NN
behavior NN, and so is changes in changes NNS in II
BNP NN
Splitting long outcomes via ‘CC’(coordinating
conjunction) and ‘,’(comma) POS tags. e.g. Serum NN
folate NN and CC vitamin NN B12 NN is split at and CC.
Stripping off square, curved or curly brackets wrapped
around outcome-spans. The content would then be
subjected to processing outlined by all the above
constraints.</p>
      <p>Outcome-spans with a sequence of words tagged as
nouns were preserved. e.g. platelet NN
thromboxane NN formation NN.
5</p>
      <sec id="sec-6-1">
        <title>Dataset and Experiments</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5.1 POS tagging and Rule-based Chunking</title>
      <p>Initially, experiments for POS tagging and Chunking are
performed on ca.70,000 outcome-spans extracted from the
EBM-NLP corpus comprising of ca.5,000 abstracts
describing RCTs annotated in detail with PICO elements8. The</p>
      <sec id="sec-7-1">
        <title>8https://ebm-nlp.herokuapp.com/index</title>
        <p>bi-LSTM</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Model</title>
      <sec id="sec-8-1">
        <title>Baseline (SVM)</title>
        <p>CNN
LSTM</p>
        <p>MNB
bi-LSTM (BM)</p>
      </sec>
      <sec id="sec-8-2">
        <title>BM - Flaw 1 BM - Flaw 2 BM - Flaw 3 BM - Flaw 4</title>
        <p>Adverse-effects
[4489/1593]
outcome-spans are annotated with six outcome types namely</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Adverse-effects, Mental, Mortality, Pain, Physical and</title>
      <p>Other. Applying techniques and constraints narrated from
section 4.1 to 4.3 narrows down the dataset to ca. 32,0009.
5.2</p>
    </sec>
    <sec id="sec-10">
      <title>Classification Model</title>
      <p>Three different neural network architectures: Long
ShortTerm Memory (LSTM) [Sundermeyer et al., 2012],
Convolutional Neural Network (CNN) [Kim, 2014], Bidirectional
Long Short-Term Memory (bi-LSTM) [Zhang et al., 2015] as
well as two bag-of-word models: Support Vector Machines
(SVM) [Wang et al., 2006] and Multinomial Naive Bayes
(MNB) [Frank and Bouckaert, 2006] are adopted to perform
classification both on the initial extract of outcome-spans ca.
70,000 and the corrected outcome-spans ca.32,000.</p>
      <p>Given a training-set, O = f(Xt; yt)gtT=1, where Xt is an
instance of an outcome-span defined as sequence of words i.e.
x = (o1; o2; : : : ; oN ) where N is the length of the
outcomespan sentence to be classified and each on is a 50-dimensional
word embedding for the respective word in the outcome-span,
Embeddings are obtained using pre-trained 840B 300d GloVe
word vectors [Pennington et al., 2014]. yt is a one-hot vector
for the corresponding label. The goal is to learn a classifier
f : X ! Y .</p>
      <p>In all experiments, five-fold cross validation is used for
evaluation, with a batch-size of 500, trained for 100 epochs
and a drop-out of 0.2 for each single fold. Note: The
bagof-words models take as input, a tf-idf vector [Yun-tao et al.,
2005] representation of the sequence of words9.
5.3</p>
    </sec>
    <sec id="sec-11">
      <title>Evaluating the experiment results</title>
      <p>Results presented in Table 6 indicate that the accuracies
increased after correcting the errors in the outcome-spans.
Moreover, the increase was not only consistent across the
five different models used, but even across prediction of the
six classes in the dataset. Notably, the bi-LSTM outperforms
all the other models, however, the bag-of-words SVM model
seems to achieve the second-highest scores. This suggests
that neural networks are most effective in learning
representations of medical literature such as outcomes in this study.
9https://github.com/MichealAbaho/pico-outcome-prediction
5.4</p>
    </sec>
    <sec id="sec-12">
      <title>Flaw Analysis</title>
      <p>In-order to examine the impact the flaws individually had
on the classification performance, the flaw correction
process was broken down to independently cater for the
different flaws one by one. The best performing model (bi-LSTM)
would then be tested on input data where only annotations
with flaw 1 had been corrected and the rest ignored. This was
repeatedly done for flaws 2, 3 and 4 as reported in the bottom
half of Table 5. Flaw 5 was not considered in this additional
analysis because of the extremely few cases it was
responsible for. Despite the largely analogous results, We observed
that corrections targeted to fix Flaw 2 alone, had a
significantly higher impact on the performance, scoring higher
F1scores for the six classes with the exception of the Physical
class. This implied that, granularity and distinctness is vitally
important when detecting not just outcomes but any relevant
clinical entities in biomedical literature. Nonetheless, neither
of the F1-scores in this analysis would match up to the
originally obtained F1-scores with all flaws corrected (line 5
Table 5).
5.5</p>
    </sec>
    <sec id="sec-13">
      <title>Conclusion</title>
      <p>Manually annotating medical data is a challenging and costly
process. As a result, crowdsourced annotations are often
noisy and inconsistent. This work performs a sanity check
on crowdsourced annotations in a public corpus EBM-NLP
revealing various flaws in the annotations. We train a spaCY
POS tagging model on Medline articles and use a rule based
chunking algorithm to fix these errors/flaws. Classification
experiments at the end justify the positive impact our
corrective approach has on the dataset.</p>
      <p>As part of future work, we aim to explore dependency
graphs to capture disjoint or entities to achieve required
granularity in outcome reporting. For instance, the outcome, chest
and abdominal pain is best detected as two independent
outcomes,chest pain and abdominal pain where pain is simply a
disjoint entity. We shall further on adopt expertly annotated
data to maximize precision, recall and quality of ground-truth
annotations and thereby, utilize transfer learning to
automatically detect outcomes. Upon satisfactorily achieving quality
annotations, we shall utilize semi-supervised learning to build
a corpus of outcomes ready to support NLP tasks in EBM.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [
          <string-name>
            <surname>Abdel-Aleem and</surname>
          </string-name>
          Abdel-aleem,
          <year>2009</year>
          <article-title>] Salah Abdel-Aleem and Salah Abdel-aleem. Design, execution, and management of medical device clinical trials</article-title>
          .
          <source>Wiley Online Library</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[Bruno</source>
          , 2018]
          <string-name>
            <given-names>Godefroy</given-names>
            <surname>Bruno</surname>
          </string-name>
          .
          <article-title>On the viability of crowdsourcing nlp annotations in healthcare</article-title>
          ,
          <year>July 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Dodd et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Susanna</given-names>
            <surname>Dodd</surname>
          </string-name>
          , Mike Clarke, Lorne Becker, Chris Mavergames, Rebecca Fish, and
          <string-name>
            <surname>Paula</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Williamson</surname>
          </string-name>
          .
          <article-title>A taxonomy has been developed for outcomes in medical research to help improve knowledge discovery</article-title>
          .
          <source>Journal of Clinical Epidemiology</source>
          ,
          <volume>96</volume>
          :
          <fpage>84</fpage>
          -
          <lpage>92</lpage>
          , 4
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[Frank and Bouckaert</source>
          , 2006]
          <string-name>
            <given-names>Eibe</given-names>
            <surname>Frank and Remco R Bouckaert.</surname>
          </string-name>
          <article-title>Naive bayes for text classification with unbalanced classes</article-title>
          .
          <source>In European Conference on PKDD</source>
          , pages
          <fpage>503</fpage>
          -
          <lpage>510</lpage>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Honnibal and Montani</source>
          , 2017]
          <article-title>Matthew Honnibal and Ines Montani. spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing</article-title>
          . To appear,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Huang et al.,
          <year>2006</year>
          ]
          <string-name>
            <given-names>Xiaoli</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Lin</surname>
          </string-name>
          , and
          <string-name>
            <surname>Dina</surname>
          </string-name>
          Demner-Fushman.
          <article-title>Evaluation of PICO as a knowledge representation for clinical questions</article-title>
          .
          <source>AMIA ... Annual Symposium proceedings. AMIA Symposium</source>
          , pages
          <fpage>359</fpage>
          -
          <lpage>63</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Kim</source>
          , 2014]
          <string-name>
            <given-names>Yoon</given-names>
            <surname>Kim</surname>
          </string-name>
          .
          <article-title>Convolutional neural networks for sentence classification</article-title>
          .
          <source>arXiv preprint arXiv:1408.5882</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Miller</source>
          ,
          <year>1995</year>
          ]
          <article-title>George A Miller. Wordnet: a lexical database for english</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>11</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Nghiem and Ananiadou</source>
          , 2018]
          <article-title>Minh-Quoc Nghiem</article-title>
          and
          <string-name>
            <given-names>Sophia</given-names>
            <surname>Ananiadou</surname>
          </string-name>
          .
          <article-title>Aplenty: annotation tool for creating high-quality datasets using active and proactive learning</article-title>
          .
          <source>In Proc. of the EMNLP'14: System Demonstrations</source>
          , pages
          <fpage>108</fpage>
          -
          <lpage>113</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Nye et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Nye</surname>
          </string-name>
          ,
          <source>Junyi Jessy Li</source>
          , Roma Patel, Yinfei Yang,
          <string-name>
            <given-names>Iain J.</given-names>
            <surname>Marshall</surname>
          </string-name>
          , Ani Nenkova, and
          <string-name>
            <surname>Byron</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Wallace</surname>
          </string-name>
          .
          <article-title>A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature</article-title>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Pennington et al.,
          <year>2014</year>
          ] Jeffrey Pennington, Richard Socher, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          . Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In Proc. of the EMNLP'14</source>
          , pages
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Sackett et al.,
          <year>1996</year>
          ] David L Sackett,
          <string-name>
            <surname>William M C Rosenberg</surname>
            ,
            <given-names>J A</given-names>
          </string-name>
          <string-name>
            <surname>Muir</surname>
            <given-names>Gray</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R Brian</given-names>
            <surname>Haynes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W Scott</given-names>
            <surname>Richardson</surname>
          </string-name>
          .
          <article-title>Evidence based medicine: what it is and what it isn't</article-title>
          . BMJ,
          <volume>312</volume>
          (
          <issue>7023</issue>
          ):
          <fpage>71</fpage>
          -
          <lpage>72</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Santos et al.,
          <year>2007</year>
          ]
          <article-title>Cristina Mame´dio da Costa Santos</article-title>
          ,
          <string-name>
            <surname>Cibele Andrucioli de Mattos Pimenta</surname>
          </string-name>
          , and
          <article-title>Moacyr Roberto Cuce Nobre. The pico strategy for the research question construction and evidence search</article-title>
          . Revista latinoamericana de enfermagem,
          <volume>15</volume>
          (
          <issue>3</issue>
          ):
          <fpage>508</fpage>
          -
          <lpage>511</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>[Smith</surname>
          </string-name>
          et al.,
          <year>2004</year>
          ]
          <string-name>
            <given-names>L</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Rindflesch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W John</given-names>
            <surname>Wilbur</surname>
          </string-name>
          . Medpost:
          <article-title>a part-of-speech tagger for biomedical text</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>20</volume>
          (
          <issue>14</issue>
          ):
          <fpage>2320</fpage>
          -
          <lpage>2321</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Sundermeyer et al.,
          <year>2012</year>
          ]
          <string-name>
            <given-names>Martin</given-names>
            <surname>Sundermeyer</surname>
          </string-name>
          , Ralf Schlu¨ter, and Hermann Ney.
          <article-title>Lstm neural networks for language modeling</article-title>
          .
          <source>In Thirteenth annual conference of the international speech communication association</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Tsuruoka et al.,
          <year>2005</year>
          ]
          <string-name>
            <given-names>Yoshimasa</given-names>
            <surname>Tsuruoka</surname>
          </string-name>
          , Yuka Tateishi,
          <string-name>
            <surname>Jin-Dong</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Tomoko Ohta,
          <string-name>
            <surname>John McNaught</surname>
            ,
            <given-names>Sophia</given-names>
          </string-name>
          <string-name>
            <surname>Ananiadou</surname>
          </string-name>
          , and
          <article-title>Jun'ichi Tsujii. Developing a robust partof-speech tagger for biomedical text</article-title>
          .
          <source>In Panhellenic Conference on Informatics</source>
          , pages
          <fpage>382</fpage>
          -
          <lpage>392</lpage>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>[Vittayakorn and Hays</source>
          , 2011]
          <string-name>
            <given-names>Sirion</given-names>
            <surname>Vittayakorn</surname>
          </string-name>
          and
          <string-name>
            <given-names>James</given-names>
            <surname>Hays</surname>
          </string-name>
          .
          <article-title>Quality assessment for crowdsourced object annotations</article-title>
          .
          <source>In BMVC</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>[Wang</surname>
          </string-name>
          et al.,
          <year>2006</year>
          ]
          <string-name>
            <surname>Zi-Qiang</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Xia Sun,
          <string-name>
            <surname>De-Xian Zhang</surname>
            , and
            <given-names>Xin</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>An optimal svm-based text classification algorithm</article-title>
          .
          <source>In 2006 International Conference on Machine Learning and Cybernetics</source>
          , pages
          <fpage>1378</fpage>
          -
          <lpage>1381</lpage>
          . IEEE,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Williamson et al.,
          <year>2017</year>
          ] Paula R Williamson, Douglas G Altman,
          <article-title>Heather Bagley</article-title>
          , Karen L Barnes,
          <string-name>
            <surname>Jane M Blazeby</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sara T Brookes</surname>
            , Mike Clarke, Elizabeth Gargon, Sarah Gorst,
            <given-names>Nicola</given-names>
          </string-name>
          <string-name>
            <surname>Harman</surname>
          </string-name>
          , et al.
          <source>The comet handbook: version 1</source>
          .0. Trials,
          <volume>18</volume>
          (
          <issue>3</issue>
          ):
          <fpage>280</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [Xu et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Boyi</given-names>
            <surname>Xu</surname>
          </string-name>
          , Ke Xu, LiuLiu Fu,
          <string-name>
            <given-names>Ling</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Weiwei</given-names>
            <surname>Xin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Hongming</given-names>
            <surname>Cai</surname>
          </string-name>
          .
          <article-title>Healthcare data analytics: using a metadata annotation approach for integrating electronic hospital records</article-title>
          .
          <source>Journal of Management Analytics</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ):
          <fpage>136</fpage>
          -
          <lpage>151</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Yang et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Yinfei</given-names>
            <surname>Yang</surname>
          </string-name>
          , Oshin Agarwal, Chris Tar, Byron C Wallace, and
          <string-name>
            <given-names>Ani</given-names>
            <surname>Nenkova</surname>
          </string-name>
          .
          <article-title>Predicting annotation difficulty to improve task routing and model performance for biomedical information extraction</article-title>
          . arXiv preprint arXiv:
          <year>1905</year>
          .07791,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [Young et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Tom</given-names>
            <surname>Young</surname>
          </string-name>
          , Devamanyu Hazarika, Soujanya Poria, and
          <string-name>
            <given-names>Erik</given-names>
            <surname>Cambria</surname>
          </string-name>
          .
          <article-title>Recent trends in deep learning based natural language processing</article-title>
          .
          <source>ieee Computational intelligence magazine</source>
          ,
          <volume>13</volume>
          (
          <issue>3</issue>
          ):
          <fpage>55</fpage>
          -
          <lpage>75</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>[</surname>
          </string-name>
          Yun-tao et al.,
          <year>2005</year>
          ] Zhang Yun-tao, Gong Ling, and
          <string-name>
            <surname>Wang</surname>
          </string-name>
          Yong-cheng.
          <article-title>An improved tf-idf approach for text classification</article-title>
          .
          <source>Journal of Zhejiang University-Science A</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>49</fpage>
          -
          <lpage>55</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Zhang et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Shu</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Dequan Zheng, Xinchen Hu, and
          <string-name>
            <given-names>Ming</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>Bidirectional long short-term memory networks for relation classification</article-title>
          .
          <source>In Proc of the 29th PACLIC'2015</source>
          , pages
          <fpage>73</fpage>
          -
          <lpage>78</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>