<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparing Word Embeddings for Document Screening based on Active Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andres Carvallo</string-name>
          <email>R@k</email>
          <email>afcarvallo@uc.cl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Denis Parra[</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department Ponti cia Universidad Catolica de Chile Santiago</institution>
          ,
          <country country="CL">Chile</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Document screening is a fundamental task within Evidencebased Medicine (EBM), a practice that provides scienti c evidence to support medical decisions. Several approaches are attempting to reduce the workload of physicians who need to screen and label hundreds or thousands of documents in order to answer speci c clinical questions. Previous works have attempted to semi-automate document screening, reporting promising results, but their evaluation is conducted using small datasets, which hinders generalization. Moreover, some recent works have used recently introduced neural language models, but no previous work have compared, for this task, the performance of di erent language models based on neural word embeddings, which have reported good results in the latest years for several NLP tasks. In this work, we evaluate the performance of two popular neural word embeddings (Word2vec and GloVe) in an active learning-based setting for document screening in EBM, with the goal of reducing the number of documents that physicians need to label in order to answer clinical questions. We evaluate these methods in a small public dataset (HealthCLEF 2017) as well as a larger one (Epistemonikos). Our experiments indicate that Word2vec have less variance and better general performance than GloVe when using active learning strategies based on uncertainty sampling.</p>
      </abstract>
      <kwd-group>
        <kwd>active learning</kwd>
        <kwd>evidence based medicine</kwd>
        <kwd>document screen- ing</kwd>
        <kwd>word embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Evidence-based Medicine (EBM) is a practice that provides scienti c evidence
to support medical decisions. This evidence nowadays is obtained from
biomedical journals, usually accessible through the portal PubMed1, a search engine
which provides free access to abstracts of biomedical research articles as well as
to the MEDLINE database. An existing problem is to nd relevant documents
within a massive volume of documents, given a clinical question or a query. As
a consequence of this, the time required for the search and screening of articles</p>
    </sec>
    <sec id="sec-2">
      <title>1 https://www.ncbi.nlm.nih.gov/pubmed/</title>
      <p>
        related to clinical questions about medical problems can take long and
sometimes it consumes a large part of a physician's workday [
        <xref ref-type="bibr" rid="ref15 ref6">15, 6</xref>
        ]. When people
conduct this repetitive task, there is a good chance of overlooking important
articles, which can have a negative impact on decisions such as the patient's
treatment [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Moreover, the publication of medical papers has grown
exponentially the last decade. Since 2005, PubMed has indexed more than 1 million
articles per year, which means that the process of searching and manual
screening of medical evidence will become increasingly more di cult for physicians
without the support of information retrieval and machine learning algorithms.
For this reason, some systems have emerged to support experts in the collection
of evidence such as Embase2, DARE3 and Epistemonikos4. In this article, we
work with data from Epistemonikos, which helps expert physicians to review
and validate scienti c evidence grouped by medical questions to facilitate its
subsequent search. Our goal is to improve the e ciency and e cacy of
document screening in the practice of EBM. In other words, we aim at reducing the
e ort made by physicians at screening documents to nd the evidence needed to
support the answers of a medical question. We use an active learning approach,
experimenting with a large dataset of medical questions, unlike previous works
which use very small datasets, some of them very recent [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In this short
paper, we contribute by: i ) Experimenting in both a large dataset (Epistemonikos,
987) and a small dataset (CLEF, 50), showing evidence of generalization of our
approaches, and ii ) comparing the performance of documents represented with
two state-of-the-art neural word embeddings (Word2vec [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and GloVe [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]) as
well as traditional relevance feedback [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
2
      </p>
      <sec id="sec-2-1">
        <title>Related Work</title>
        <p>
          The task of nding relevant documents related to a medical question through
citation screening has been studied and it is known as the total recall problem:
given a medical topic or question, nd all the documents that are relevant about
a particular topic. Recently, the CLEF task 2 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] is a challenge that calls for
solving the problem of prioritizing which documents to screen to reduce work
overload for experts. They provide a public dataset with medical topics and a
set of candidate documents; participants have to rank documents by relevance
for every speci c medical subject in the minimum of iterations to make more
e cient the document screening process. In the literature, the approaches to
solving this problem are based on two general lines: information retrieval and
machine learning methods.
        </p>
        <p>
          In the information retrieval area, there have been many attempts to solve the
problem using techniques such as relevance feedback [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], query expansion [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ],
ranking and inference based on external knowledge [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. However, they do not
2 https://www.elsevier.com/solutions/embase-biomedical-research
3 https://www.crd.york.ac.uk/CRDWeb/
4 https://www.epistemonikos.org/en
ensure a level of recall necessary to capture all the evidence related to a medical
question.
        </p>
        <p>
          From the machine learning community, the approach is to automate or
semiautomate the screening process or review of medical articles that were previously
selected as relevant to a medical question by learning the pattern of physicians
conducting a document survey. There have been e orts to solve this problem
by using automatic classi cation [
          <xref ref-type="bibr" rid="ref1 ref16 ref2 ref21 ref3">2, 3, 1, 16, 21</xref>
          ]. Where they compared classi ers
such as Naive Bayes, K-NN, and SVM, using di erent ways to represent text,
such as word embeddings and bag-of-clinical terms from titles and abstracts.
There is also literature that has used active learning [
          <xref ref-type="bibr" rid="ref15 ref22 ref7 ref9">9, 7, 22, 15</xref>
          ] for medical topic
detection and clinical text classi cation. Moreover, a few of deep learning models
have been proposed for the classi cation of relevant evidence and categorization
of documents in medical questions [
          <xref ref-type="bibr" rid="ref10 ref4">4, 10</xref>
          ]. Generally, the majority of work done
has used datasets of up to 50 medical topics/questions and 200,000 documents,
and in this case, we work with a dataset close to 1; 000 medical questions and
370; 000 potential documents, allowing models to generalize and obtain better
e cacy results compared to the state of the art. In addition, unlike previous
work, we compare two neural word embedding models for document
representation (Word2vec [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and GloVe [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]) in order to asses their performance for the
biomedical document screening task.
3
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Proposed Solution</title>
        <p>The process of nding documents that answer a clinical question requires rst
retrieving a set of candidate documents. Then, physicians perform the document
screening where they verify that abstracts and titles of each document are related
to the medical question, this particular process may involve much time and
cognitive e ort from experts.</p>
        <p>Problem formulation: Given a medical question q and a set of candidate
documents C = fc1; c2; :::cng we need to ask an oracle (physician) O to label
these documents as relevant or not relevant to q. We want to avoid asking the
labeling of every document, so we select an informative sample to be labeled
by the expert. With these labels, we train a predictive model M . It might be
necessary to ask for labels in many iterations in order to re ne the model, ending
up with several models M0; M1; :::; Mk.</p>
        <p>
          In our case, we use an active learning (AL) approach [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Using an AL strategy
A (e.g., uncertainty sampling, query-by-committee, etc.), we sample a set of
unlabeled documents X from C, in order to ask the oracle O to label them.
With the labeled items, we then train a machine learning model Mi(X; Y ) with
the new observations X C with labels Y (binary, Y = 1 means relevant
document and Y = 0 means not relevant) given by O. Then, we use the trained
model Mi to predict relevance labels for unobserved documents, and using the
active learning strategy A we select new items to be labeled by O in order to
create an updated version of the model Mi+1. In each iteration, we evaluate the
model (precision, recall, ), and we can stop until a xed number of iterations or
after the model converges.
        </p>
        <p>To address the problem above, we developed a system, where we start with
a small proportion of labeled documents as relevant or not relevant for each
medical question to train a rst version of the machine learning model Mi.
Then, using the active learning strategy we chose instances to be labeled by a
physician based on the title and abstract text features represented internally as
word embeddings (GloVe and Word2vec). After the physician adds the labels,
they are used to train a machine learning model Mi+1 to predict the relevance
of new unlabeled documents and thus begin a new iteration.</p>
        <p>
          The performance of our approach rst depends on the machine learning
algorithm chosen and second, on the active learning strategy that chooses unlabeled
examples to create a labeled dataset as input for supervised learning algorithms.
The strategies used in this experiment are uncertainty sampling and random
sampling, given their lower complexity compared to others such as error-based,
gradient-based and variable reduction [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. The machine learning algorithms
that are considered to be trained with new labeled examples are random forests,
logistic regression, and neural networks.
        </p>
        <p>With respect to the active learning sampling strategies, random sampling: chooses
random documents to train the machine learning model and it is usually used as
a comparison baseline against other approaches. On the other side, uncertainty
sampling: looks for records which have higher label prediction uncertainty,
making them potentially more informative for collecting their actual labels and then
training or updating a model.
4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Experiments</title>
        <p>Dataset. For the experiments, we used two datasets: CLEF5 and Epistemonikos.
Both of them have a similar distribution of documents per question, where the
majority of medical questions contain an approximate number of 200 relevant
documents. On the one hand, CLEF dataset contains only 50 medical questions
and 200,000 documents related to them that were crawled from PubMed using
each document id. On the other hand, the Epistemonikos Evidence Synthesis
Project is a collaborative initiative established in 2012 with the objective of
collecting, organizing and comparing all relevant evidence for health
decisionmaking, through a multilingual platform. This dataset is composed of 987
medical questions and 372,829 potential documents. In both datasets each medical
question is associated to a Systematic Review (hereinafter, SR), which is a type
of article that collects and synthesizes the most relevant primary studies and
trials related to a question. The information of documents from both datasets
consists of the title, abstract, author, year and the label if it is relevant (or not)
to the question or medical subject. In the case of Epistemonikos data, the labels
were previously curated by senior medical students, in which they had to select
papers related to a set of medical questions. Document representation: for each</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5 https://sites.google.com/site/clefehealth2017/task-2</title>
      <p>
        document we lower case the concatenation of title and abstract, then remove
stop words, and we use GloVe [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and Word2vec [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to obtain an embedding
representation of 300 dimensions of each word. Finally, using average pooling we
obtain a vector for each document.
4.1
      </p>
      <p>
        O ine Active Learning Setup
We experimented doing a simulation of the active learning labeling process of
documents for medical questions. As each medical question has a di erent
number of relevant documents, we sample documents that are not relevant where
the total of relevants corresponds to the 5%, so that the distribution of
documents is similar to the CLEF dataset [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. We ltered out some of the medical
questions, keeping those that have more than ve relevant documents and less
than 2,000 relevant documents, ending up with 987. We compared the results of
applying active learning on the CLEF dataset that contains 50 SR (Systematic
Reviews) with Epistemonikos. For each medical question, we hide the document
labels and we leave only ve with their respective labels to start building the
model and then iterating with active learning to receive feedback from the
oracle. For each prediction made by the machine learning model in each iteration,
we sorted the results depending on the predicted probability of being relevant
for each model, so the evaluation metrics were calculated with the ranked list of
potential candidates given by each strategy. The parameters chosen for machine
learning algorithms were: for neural networks we used ve hidden layers, ReLu
activation function, learning rate of 1e-05, momentum of 0.9, 100 neurons per
layer and Adam optimization function. For the random forest, we used 100
estimators. Experiments were programmed in Python3 using libact [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], scikit-learn
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], pandas and gensim libraries. Code for these experiments will be published
in a github repository after noti cation.
      </p>
      <p>
        Evaluation metrics. We evaluated our proposed active learning strategies with
traditional IR metrics also used by Lee et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]: precision@k, recall@k and mean
average precision (MAP). We report the metrics obtained after ten iterations,
with ten documents labeled per iteration.
5
      </p>
      <sec id="sec-3-1">
        <title>Results and Discussion</title>
        <p>Average Precision (MAP) performance measured in Epistemonikos and CLEF datasets using active
learning strategies (US: uncertainty sampling, RS: random sampling) using a batch of 10 documents
per feedback iteration for Word2vec and GloVe representation.</p>
        <p>Dataset
Epistemonikos</p>
        <p>987 SRs
GloVe 300 dim
Epistemonikos</p>
        <p>987 SRs
Word2vec 300 dim</p>
        <p>CLEF
50 SRs
GloVe 300 dim</p>
        <p>CLEF
50 SRs
Word2vec 300 dim</p>
        <p>AL-Model</p>
        <p>MAP
US-NN
US-RF
US-LR
RS-LR
US-NN
US-RF
US-LR
RS-LR
US-NN
US-RF
US-LR
RS-LR
US-NN
US-RF
US-LR
RS-LR</p>
      </sec>
      <sec id="sec-3-2">
        <title>6 Conclusion and Future Work</title>
        <p>
          In this article we supported results from previous studies in terms of showing
that active learning with an uncertainty sampling strategy yields good results
for the task of biomedical document screening. Moreover, we contribute by
comparing two popular word embeddings to represent documents: Word2vec and
GloVe. The best results were obtained using Word2vec document representation
and random forests as the learning algorithm. GloVe document representation
also yields competitive results, but it seems more sensitive two the classi cation
model used: it performs well with random forests but shows poor performance
with neural networks and logistic regression. Moreover, our experiments indicate
that these results are consistent in both the small public dataset of HealthCLEF
and the larger dataset of Epistemonikos, giving evidence of generalization.
For future work, we will try other machine learning models, active learning
strategies and evaluate the results using CLEF metrics [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. We will also test
other paradigms for more scalable learning, such as weak supervision. With
respect, to embeddings, we will test di erent values of sensitive parameters, as
mentioned by Roy et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Finally, we will conduct a user study with actual
physicians in order to evaluate online the performance of our approach.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>7 Acknowledgements</title>
        <p>We acknowledge Epistemonikos Foundation, the Chilean research agency
Conicyt, Fondecyt grant 1191791 and the Millenium Institute IMFD.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Adeva</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atxa</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrillo</surname>
            ,
            <given-names>M.U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zengotitabengoa</surname>
            ,
            <given-names>E.A.</given-names>
          </string-name>
          :
          <article-title>Automatic text classi cation to support systematic reviews in medicine</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>41</volume>
          (
          <issue>4</issue>
          ),
          <volume>1498</volume>
          {
          <fpage>1508</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bekhuis</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tseytlin</surname>
            , E., Mitchell,
            <given-names>K.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Feature engineering and a proposed decision-support system for systematic reviewers of medical evidence</article-title>
          .
          <source>PloS one 9</source>
          (
          <issue>1</issue>
          ),
          <year>e86277</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ryu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Combining relevancy and methodological quality into a single ranking for evidence-based medicine</article-title>
          .
          <source>Information Sciences</source>
          <volume>214</volume>
          ,
          <volume>76</volume>
          {
          <fpage>90</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Del</given-names>
            <surname>Fiol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Michelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Iorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Cotoi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Haynes</surname>
          </string-name>
          , R.B.:
          <article-title>A deep learning method to automatically identify reports of scienti cally rigorous clinical research from the biomedical literature: Comparative analytic study</article-title>
          .
          <source>J Med Internet Res</source>
          <volume>20</volume>
          (
          <issue>6</issue>
          ) (
          <year>Jun 2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Donoso-Guzman</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parra</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>An interactive relevance feedback interface for evidence-based health care</article-title>
          .
          <source>In: 23rd International Conference on Intelligent User Interfaces</source>
          . pp.
          <volume>103</volume>
          {
          <fpage>114</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Elliott</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turner</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clavisi</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Higgins</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mavergames</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gruen</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          :
          <article-title>Living systematic reviews: an emerging opportunity to narrow the evidence-practice gap</article-title>
          .
          <source>PLoS medicine 11(2)</source>
          ,
          <year>e1001603</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Figueroa</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng-Treitler</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngo</surname>
            ,
            <given-names>L.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goryachev</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiechmann</surname>
            ,
            <given-names>E.P.</given-names>
          </string-name>
          :
          <article-title>Active learning for clinical text classi cation: is it better than random sampling</article-title>
          ?
          <source>Journal of the American Medical Informatics Association</source>
          <volume>19</volume>
          (
          <issue>5</issue>
          ),
          <volume>809</volume>
          {
          <fpage>816</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Goodwin</surname>
            ,
            <given-names>T.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harabagiu</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          :
          <article-title>Knowledge representations and inference techniques for medical question answering</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology (TIST) 9</source>
          (
          <issue>2</issue>
          ),
          <volume>14</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hashimoto</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontonatsios</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miwa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Topic detection using paragraph vectors to support active learning in systematic reviews</article-title>
          .
          <source>Journal of biomedical informatics 62</source>
          ,
          <volume>59</volume>
          {
          <fpage>65</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hughes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kotoulas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suzumura</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Medical text classi cation using convolutional neural networks</article-title>
          .
          <source>Stud Health Technol Inform</source>
          <volume>235</volume>
          ,
          <issue>246</issue>
          {
          <fpage>50</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azzopardi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spijker</surname>
          </string-name>
          , R.:
          <article-title>Clef 2017 technologically assisted reviews in empirical medicine overview</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          . vol.
          <year>1866</year>
          , pp.
          <volume>1</volume>
          {
          <issue>29</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Keselman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          :
          <article-title>A classi cation of errors in lay comprehension of medical documents</article-title>
          .
          <source>Journal of biomedical informatics 45(6)</source>
          ,
          <volume>1151</volume>
          {
          <fpage>1163</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Seed-driven document ranking for systematic reviews in evidence-based medicine</article-title>
          .
          <source>In: The 41st International ACM SIGIR Conference on Research &amp; Development in Information Retrieval</source>
          . pp.
          <volume>455</volume>
          {
          <fpage>464</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>3111</volume>
          {
          <issue>3119</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Miwa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>OMara-Eves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Reducing systematic review workload through certainty-based screening</article-title>
          .
          <source>Journal of biomedical informatics 51</source>
          ,
          <volume>242</volume>
          {
          <fpage>253</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mo</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontonatsios</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Supporting systematic reviews using lda-based document representations</article-title>
          .
          <source>Systematic reviews 4(1)</source>
          ,
          <volume>172</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , et al.:
          <article-title>Scikit-learn: Machine learning in python</article-title>
          .
          <source>Journal of machine learning research 12(Oct)</source>
          ,
          <volume>2825</volume>
          {
          <fpage>2830</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          . pp.
          <volume>1532</volume>
          {
          <issue>1543</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganguly</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhatia</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bedathur</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Using word embeddings for information retrieval: How collection and term normalization choices a ect performance</article-title>
          .
          <source>In: Proceedings of the 27th ACM International Conference on Information and Knowledge Management</source>
          . pp.
          <year>1835</year>
          {
          <year>1838</year>
          . ACM (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Settles</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Active learning</article-title>
          .
          <source>Synthesis Lectures on Arti cial Intelligence and Machine Learning</source>
          <volume>6</volume>
          (
          <issue>1</issue>
          ),
          <volume>1</volume>
          {
          <fpage>114</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Wallace</surname>
            ,
            <given-names>B.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Small</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brodley</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lau</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bertram</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lill</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>J.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trikalinos</surname>
            ,
            <given-names>T.A.</given-names>
          </string-name>
          :
          <article-title>Toward modernizing the systematic review pipeline in genetics: e cient updating via data mining</article-title>
          .
          <source>Genetics in medicine 14(7)</source>
          ,
          <volume>663</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Wallace</surname>
            ,
            <given-names>B.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Small</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brodley</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trikalinos</surname>
            ,
            <given-names>T.A.</given-names>
          </string-name>
          :
          <article-title>Active learning for biomedical citation screening</article-title>
          .
          <source>In: Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          . pp.
          <volume>173</volume>
          {
          <fpage>182</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chung</surname>
            ,
            <given-names>Y.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>T.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
          </string-name>
          , H.T.:
          <article-title>libact: Poolbased active learning in python (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>