<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>QUT ielab at CLEF eHealth 2017 Technology Assisted Reviews Track: Initial Experiments with Learning To Rank</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Harrisen Scells</string-name>
          <email>harrisen.scells@hdr.qut.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guido Zuccon</string-name>
          <email>g.zuccon@qut.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anthony Deacon</string-name>
          <email>aj.deacon@qut.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bevan Koopman</string-name>
          <email>bevan.koopman@csiro.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Australian E-Health Research Centre, CSIRO</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Queensland University of Technology</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe our participation to the CLEF eHealth 2017 Technology Assisted Reviews track (TAR). This track aims to evaluate and advance search technologies aimed at supporting the creation of biomedical systematic reviews. In this context, the track explores the task of screening prioritisation: the ranking of studies to be screened for inclusion in a systematic review. Our solution addresses this challenge by developing ranking strategies based on learning to rank techniques and exploiting features derived by the use of the PICO framework. PICO (Population, Intervention, Control or comparison and Outcome) is a technique used in evidence based practice to frame and answer clinical questions and is used extensively in the compilation of systematic reviews. Our experiments show that the use of the PICO-based feature within learning to rank provides improvements over the use of baseline features alone.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>A systematic review is a type of literature review that appraises and synthesises
the work of primary research studies to answer one or more research questions.
Most authors follow the Preferred Reporting Items for Systematic Reviews and
Meta-Analyses (PRISMA) method for conducting and reporting these reviews.
This includes the de nition of a formal search strategy to retrieve studies which
are to be considered for inclusion in the review.</p>
      <p>Given a research question and a set of inclusion/exclusion criteria, researchers
undertaking a systematic review de ne a search strategy (the query) to be issued
to one or more search engines that index published literature (e.g. PubMed). In
medical and biomedical research, search strategies are commonly expressed as
(large) boolean queries. After the search strategy has been executed, the title,
and then abstract, of studies retrieved by it are reviewed in a process known as
screening. Where the study appears relevant the full-text is then retrieved for
more detailed examination.</p>
      <p>
        The compilation of systematic reviews can take signi cant time and resources,
hampering their e ectiveness. Tsafnat et al. report that it can take several years
to complete and publish a systematic review [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. When systematic reviews take
such signi cant time to complete, they can become out-of-date even at time of
publishing. While the compilation of a systematic review involves several steps,
one of the most time-consuming is screening. Thus, the development of IR
methods that decrease the number of documents to be screened, would have a major
impact on the time and resources required to undertake systematic reviews.
Similarly, the ordering of studies to be screened according to the likelihood of
satisfying the inclusion criteria of the systematic reviews (screening prioritisation)
would allow relevant studies to be identi ed early on in the screening process,
thus providing a feedback loop to improve the development of search strategies.
Screening prioritisation is typically done as a two-stage process. An initial set
of studies are retrieved using a boolean retrieval process; these are then ranked
according to some relevance measure.
      </p>
      <p>
        The challenge of compiling systematic reviews can be fertile ground for
information retrieval (IR) research, as this can provide techniques to improve current
screening and screening prioritisation processes. The CLEF eHealth 2017
Technology Assisted Reviews track (TAR) [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ] joins our recent work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] in devising
evaluation resources for evaluation of information retrieval techniques that
attempt to automate and improve processes involved in the creation of systematic
reviews. The TAR track considers two tasks: (1) to produce an the e cient
ordering of studies retrieved by a boolean search strategy, such that all of the
relevant abstracts are retrieved as early as possible, and (2) to identify a subset
of the ranked studies which contains all or as many of the relevant abstracts for
the least e ort (i.e. total number of abstracts to be assessed). In our submissions,
we tackle the rst task, and use learning to rank to produce a re-ranking of the
initial set of studies retrieved for screening by the systematic review's boolean
search strategy.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Our Approach for TAR</title>
      <p>We trained a learning to rank model using domain speci c features to provide
an e cient ordering of studies retrieved by a systematic review. Speci cally,
we aim to observe what e ect PICO features have with respect to learning to
rank algorithms. PICO (Population, Intervention, Control or comparison and
Outcome) is a technique used in evidence based practice to frame and answer
clinical questions and is used extensively in the compilation of systematic
reviews. We investigated several learning to rank algorithms and observed the
e ect queries annotated with PICO elements had on the reordering of results
compared to the original Boolean queries.</p>
      <p>
        We trained two learning to rank models using both the original queries
provided by the task organiser, and another modi ed set of queries which contains
annotations from the PICO framework. In total, we used seven features to train
our learning to rank model. Table 1 summarises the features used. The rst four
features (IDFSum, IDFStd, IDFMax, and IDFAvg) calculate the inverse
document frequency (idf) for each of the terms in the document that also appear
Id Feature
1 IDFSum
2 IDFStd
3 IDFMax
4 IDFAvg
5 PopulationCount
6 InterventionCount
7 OutcomeCount
in the query. IDFSum sum of all idf scores, IDFStd is the standard deviation
of the idf scores, IDFMax is the maximum idf score and IDFAvg is the mean
idf score. The other three features (PopulationCount, InterventionCount, and
OutcomeCount) are the number of terms in the document and in the query
that also appear in the respective PICO annotation. PICO annotations for
documents were automatically extracted using RobotReviewer [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This automatic
process only annotates the Population, Intervention, and Outcome for studies
(the Control element is not annotated). PICO annotations for queries were
manually collected by one of the team members, who is a clinician (AD). Afterwards,
search strategies (both the original boolean query, and the new boolean query
with PICO annotations) were manually transformed into Elasticsearch queries.
The result is two Elasticsearch queries per topic | one which is
representative of the original query made by the systematic review authors, and another
annotated with PICO elements.
      </p>
      <p>
        Initial testing on a recent collection we developed [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] allowed us to select a
number of candidate learning to rank algorithms that may be e ective in the
screening prioritisation of systematic reviews.
      </p>
      <p>We then empirically evaluated the selected ve learning to rank algorithms
listed in Table 2 and found that Coordinate Ascent provided us with the best
MAP score on validation data for CLEF eHealth 2017 over the other
models. Each time we trained a model, we used the default values for that model1
and we set aside the same 30% of queries for validation. Table 2 summarises
the NDCG@10 and average precision (AP) scores for both the original Boolean
queries and the annotated PICO queries. We found that Coordinate Ascent
was the best algorithm for learning to rank these types of studies. Additionally,
we found that Random Forests and MART methods both had similar levels of
NCG@10 and AP.</p>
      <p>Additionally, we also used Elasticsearch (version 5.3) to produce a re-ranking.
We did this by issuing the Boolean and PICO query to Elasticsearch and limited
the results to only the PubMed identi ers contained in the topic le for each
query. We then let Elasticsearch rank these documents using BM25 with the
default settings2. We considered the Elasticsearch runs as our baseline.
1 The default values for each model can be found at the following URL https://
sourceforge.net/p/lemur/wiki/RankLib%20How%20to%20use/
2 k1 = 1:2, b = 0:75.</p>
      <sec id="sec-2-1">
        <title>Elasticsearch</title>
      </sec>
      <sec id="sec-2-2">
        <title>MART</title>
        <p>AdaRank
Coordinate Ascent*
LambdaMART
Random Forests*
Boolean
0.397
We found that a learning to rank approach to re-ranking studies for
systematic reviews shows promising results. Table 2 illustrates our submitted runs
compared to the baseline Elasticsearch ranking and additional runs performed
post-submission. The models trained using the search strategies annotated using
PICO achieved slightly better results than the provided Boolean search
strategies. None of our models were able to score higher than the baseline in NCG@10,
however the Coordinate Ascent model trained using PICO annotations
outperformed the baseline in AP.</p>
        <p>Additionally, we report AP, NCG@10, WSS@100 and the position of the
last relevant document (last rel) in Figure 1. These visualisations show that the
Coordinate Ascent model provides the most e ective ranking of documents (in
terms of AP and WSS@100) and scores the highest amongst the learning to
rank models for recall based measurements (NCG@10). Figure 1c shows that
learning to rank models trained on the Boolean search strategies positioned the
last relevant document in the re-ranked list the highest; and that the baseline
Elasticsearch runs do not do this as well.</p>
        <p>Figure 2 examines the e ect PICO had on re-ranking. The e ect appears
negligible on the baseline, however, we notice an increase in precision when
PICO annotations are used as training data for learning to rank models. This
suggests that the use of PICO provides a trade o between precision and recall.
Our results illustrate this clearly when precision-based measures are compared
against recall-based measures.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Future Work</title>
      <p>We plan to further increase the precision of our experiments by tuning the hyper
parameters of the best performing learning to rank models. Our learning to rank
models were trained using only a small number of features. We will investigate
the e ects of other features that are commonly used for learning to rank, and
explore more domain speci c features in addition to PICO.
pico_es
bool_es
pico_CA
pico_LambdaMART
pico_MART
bool_CA
pico_RF
bool_LambdaMART</p>
      <p>bool_RF
bool_MART
pico_AdaRank
pico_RankNet
bool_RankNet
bool_AdaRank
0.0
pico_CA
bool_es
pico_es
pico_LambdaMART</p>
      <p>pico_RF
pico_MART</p>
      <p>bool_CA
bool_LambdaMART</p>
      <p>bool_MART
pico_AdaRank</p>
      <p>bool_RF
pico_RankNet
bool_RankNet
bool_AdaRank
0.00
(c) Position in the re-ranked list of the last (d) Position in the re-ranked list of the last
study retrieved (including baselines: bool es study retrieved (including baselines: bool es
and pico es). and pico es).</p>
      <p>Fig. 1: Comparison of the e ects each algorithm had on di erent measures.
0.25
0.0
0.2
0.4</p>
      <p>0.6</p>
      <p>Recall
(a) Elasticsearch</p>
      <sec id="sec-3-1">
        <title>Boolean PICO</title>
        <p>0.25
n
o
is0.20
i
c
e
r
dP0.15
e
t
a
l
rop0.10
e
t
n
I0.05
0.8
1.0
0.0
0.4</p>
        <p>0.6</p>
        <p>Recall
(b) Coordinate Ascent
0.8</p>
        <p>1.0</p>
      </sec>
      <sec id="sec-3-2">
        <title>Boolean</title>
        <p>PICO
0.0
0.2
0.4
0.6
Recall
0.8
1.0
0.0
0.4
0.6
Recall</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spijker</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>CLEF 2017 ehealth evaluation lab overview</article-title>
          .
          <source>In: CLEF 2017 - 8th Conference and Labs of the Evaluation Forum. Lecture Notes in Computer Science (LNCS)</source>
          , Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azzopardi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spijker</surname>
          </string-name>
          , R.:
          <article-title>CLEF 2017 technologically assisted reviews in empirical medicine overview</article-title>
          . In: Working Notes of CLEF 2017 -
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          , Dublin, Ireland,
          <source>September 11-14</source>
          ,
          <year>2017</year>
          . CEUR Workshop Proceedings, CEUR-WS.org (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Scells</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deacon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azzopardi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Geva</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A test collection for evaluating retrieval of studies for inclusion in systematic reviews</article-title>
          .
          <source>In: Proceedings of SIGIR '17</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Tsafnat</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glasziou</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choong</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunn</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galgani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coiera</surname>
          </string-name>
          , E.:
          <article-title>Systematic review automation technologies</article-title>
          .
          <source>Systematic reviews 3(1)</source>
          ,
          <volume>74</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Wallace</surname>
            ,
            <given-names>B.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>I.J.</given-names>
          </string-name>
          :
          <article-title>Extracting PICO sentences from clinical trial reports using supervised distant supervision</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>17</volume>
          (
          <issue>132</issue>
          ),
          <volume>1</volume>
          {
          <fpage>25</fpage>
          (
          <year>2016</year>
          )
          <article-title>0.2 0.2 0.8 1.0 (c) Random Forests (d) LambdaMART Fig. 2: Precision-recall curves for the Elasticsearch baselines and the Coordinate Ascent models</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>