<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Partially automated literature screening for systematic reviews by modelling non-relevant articles</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Henry Petersen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Josiah Poon</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simon Poon</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Clement Loy</string-name>
          <email>clement.loy@sydney.edu.au</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariska Lee ang</string-name>
          <email>m.m.leeflang@amc.uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Academic Medical Center, University of Amsterdam</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information Technologies, University of Sydney</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Public Health, University of Sydney</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>3</fpage>
      <lpage>4</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Systematic reviews are widely considered as the highest form of medical
evidence, since they aim to be a repeatable, comprehensive, and unbiased summary
of the existing literature. Because of the high cost of missing relevant studies,
review authors go to great lengths to ensure all relevant literature is included.
It is not atypical for a single review to be conducted over the course of months
or years, with multiple authors screening thousands of articles in a multi-stage
triage process; rst on title, then on title and abstract, and nally on full text.
Figure 1a shows a typical literature screening process for systematic reviews.</p>
      <p>
        In the last decade, the information retrieval (IR) and machine learning (ML)
communities have shown increasing interest in literature searches for systematic
reviews [1{3]. Literature screening for systematic reviews can be characterised
as a classi cation task with two de ning features; a requirement for near perfect
recall on the class of relevant studies (the high cost of missing relevant evidence),
and highly imbalanced training data (review authors are often willing to screen
thousands of citations to nd less than 100 relevant articles). Previous attempts
at automating literature screening for systematic reviews have primarily focused
on two questions; how to build a suitably high recall model for the target class
in a given review under the conditions of highly imbalanced training data [
        <xref ref-type="bibr" rid="ref1 ref3">1, 3</xref>
        ],
and how best to integrate classi cation into the literature screening process [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>When screening articles, reviewers exclude studies for a number of reasons
(animal populations, incorrect disease etc.). Additionally, in any given triage
stage a study may not be relevant but still progress to the next stage as the
authors have insu cient information to exclude it (i.e. the title may not indicate
a study was performed with an animal population, however this may become
apparent upon reading the abstract). We meet the requirement for near perfect
recall on relevant studies by inverting the classi cation task and identifying
subsets of irrelevant studies with near perfect precision. We attempt to identify
such studies by training the classi er using the labels assigned at the previous
triage stage (see Figure 1c). The seamless integration with the existing manual
screening process is an advantage of our approach.</p>
      <p>The classi er is built by rst selecting terms from the title and abstracts with
the greatest information gain on labels assigned in the rst triage stage. Articles
Reviewer 1
Screens on
Ful Text</p>
      <p>Reviewer 2
Screens on</p>
      <p>Ful Text
(a)
{ 'neutropenia', but
'infection' or 'thorax'</p>
      <p>not
{ 'skin' but not 'thorax'
{ 'immunoglobulin g'
{ 'animals'
{ 'drug therapy', but not
'risk' or 'infection'
(b)</p>
      <p>Exclude
Obtain Title
and Abstract</p>
      <p>Initial
Screening on
Title Alone
Obtain Title
and Abstract
Reviewer 1
Screens on
Title and
Abstract</p>
      <p>Build and Run</p>
      <p>Classifier</p>
      <p>Resolve Both Exclude Exclude
One or more Include</p>
      <p>Reviewer 2
Screens on
Title and
Abstract
DisRcreespoalvnecies Exclude
Obtain Ful</p>
      <p>Text
(c)
are then represented as Boolean statements over these terms, and interpretable
rules are then generated using Boolean minimisation (examples of rules are given
in 1b Review authors can then re ne the classi er by selecting only those rules
most likely to describe non-relevant studies, maximising overall precision.</p>
      <p>Preliminary experiments simulating the process outlined in Figure 1c on a
previously conducted systematic review indicate that as many as 25% of articles
can be safely eliminated without the need for screening by a second reviewer.
The evaluation does assume that all false positives (studies erroneously excluded
by the generated rules) were included by the rst reviewer. Such an assumption
is reasonable; the reason for multiple reviewers is that even human experts make
mistakes. A study comparing the precision of our classi er to human reviewers is
planned. In addition, future work will focus on improving the quality of the
generated rules by trying to better capture reasons for excluding studies matching
those used by human reviewers.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aaron</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kyle H. Ambert</surname>
          </string-name>
          , and
          <string-name>
            <surname>Marian McDonagh</surname>
          </string-name>
          . Research paper:
          <article-title>Crosstopic learning for work prioritization in systematic review creation and update</article-title>
          .
          <source>JAMIA</source>
          ,
          <volume>16</volume>
          (
          <issue>5</issue>
          ):
          <volume>690</volume>
          {
          <fpage>704</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Oana</given-names>
            <surname>Frunza</surname>
          </string-name>
          , Diana Inkpen, Stan Matwin,
          <string-name>
            <given-names>William</given-names>
            <surname>Klement</surname>
          </string-name>
          , and
          <article-title>Peter OBlenis. Exploiting the systematic review protocol for classi cation of medical abstracts</article-title>
          .
          <source>Arti cial Intelligence in Medicine</source>
          ,
          <volume>51</volume>
          (
          <issue>1</issue>
          ):
          <volume>17</volume>
          {
          <fpage>25</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Stan</given-names>
            <surname>Matwin</surname>
          </string-name>
          , Alexandre Kouznetsov, Diana Inkpen, Oana Frunza, and
          <string-name>
            <surname>Peter O'Blenis</surname>
          </string-name>
          .
          <article-title>A new algorithm for reducing the workload of experts in performing systematic reviews</article-title>
          .
          <source>JAMIA</source>
          ,
          <volume>17</volume>
          (
          <issue>4</issue>
          ):
          <volume>446</volume>
          {
          <fpage>453</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>