<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ShARe/CLEF eHealth Evaluation Lab 2013, Task 3: Information Retrieval to Address Patients' Questions when Reading Clinical Reports</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J.F. Jones</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liadh Kelly</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johannes Leveling</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Allan Hanbury</string-name>
          <email>hanbury@ifs.tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Muller</string-name>
          <email>henning.mueller@hevs.ch</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sanna Salantera</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hanna Suominen</string-name>
          <email>hanna.suominen@nicta.com.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guido Zuccon</string-name>
          <email>guido.zuccon@csiro.au</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NICTA and The Australian National University</institution>
          ,
          <addr-line>ACT</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>SO</institution>
          ,
          <addr-line>Sierre</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The Australian e-Health Research Centre</institution>
          ,
          <addr-line>CSIRO, QLD</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Turku</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Vienna University of Technology</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the results of task 3 of the ShARe/CLEF eHealth Evaluation Lab 2013. This evaluation lab focuses on improving access to medical information on the web. The task objective was to investigate the e ect of using additional information such as the discharge summaries and external resources such as medical ontologies on the IR e ectiveness. The participants were allowed to submit up to seven runs, one mandatory run using no additional information or external resources, and three each using or not using discharge summaries.</p>
      </abstract>
      <kwd-group>
        <kwd>Information retrieval</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Medical information retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The goal of the ShARe/CLEF (Cross-Language Evaluation Forum) eHealth
Evaluation Lab is to evaluate systems that support laypeople in searching for
and understanding their health information [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It comprises three tasks. The
speci c use case considered is as follows: before leaving the hospital, a patient
receives a discharge summary. This describes the diagnosis and the treatment
that they received in the hospital. The rst task considered in CLEF eHealth
aims at extracting names of disorders from the discharge summaries, while the
second task requires normalisation and expansion of abbreviations and acronyms
? In alphabetical order, LG, GJFJ, LK &amp; JL led Task 3; AH, HM, SS, HS &amp; GZ were
on the Task 3 organising committee.
present in the discharge summaries. The use case then postulates that, given the
discharge summaries and the diagnosed disorders, patients often have questions
regarding their health condition. The goal of the third task is to provide
valuable and relevant documents to patients, so as to satisfy their health-related
information needs. To evaluate systems that tackle this third task, we provide
potential patient queries and a document collection containing various health
and biomedical documents for task participants to create their search system.
As is common in evaluation of information retrieval (IR), the test collection
consists of documents, queries, and corresponding relevance judgements.
      </p>
      <p>
        Searching for health advice is a common and important task performed by
individuals on the web. Nearly 70% of search engine users in the US have
conducted a web search for information about a speci c disease or health problem [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
While health IR is often considered as a domain-speci c task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], it is performed
by a large variety of users, including various healthcare workers, but also, and
increasingly commonly, by laypeople (e.g., patients and their relatives). This
variety of potential information seekers, each characterised by di erent health
knowledge, implies a broad range of information needs, and consequently a
requirement for retrieval systems able to satisfy the health information needs of
di erent categories of users.
      </p>
      <p>
        The growing importance of health IR has provided the motivation for a
number of evaluation campaigns focusing on health information. For example, the
TREC (Text REtrieval Conference) Medical Records Tracks aim at identifying
patient cohorts from medical reports to recruit for user studies [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In this task,
topics include a particular disease/condition set and a particular
treatment/intervention set; demographics or other characteristics may also be part of the
topics (e.g., age group and hospitalisation status). Moreover, the ImageCLEFmed
tracks of the CLEF Initiative (Conference and Labs of the Evaluation Forum,
formerly known as Cross-Language Evaluation Forum) have created resources
for the evaluation of image search in online resources or biomedical journal
articles [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. However, while addressing di erent information needs (e.g., nding
similar clinical cases vs. journal papers), these previous campaigns have targeted
speci c groups of users with expert health knowledge (e.g., clinicians and health
researchers). The ShARe/CLEF eHealth Task 3 resembles other ad-hoc
information retrieval tasks but with a focus on the information needs of laypeople
and the types of queries they pose to express these needs.
      </p>
      <p>The rest of this paper is organized as follows: Section 2 outlines the main
evaluation campaigns on health IR. Section 3 describes the creation of the CLEF
eHealth dataset, that is, the document collection, query generation, and
relevance assessment. Section 4 presents the result sets and their evaluation and
Section 5 the approaches used by task participants. Finally Section 6 concludes
the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Previous research has considered the information needs of individuals seeking
health advice on the web, but these studies mainly analysed query logs from
large commercial search engines [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. To the best of our knowledge, no evaluation
campaign has considered the information needs that patients may have regarding
their health conditions and provided resources for evaluating IR systems for
this task. Such lack of attention to this task arises, at least partially, due to
the complexity of assessing the information needs: laypeople that search for
health information on the web have very varied pro les, and their queries and
searching time tend to be much shorter than those considered in past health IR
benchmarks [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
      </p>
      <p>
        OHSUMED, published in 1994, was the rst collection containing medical
data used for IR evaluation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The collection contained around 350,000
abstracts from medical journals on the MEDLINE database over a period of ve
years (1987{1991) and two sets of topics: 63 topics manually generated and
around 5,000 topics based on the controlled vocabulary thesaurus of the
Medical Subject Headings7 (concept name and de nition). The collection was created
for the TREC 2000 Filtering Track but also used for other research on health
IR [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ].
      </p>
      <p>
        The TREC Medical Records Track ran in 2011 and 2012 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. It was based on
a collection of de-identi ed medical records (93,551 medical reports mapped into
17,264 visits) and queries (35 queries in 2011 and 50 in 2012) that resembled
eligibility criteria of clinical studies. Records were grouped into visits,
corresponding to a patient admission in the hospital; visits ranged in length from a
few hours to in excess of a year. The goal of the track was to nd patient cohorts
that are relevant to the criteria for recruitment as populations in comparative
e ectiveness studies.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Task 3 Description</title>
      <p>
        The data set provided to participants comprises a document collection of around
one million documents (web pages from medical web sites), 50 topics, which were
developed by medical experts, and the corresponding relevance information [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
In addition to TREC-style title and description elds, the topics contain an
additional eld discharge-summary, which contains the discharge report which
the patient's query stemmed from.
      </p>
      <p>The data was provided to participants after signing an agreement, through
the PhysioNet website. As test data, ve training topics together with
corresponding relevance assessment were released.</p>
      <p>We describe in this section each part of the task dataset.
7 http://www.ncbi.nlm.nih.gov/mesh/</p>
      <sec id="sec-3-1">
        <title>Document Collection</title>
        <p>
          A large web crawl of health resources is used as the corpus for this task. The
crawl contains about one million documents, which have been made available
to CLEF eHealth through the Khresmoi project [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. This collection consists of
web pages covering a broad range of health topics, targeted at both the general
public and healthcare professionals. These domains consist predominantly of
health and medicine websites that have been certi ed by the Health on the Net
(HON) Foundation8 as adhering to the HONcode principles9 (approximately
60{70% of the collection), as well as other commonly used health and medicine
websites such as Drugbank10, Diagnosia11 and Trip Answers12. The crawled
documents are provided in the dataset in their raw HTML (Hyper Text Markup
Language) format along with their uniform resource locators (URL). The dataset
is made available for download on the web to registered participants on a secure
password-protected server.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Discharge Summaries</title>
        <p>Novel methods to generate contextualised statements of patient information
needs were used. These are based on realistic short query statements created
in the context of patient discharge summaries. The discharge summaries can be
considered as a description of the context in which the patient has been
diagnosed with a given disorder and has written a query. The discharge summaries
originate from the de-identi ed MIMIC-II database13 (Multiparameter
Intelligent Monitoring in Intensive Care, Version 2.5).</p>
        <p>Discharge summaries are semi-structured reports with the following
appearance:
Admission Date : [ 2014 03 28 ]
D i s c h a r g e Date : [ 2014 04 08 ]
Date o f B i r t h : [ 1930 09 21 ]
Sex : F
S e r v i c e : CARDIOTHORACIC
A l l e r g i e s :
P a t i e n t r e c o r d e d as having No Known A l l e r g i e s to Drugs
Attending : [ Attending I n f o 5 6 5 ]
C h i e f Complaint : Chest pain
Major S u r g i c a l or I n v a s i v e Procedure :
Coronary a r t e r y bypass g r a f t 4 .</p>
        <p>H i s t o r y o f P r e s e n t I l l n e s s :
83 year o l d woman , p a t i e n t o f Dr . [ F i r s t Name4
( NamePattern1 ) ] [ Last Name ( NamePattern1 ) 5 0 0 5 ] ,
Dr . [ F i r s t Name ( S T i t l e ) 5 8 0 4 ] [ Name ( S T i t l e )
2 2 7 5 ] , with i n c r e a s e d SOB with a c t i v i t y , l e f t s h o u l d e r
b l a d e / back pain at r e s t , + MIBI , r e f e r r e d f o r c a r d i a c
cath . This p l e a s a n t 83 year o l d p a t i e n t n o t e s becoming
8 http://www.healthonnet.org
9 http://www.hon.ch/HONcode/Patients-Conduct.html
10 http://www.drugbank.ca/
11 http://www.diagnosia.com/
12 http://www.tripanswers.org/
13 http://mimic.physionet.org
SOB when walking up h i l l s or i n c l i n e s about one year
ago . This SOB has p r o g r e s s i v e l y worsened and she i s now
SOB when walking [ 01 19 ] c i t y b l o c k ( f l a t s u r f a c e ) .
[ . . . ]
Past Medical H i s t o r y :
a r t h r i t i s ; c a r p a l t u n n e l ; s h i n g l e s r i g h t arm 2 0 0 0 ;
needs r i g h t knee r e p l a c e m e n t ; l e f t knee r e p l a c e m e n t
i n [ 2 0 1 0 ] ; thyroidectomy 1 9 7 8 ; c h o l e c y s t e c t o m y
[ 1 9 8 1 ] ; hysterectomy 2 0 0 1 ; h/o LGIB 2000 2001
a f t e r t a k i n g baby ASA; 81 QOD
[ . . . ]
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Topics</title>
        <p>
          The queries used in the task aim to model those used by laypeople (i.e., patients,
their relatives or other representatives) to nd out more about their disorders,
once they have examined a discharge summary. Disorders have been identi ed
within discharge summaries and linked to the matching UMLS (Uni ed Medical
Language System) concept in the CLEF eHealth Task 1 [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>Previous evaluation tasks in health IR have used MeSH entries (the MeSH
ontology is contained in the UMLS meta-ontology) as queries (see Section 2).
However, the queries considered by the task presented here are intended to be
representative of real patients' information needs and statements. Thus the
possibility of issuing concept-queries is discarded. Layperson queries tend to be
short, with an average length less than two words. However, di erent patients
can have di erent information needs associated with the same query statement.
For example, a patient that receives a cancer diagnosis for the rst time would
have a di erent information need than a patient at a terminal cancer stage. This
type of contextual information related to the patient history is contained in the
discharge summary. Thus, the discharge summaries can be used for contextually
focused generation of queries. The information in a discharge summary can then
be used to determine the relevance of retrieved information to the speci c user.</p>
        <p>A query is generated for a given disorder and a discharge summary. To better
structure the query generation process, patients' information needs have been
grouped into three main scenarios:
1. the patient has a short-term disease, or has been hospitalised after an
accident (little to no knowledge of the disorder, short-term treatment),
2. the patient has a chronic disease or a long-term disease that has just been
diagnosed (little to no knowledge of the disorder, long-term treatment), and
3. the patient has a chronic or long-term disease, and this is the n-th diagnosis
(potentially good knowledge of the disorder, long-term treatment).</p>
        <p>Queries to be used in this task have been created by experts (each expert
was a registered nurse and clinical documentation researcher) involved in the
CLEF eHealth consortium. This solution has been chosen in place of recruiting
patients because of the issues involved with recruitment and privacy. We believe
that, being on a daily basis in contact with patients receiving treatments and
discharge summaries, nurses are familiar with patients' information needs and
patient pro les.</p>
        <p>65 disorders have been randomly selected from the set of 1,006 disorders
identi ed in the CLEF eHealth Task 1. For each disorder, a discharge summary
containing the disorder itself has been randomly selected. Using the pairs of
disorder and associated discharge summary, the experts have developed a set
of patient queries (and criteria for judging the relevance of documents to the
queries, for use in the relevance assessment task described in the next section).
Queries are provided in a standard TREC format, consisting of a topic title (text
of the query), description (longer description of what the query means), and a
narrative (expected content of the relevant documents).</p>
        <p>The following example outlines a query:
&lt;query&gt;
&lt; t i t l e &gt; thrombocytopenia treatment c o r t i c o s t e r o i d s</p>
        <p>l e n g t h &lt;/ t i t l e &gt;
&lt;desc&gt; How l o n g s h o u l d be the c o r t i c o s t e r o i d s treatment</p>
        <p>to c u r e thrombocytopenia ? &lt;/desc&gt;
&lt;narr&gt; Documents s h o u l d c o n t a i n i n f o r m a t i o n about
t r e a t m e n t s o f thrombocytopenia , and e s p e c i a l l y
c o r t i c o s t e r o i d s . I t s h o u l d d e s c r i b e the treatment ,
i t s d u r a t i o n and how the d i s e a s e i s cured u s i n g i t .
&lt;s c e n a r i o &gt; The p a t i e n t has a s h o r t term d i s e a s e , or
has been h o s p i t a l i s e d a f t e r an a c c i d e n t ( l i t t l e to
no knowledge o f the d i s o r d e r , s h o r t term treatment )
&lt;/s c e n a r i o &gt;
&lt;p r o f i l e &gt; P r o f e s s i o n a l f e m a l e &lt;/ p r o f i l e &gt;
&lt;/narr&gt;
&lt;/query&gt;</p>
        <p>
          With this approach, ve training and fty test queries have been generated
for use in the task. 65 disorders have been selected (i.e. more than the targeted
number of queries) because some disorders/queries may not be answerable using
web pages from the document collection. During the query generation process,
the experts manually removed disorders from the list of 65 that do not allow
for realistic query generation. A real log containing queries issued by the general
public on the HON website has also been used to exclude candidate queries which
are unrealistic of the type of query that a patient would typically enter. For each
query, an IR system that implements a standard BM25 weighting scheme [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]
has been used to retrieve a shallow pool of documents. This has been used to
assess whether a standard retrieval system could match at least one relevant
document to a candidate query. Queries with no relevant documents retrieved
in the shallow pool have been removed.
3.4
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Relevance Assessment</title>
        <p>
          Relevance assessment has been performed by domain experts and IR experts. We
used the Relevation system14 for collecting relevance assessments of documents
contained in the assessment pools. Relevation is a system for performing
relevance judgements for Information Retrieval system evaluation [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Documents
14 http://ielab.github.io/relevation/
and queries can be uploaded to the system via the web interface; relevance
judges can browse the uploaded documents and queries and provide their
relevance assessments. The system is open source and based on Python's Django
web framework. Relevation used a simple Model-View-Controller model that is
designed for easy customisation and extension.
        </p>
        <p>As we received many run submissions, we had to limit the pool depth and
distributed the relevance assessment workload between medical experts and IR
experts. The relevance assessment for these test queries has been conducted
by six Finnish nursing professionals and ve Australian nursing professionals or
students in health sciences (domain experts), and three Irish, one Australian, and
one Swedish senior researcher in clinical NLP (Natural Language Processing) and
ML (Machine Learning), all technological experts. Each document was assessed
by one person.</p>
        <p>We pooled the top ten documents obtained from the participants' baseline
runs, and both their top-priority run using discharge summaries and their
toppriority run not using discharge summaries. This resulted in a pool of 6,391
documents. The relevance assessment was based on a four point scale, which
is mapped to a binary scale. The graded relevance assessment yielded 0: 4,316,
1: 197, 2: 1,439, 3: 439 documents. The binary relevance assessments yielded 0:
4,513 non-relevant and 1: 1,878 relevant documents.</p>
        <p>Thus, there have been 37.56 relevant documents per topic on average and
127.82 documents assessed per topic. Table 1 shows the relevance assessment
coverage. (*) indicates runs for which results up to rank 10 were completely
assessed.</p>
        <p>
          Relevance assessments for the ve training queries were formed based on
pooled sets generated using the Vector Space Model [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] and Okapi BM25 [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
Assessments for these ve training queries were conducted by two Finnish nurses.
Each document was assessed by one person.
Run
1 (*)
2 (*)
3
4
5 (*)
6
7
9
5
5
5
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>For this task, the participants were allowed to submit up to seven runs, one
mandatory run using no additional information or external resources (run 1),
three using the discharge summary and any other external resource (runs 2-4),
and three using external resources but not using the discharge summaries (run
5-7). Among each set of additional runs, one had to use only the title and the
description elds of the query. Participants were also asked to rank their runs
2-4 and 5-7 according to their importance.
4.1</p>
      <sec id="sec-4-1">
        <title>Participants</title>
        <p>13 groups registered for task 3 and 9 groups submitted runs. The groups are
from 5 countries and 4 continents: Asia (2), Australia (2), USA (4), Europe (1).
The groups are listed in Table 2. Although CLEF has mainly been the focus of
European groups, only one group from Europe submitted runs to this lab.</p>
        <p>Teams submitted in total 48 runs, including: 9 baseline runs, 15 runs using
discharge summaries (from 5 teams), and 24 runs not using them.
We examined all documents in runs 1, 2 and 5 up to rank 10 for relevance. The
two major evaluation metrics are therefore metrics at a cut-o of up to 10
documents, i.e. P@5, P@10, NDCG@5, and NDCG@10. In addition, we considered
MAP as an evaluation metric, but we are aware that MAP is unreliable because
only the top ten documents have been assessed. Nevertheless, we wanted to
report a measure covering the full set of up to 1000 retrieved documents. We also
report the number of relevant and retrieved documents in the top 1000 results
as a more recall-oriented measure. Table 1 provides details on the coverage of
assessments for each run up to rank 10.</p>
        <p>Performance metrics are computed with the standard trec eval tool15 using
the following options:
{ trec eval -c -M1000
{ trec eval -c -M1000 -m ndcg cut</p>
        <p>We are aware that the performance metrics for other runs might be unreliable
compared to that of runs 1, 2, and 5. However, this situation is common for IR
lab evaluations, where additional experiments on an existing data set typically
do not include re-assessment of documents previously not retrieved or relevance
assessment of additional documents.
4.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Baseline System</title>
        <p>
          For comparison, we created our own baseline experiment, which is based on the
BM25 retrieval model. This experiment does not incorporate any domain-speci c
adaptations. We used the JSoup library16 to extract the textual content from
the web pages and applied standard normalization (e.g. replacing named HTML
characters) and spelling correction based on a xed list of frequent spelling errors.
This approach has been employed before for experiments on medical records
retrieval [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and was shown to represent a very strong baseline.
        </p>
        <p>
          This baseline system uses the BM25 retrieval model and standard blind
relevance feedback for BM25 [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The system uses a standard stop-word list
containing the Okapi stop-words (222 stop-words) for stop-word removal.
        </p>
        <p>
          The baseline system performs only two types of document preprocessing:
character normalization (i.e. mapping characters with diacritical marks to the
equivalent characters without) and word normalization (e.g. correcting frequent
spelling errors). Spelling correction is based on a list of 9533 spelling errors from
medical documents [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], which was added to a list of 4192 frequent spelling errors
compiled from Wikipedia. During indexing, misspelled words are replaced with
their corrections from this list.
        </p>
        <p>We generated two baseline experiments: one with standard retrieval using the
BM25 model (BM25 baseline) and a second experiment using blind relevance
feedback (BM25 FB baseline). For blind relevance feedback, the top T terms
from the top ranked D documents are added to the query to retrieve the nal
result set of documents. For our baseline experiments, we used T = 10 and
D = 10.
4.4</p>
      </sec>
      <sec id="sec-4-3">
        <title>Evaluation Results</title>
        <p>The o cial results for all submitted runs and for our baseline experiments are
shown in Table 3.</p>
        <p>Comparing the participants' results wrt. P@10 and NDCG@10, we found
that for 4 teams (ohsu, THCIB, teamAEHRC, uog) run 5 (using discharge
summaries) achieves the best P@10, while for 2 teams (QUT-TOPSIG, TeamMayo),
15 http://trec.nist.gov/trec eval/
16 http://jsoup.org/
run 2 (not using summaries) achieves the best P@10. For 3 teams (TeamKC,
MEDINFO, UTHealth), the baseline runs (run 1) performs best.</p>
        <p>Only 4 runs (from the same group) outperform our baseline experiment using
standard blind relevance feedback wrt. P@10. Runs from two groups outperform
the baseline experiment wrt. NDCG@10. This illustrates that our baseline
system based on BM25 is a very strong baseline.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Approaches Used</title>
      <p>We describe in this section the approaches used by each team, and summarize
ndings from their analysis.</p>
      <p>
        Team THCIB [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] implemented ad hoc retrieval using Lucene for indexing and
retrieval and HITS for ranking. They submitted 7 runs. They investigated query
annotations with UMLS, and the use of various elds of the queries. They also
used query expansion based on the concepts and/or the discharge summaries,
as well as concept-based re-ranking. Their best run scores 0.42 P@10 (run 5),
and uses their baseline system, with query expansion without the discharge
summaries on the title and desc topic elds.
      </p>
      <p>
        Team TOPSIG [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] used the TopSig open source tool, which implements
signaturebased approaches for information retrieval. Six runs were submitted, using
various parts of the topic les as queries, and using the text of the discharge reports
for query re nement. Submissions used TopSig normally for query-only runs and
with a new 're ne' mode for discharge summary runs. This means the query is
used for searching the collection and the results are then re-ranked using the
discharge summaries, similarly to how blind feedback works. This produced much
better results than simply using the discharge summaries in the original query
(P@10 of 0.36), which tend to be too noisy to produce e ective results.
Team AEHRC [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] used a Dirichlet-smoothed language modelling as provided
by the Terrier system for retrieval. They experimented with setting prior
probabilities depending on document readability and authorativeness information
derived from a static list of web sites. They also explored spelling correction
(using Google) and acronym expansion. They submitted 4 runs, and did not use
the discharge summaries. Their best result (P@10 of 0.48) is obtained with the
baseline and the acronym expansion and spelling correction in the query.
Team KC [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] approach is based on statistical topic models. Their baseline used
Language models for retrieval, with Indri index engine. They submitted 6 runs,
exploring query expansion through the extraction and selection of noun phrases
from documents and discharge summaries, using probabilistic topic modelling
and language models. They scored their best result (P@10 of 0.40) with the
baseline.
Team Mayo [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] used a Query-Likelihood Model as their baseline, and further
runs using a Markov Random Field Model. Query expansion was carried out
using a Mixture of Relevance Models and using the MeSH ontology. They also
investigated concept-based re-ranking using medical concepts identi ed in the
discharge summaries, and weighting concepts using attributes such as negation
or semantic groups. Their run # 2 got the best results (P@10 of 0.52), using
a Markov random eld to model the dependency between query terms,
expansion of query terms through 4 medical/genomics collections and a concept-space
search ranking (based on query and discharge summary).
      </p>
      <p>
        Team SNUMEDINFO [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] used unigram language model with Dirichlet prior
smoothing on indri search engine as a baseline. Their runs without discharge
summaries are using passage based language model, combining max-scoring
passage-based relevance score with unigram language model score, with di
erent weighting parameters. The following runs used lexical query expansion with
UMLS concepts and preferred terms. They also explored the use of di erent
elds of the query. Their best results are obtained with the baseline (P@10 of
0.48).
      </p>
      <p>
        Team OHSU [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] submission included runs from two di erent retrieval systems.
One set used the Lucene Vector Space Retrieval model and extensions which
made use of MetaMap for query expansion. The second used novel statistical
language modelling techniques. Valuable insights for future system improvement
were gained, such as insights into selective indexing due to the large amount of
noise in web page content, and the need for sophisticated query parsing and
expansion techniques. They obtained best results with run #5, using a language
model to attempt to nd the documents whose word distributions would have
been most likely to generate the query, together with an adaptive perplexity
threshold.
      </p>
      <p>
        Team UThealth [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] submission contained runs using the vector space model
and the semantic vector space model. Best performance was obtained using the
vector space model. They also explored the use of di erent elds of the topics
as queries. Their best result was obtained with the baseline (P@10 of 0.37).
Team UOG [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] used the Terrier information retrieval framework. Divergence
from Randomness and pseudo relevance feedback were employed on the retrieval
of medical web pages. They also investigated query expansion using corpus of
MEDLINE abstracts or Wikipedia collection. Their best run (P@10 of 0.44) used
the baseline system and pseudo relevance feedback.
      </p>
      <p>The best result overall was obtained by team Mayo, using a retrieval model
and external resources for query expansion, as well as re-ranking based on
concepts from the query and the discharge summary. However, this team, together
with Topsig, are the only teams having improved their baseline with the use of
the discharge summaries, and in both cases for re-ranking purposes.</p>
      <p>Concept re-ranking has also been used by team THCIB and given their best
results. For ve teams, the best scores are obtained with their baseline
(coupled with pre-processing such as acronym expansion and spelling correction for
team AEHRC). Combination of methods and external resources such as those
employed by team Mayo seem to be e cient, but the overall results here show
that research still needs to be conducted to make the best out of the external
resources.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this rst year of the ShARe/CLEF eHealth2013 evaluation lab Task 3, there
was strong take-up in the community with 9 groups submitting runs to the task.
The challenge of developing retrieval techniques for layperson medical queries
proved di cult. Overall, BM25 is a very strong baseline, which proved hard for
teams to beat. BM25 with relevance feedback, as opposed to standard BM25,
proved to be the strongest baseline. Here P@10 of 0.4860 was obtained, and
NDCG@10 of 0.4328. One team, Mayo, had four runs which performed better
than BM25 for P@10. P@10 achieved for these runs ranged from 0.5180 { 0.4880.
Of these runs, the best performing used information in the discharge summaries.
The only other team that improved over the BM25 baseline was SNUBME for
their baseline run, where NDCG@10 of 0.4377 was achieved.</p>
      <p>Despite the results obtained being mostly lower than the baseline in this rst
year of the task, many valuable insights into retrieval technique development
for the domain were gained. This forms a good basis for future exploration in
the domain in general, and speci cally for further technique development for the
2014 CLEF eHealth Task 3 lab.</p>
      <p>Given the success of this year 1 of the task, we anticipate even more interest in
next year's campaign. In the second year of the task, we will seek to remove noise
from the document collection. This year, as task participants were informed,
there were some non-English web pages in the collection and duplicate web
pages. That is there were occurrences in the document collection of the same
web page, with the same URL but di erent unique identi er. On top of this,
we also identi ed several occurrences of the same web pages with di erent URL
pre xes. This occurs for example, when a dropdown menu is expanded. We will
also explore the possibility of removing or highlighting such duplication in the
collection. In next year's campaign we also intend exploring new query generation
techniques and means to improve the relevance assessment work ow. In query
generation for example, we will look at using the main disease in the discharge
summary the query (information need) stems from instead of using a randomly
selected disease from within the discharge summary. The goal here being to
increase the relevance of using discharge summaries.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgement</title>
      <p>Task 3 of the ShARe/CLEFeHealth2013 evaluation lab has been supported in
part by the Khresmoi project, funded by the European Union Seventh
Framework Programme (FP7/2007-2013) under grant agreement no 257528, and NICTA,
funded by the Australian Government as represented by the Department of
Broadband, Communications and the Digital Economy and the Australian
Research Council through the ICT Centre of Excellence program. The relevance
judgements were in part funded by the ESF project ELIAS. We acknowledge
the time given to perform the relevance assessment task. We want to thank
the following individuals: Maricel Angel (NICTA, Australia), Rmi Bois (Dublin
City University), Riitta Danielsson-Ojala (University of Turku, Finland), Lotta
Kauhanen (University of Turky, Finland), Hugo Mougard (Dublin City
University), Laura-Maria Murtola (University of Turku, Finland), Heidi Parisod
(University of Turku, Finland), Josh Robertson (Dublin City University,
Ireland), Eriikka Siirala (University of Turku, Finland), Timothy Sladden
(University of Queensland, Australia), Thomas Souchen (AEHRC, Australia), Sumithra
Velupillai (Stockholm University, Sweden).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Salantera,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Velupillai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.W.</given-names>
            ,
            <surname>Savova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Elhadad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Mowery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Leveling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Martinez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>ShARe/CLEF eHealth Evaluation Lab 2013: Three shared tasks on natural language processing and machine learning to make clinical reports easier to understand for patients</article-title>
          .
          <source>In: CLEF 2013. Lecture Notes in Computer Science (LNCS)</source>
          , Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Health topics:
          <volume>80</volume>
          %
          <article-title>of internet users look for health information online</article-title>
          .
          <source>Technical report</source>
          , Pew Research Center (
          <year>February 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Medical information retrieval: an instance of domain-speci c search</article-title>
          .
          <source>In: Proceedings of SIGIR</source>
          <year>2012</year>
          .
          <article-title>(</article-title>
          <year>2012</year>
          )
          <volume>1191</volume>
          {
          <fpage>1192</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tong</surname>
            ,
            <given-names>R.M.:</given-names>
          </string-name>
          <article-title>Overview of the TREC 2011 medical records track</article-title>
          .
          <source>In: Proceedings of TREC, NIST</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kalpathy-Cramer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Bedrick</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eggel</surname>
            , I., de Herrera,
            <given-names>A.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The CLEF 2011 medical image retrieval and classi cation tasks</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          (
          <article-title>Cross Language Evaluation Forum)</article-title>
          . (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Muller, H.,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deselaers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caputo</surname>
          </string-name>
          , B., eds.: ImageCLEF |
          <article-title>Experimental Evaluation in Visual Information Retrieval</article-title>
          . Volume
          <volume>32</volume>
          of The Information Retrieval Series. Springer (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horvitz</surname>
          </string-name>
          , E.:
          <article-title>Cyberchondria: Studies of the escalation of medical concerns in web search</article-title>
          .
          <source>Technical report, Microsoft Research</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Boyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gschwandtner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kritz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pletneva</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vargas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Use case de nition including concrete data requirements (D8.2). public deliverable, Deliverable of the Khresmoi EU project (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Suominen</surname>
          </string-name>
          , H., ed.:
          <source>The Proceedings of the CLEFeHealth2012 { the CLEF 2012 Workshop on Cross-Language Evaluation of Methods</source>
          , Applications, and
          <article-title>Resources for eHealth Document Analysis</article-title>
          .
          <source>NICTA</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hersh</surname>
            ,
            <given-names>W.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leone</surname>
            ,
            <given-names>T.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hickam</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          :
          <article-title>OHSUMED: An interactive retrieval evaluation and new large test collection for research</article-title>
          .
          <source>In: Proceedings of SIGIR '94</source>
          . (
          <year>1994</year>
          )
          <volume>192</volume>
          {
          <fpage>201</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Claveau</surname>
          </string-name>
          , V.:
          <article-title>Unsupervised and semi-supervised morphological analysis for information retrieval in the biomedical domain</article-title>
          .
          <source>In: Proceedings of COLING</source>
          . (
          <year>2012</year>
          )
          <volume>629</volume>
          {
          <fpage>645</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bruza</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sitbon</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawley</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An evaluation of corpus-driven measures of medical concept similarity for information retrieval</article-title>
          .
          <source>In: Proceedings of CIKM</source>
          <year>2012</year>
          .
          <article-title>(</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Leveling</surname>
          </string-name>
          , J.:
          <article-title>Creation of a new evaluation benchmark for information retrieval targeting patient information needs</article-title>
          . In Song, R.,
          <string-name>
            <surname>Webber</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kishida</surname>
          </string-name>
          , K., eds.
          <source>: Proceedings of the 5th International Workshop on Evaluating Information Access (EVIA)</source>
          ,
          <source>a Satellite Workshop of the NTCIR-10 Conference</source>
          , Tokyo/Fukuoka, Japan, National Institute of Informatics/Kijima Printing (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H.:
          <article-title>Khresmoi { multimodal multilingual medical information search</article-title>
          . In:
          <article-title>MIE village of the future</article-title>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>S.</given-names>
            <surname>Pradhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Elhadad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.S.</given-names>
            <surname>D.M.L.C.H.W.C.G.S.</surname>
          </string-name>
          <article-title>: Task 1: Share/clef ehealth</article-title>
          .
          <source>In: Proceedings of the CLEF conference</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Simple, proven approaches to text retrieval</article-title>
          .
          <source>Technical Report 356</source>
          , University of Cambridge (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>Relevation! an open source system for information retrieval relevance assessment</article-title>
          .
          <source>arXiv preprint</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.S.:</given-names>
          </string-name>
          <article-title>A vector space model for automatic indexing</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>18</volume>
          (
          <issue>11</issue>
          ) (
          <year>1975</year>
          )
          <volume>613</volume>
          {
          <fpage>620</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.F.</given-names>
          </string-name>
          : DCU@
          <article-title>TRECMed 2012: Using ad-hoc baselines for domain-speci c retrieval</article-title>
          .
          <source>In: Proceedings of TREC</source>
          <year>2012</year>
          ,
          <string-name>
            <surname>NIST</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Zhong</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Na</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Concept-based medical document retrieval: THCIB at CLEF eHealth lab 2013 task 3</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Chappell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Geva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Working notes for topsig at share/clef ehealth 2013</article-title>
          .
          <article-title>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</article-title>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Retrieval of health advice on the web: AEHRC at ShARe/CLEF eHealth evaluation lab task 3</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Barajas</surname>
            ,
            <given-names>K.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akella</surname>
          </string-name>
          , R.:
          <article-title>Incorporating statistical topic models in the retrieval of health care documents</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>James</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carterette</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , H.:
          <article-title>Using discharge summaries to improve information retrieval in clinical domain</article-title>
          .
          <source>In: Proceedings of the ShARe/- CLEF eHealth Evaluation Lab</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>SNUMedinfo at CLEFeHealth2013 task 3</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Bedrick</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheikhshabbafghi</surname>
          </string-name>
          , G.:
          <article-title>Lucene, metamap, and language modeling: OHSU at CLEF eHealth 2013</article-title>
          .
          <article-title>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</article-title>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Y.,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Evaluation of vector space models for medical disorders information retrieval</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Limsopatham</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macdonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ounis</surname>
          </string-name>
          , I.: University of glasgow at CLEF 2013:
          <article-title>Experiments in eHealth task 3 with terrier</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>