<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ShARe/CLEF eHealth Evaluation Lab 2014, Task 3: User-centred health information retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liadh Kelly</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wei Li</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joao Palotti</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavel Pecina</string-name>
          <email>pecina@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guido Zuccon</string-name>
          <email>g.zuccon@qut.edu.au</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Allan Hanbury</string-name>
          <email>hanbury@ifs.tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J.F. Jones</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Muller</string-name>
          <email>henning.mueller@hevs.ch</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University in Prague</institution>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Queensland University of Technology</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>SO</institution>
          ,
          <addr-line>Sierre</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Vienna University of Technology</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>43</fpage>
      <lpage>61</lpage>
      <abstract>
        <p>This paper presents the results of task 3 of the ShARe/CLEF eHealth Evaluation Lab 2014. This evaluation lab focuses on improving access to medical information on the web. The task objective was to investigate the e ect of using additional information such as a related discharge summary and external resources such as medical ontologies on the e ectiveness of information retrieval systems, in a monolingual (Task 3a) and in a multilingual (Task 3b) context. The participants were allowed to submit up to seven runs for each language (English, Czech, French, German), one mandatory run using no additional information or external resources, and three each using or not using discharge summaries.</p>
      </abstract>
      <kwd-group>
        <kwd>Information retrieval</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Medical information retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The goal of the ShARe/CLEF (Cross-Language Evaluation Forum) eHealth
Evaluation Lab is to evaluate systems that support laypeople in searching for
and understanding their health information [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It comprises three tasks. The
speci c use case considered is as follows: upon leaving the hospital, a patient
receives a discharge summary. This describes the diagnosis and the treatment
that they received in the hospital. Task 1 focuses on visual-interactive search
and exploration of eHealth data. Its aim is to help patients (or their next-of-kin)
in readability issues related to their hospital discharge documents and related
information search on the Internet. Task 2 explores information extraction from
clinical reports. Finally, this year's Task 3 further extends the 2013
information retrieval task, by cleaning the 2013 document collection and introducing a
new query generation method and multilingual topics. This year then, Task 3
is split into Task 3a and Task 3b. Task 3a, similar to last year's Task 3, is a
monolingual English retrieval task. Task 3b, adds a cross-lingual retrieval
challenge to the lab, where participants must rst translate parallel German, French
and Czech queries into English before performing retrieval. The overall goal of
Task 3 is to provide valuable and relevant documents to patients, so as to
satisfy their health-related information needs. To evaluate systems that tackle this
third task, we provide potential patient queries and a document collection
containing various health and biomedical documents for task participants to create
their search system. As is common in evaluation of information retrieval (IR),
the test collection consists of documents, topics6, and corresponding relevance
judgements.
      </p>
      <p>
        Searching for health advice is a common and important task performed by
individuals on the web. Nearly 70% of search engine users in the US have
conducted a web search for information about a speci c disease or health problem [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
While health IR is often considered as a domain-speci c task, it is performed
by a large variety of users, including various healthcare workers, but also, and
increasingly commonly, by laypeople (e.g., patients and their relatives). This
variety of potential information seekers, each characterized by di erent health
knowledge, implies a broad range of information needs, and consequently a
requirement for retrieval systems able to satisfy the health information needs of
di erent categories of users.
      </p>
      <p>
        The growing importance of health IR has provided the motivation for a
number of evaluation campaigns focusing on health information. For example, the
TREC (Text REtrieval Conference) Medical Records Track aims at identifying
patient cohorts from medical reports to recruit for clinical trials [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this task,
topics include a particular disease/condition set and a particular
treatment/intervention set; demographics or other characteristics may also be part of the
topics (e.g., age group and hospitalization status). Moreover, the ImageCLEFmed
tracks of the CLEF Initiative (Conference and Labs of the Evaluation Forum,
formerly known as Cross-Language Evaluation Forum) have created resources
for the evaluation of image search in online resources or biomedical journal
articles [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. However, while addressing di erent information needs (e.g., nding
similar clinical cases vs. journal papers), these previous campaigns have targeted
speci c groups of users with expert health knowledge (e.g., clinicians and health
researchers). The ShARe/CLEF eHealth Task 3 resembles other ad-hoc
information retrieval tasks but with a focus on the information needs of laypeople
and the types of queries they pose to express these needs. Results from the 2013
task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] showed that this was a challenging task, with space for improvement
and innovative techniques. Results from this year show considerable
improvement over last year's results, both for the team submissions and the baseline,
albeit on a new query set.
6 A topic is considered to be an enriched version of a query, but both terms are used
to refer to a topic in the paper.
      </p>
      <p>The rest of this paper is organized as follows: Section 2 outlines the main IR
evaluation campaigns on health topics. Section 3 describes the creation of the
CLEF eHealth dataset, that is, the document collection, query generation, and
relevance assessment. Section 4 presents the result sets and their evaluation and
Section 5 the approaches used by task participants. Finally Section 6 concludes
the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Previous research has considered the information needs of individuals seeking
health advice on the web, but these studies mainly analyzed query logs from
large commercial search engines [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. To the best of our knowledge, no evaluation
campaign has considered the information needs that patients may have regarding
their health conditions and provided resources for evaluating IR systems for
this task. Such lack of attention to this task arises, at least partially, due to
the complexity of assessing the information needs: laypeople that search for
health information on the web have very varied pro les, and their queries and
searching time tend to be much shorter than those considered in past health IR
benchmarks [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
      </p>
      <p>
        OHSUMED, published in 1994, was the rst collection containing medical
data used for IR evaluation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The collection contained around 350,000
abstracts from medical journals on the MEDLINE database over a period of ve
years (1987{1991) and two sets of topics: 63 topics manually generated and
around 5,000 topics based on the controlled vocabulary thesaurus of the
Medical Subject Headings7 (concept name and de nition). The collection was created
for the TREC 2000 Filtering Track but also used for other research on health
IR [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ].
      </p>
      <p>
        The TREC Medical Records Track ran in 2011 and 2012 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. It was based
on a collection of de-identi ed medical records (93,551 medical reports mapped
into 17,264 visits) and queries (35 queries in 2011 and 50 in 2012) that
resembled eligibility criteria of clinical studies. Records were grouped into visits,
corresponding to a patient admission in the hospital; visits ranged in length
from a few hours to in excess of a year. The goal of the track was to nd
patient cohorts that are relevant to the criteria for recruitment as populations in
comparative e ectiveness studies. In 2014, TREC organized a new medical
evaluation challenge, called TREC Clinical Decision Support Track8. The focus of
the track is the retrieval of biomedical articles relevant for answering generic
clinical questions about medical records. Participants are provided with short
case reports, as idealized representations of actual medical records. They have
to retrieve biomedical articles that answer questions related to several types of
clinical information needs based on the report.
      </p>
      <p>
        In 2013, CLEF hosted a workshop and challenge focusing on multilingual
biomedical named entities recognition, CLEF-ER[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Their challenge was based
7 http://www.ncbi.nlm.nih.gov/mesh/
8 http://www.trec-cds.org/
on a parallel corpus in English, French, German, Spanish, and Dutch, composed
of patent texts, titles of Medline abstracts and EMEA documents. The goal of
the task was to identify concepts by their CUIs (Concept Unique Identi ers)
in the documents, using biomedical terminological resources, and an annotated
English corpus.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Task 3 Description</title>
      <p>The data set provided to participants comprises a document collection of around
one million documents (web pages from medical web sites), 50 parallel topics
(in English (EN), Czech (CS), French (FR), and German (DE)), which were
developed by medical experts in English and translated into CS, FR and DE,
and the corresponding relevance information. In addition to TREC-style title
and description elds, the topics contain an additional eld discharge-summary,
which contains the discharge report which the patient's query stemmed from.</p>
      <p>The data was provided to participants after signing an agreement, through
the PhysioNet website. As test data, ve parallel training topics (in EN, CS,
FR, and DE) together with corresponding relevance assessment were released.</p>
      <p>In this section we describe each part of the task dataset.
3.1</p>
      <sec id="sec-3-1">
        <title>Document Collection</title>
        <p>A large web crawl of health resources is used as the corpus for this task. This is an
updated version of the web crawl released for CLEFeHealth Task 3 2013. In this
updated version further e orts have been made to clean the document collection,
by removing duplicate documents with the same URL and xing detected errors
in HTML.</p>
        <p>
          The crawl contains about one million documents, which have been made
available to CLEF eHealth through the Khresmoi project [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. This collection
consists of web pages covering a broad range of health topics, targeted at both
the general public and healthcare professionals. These domains consist
predominantly of health and medicine websites that have been certi ed by the Health
on the Net (HON) Foundation9 as adhering to the HONcode principles10
(approximately 60{70% of the collection), as well as other commonly used health
and medicine websites such as Drugbank11, Diagnosia12 and Trip Answers13.
The crawled documents are provided in the dataset in their raw HTML
(Hyper Text Markup Language) format along with their uniform resource locators
(URL). The dataset is made available for download on the web to registered
participants on a secure password-protected server.
9 http://www.healthonnet.org
10 http://www.hon.ch/HONcode/Patients-Conduct.html
11 http://www.drugbank.ca/
12 http://www.diagnosia.com/
13 http://www.tripanswers.org/
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Discharge Summaries</title>
        <p>
          Novel methods to generate contextualized statements of patient information
needs were used. These are based on realistic short query statements created
in the context of patient discharge summaries. The discharge summaries can
be considered as a description of the context in which the patient has been
diagnosed with a given disorder and has written a query. The discharge
summaries originate from the de-identi ed MIMIC-II database14 (Multiparameter
Intelligent Monitoring in Intensive Care, Version 2.5). They are, together with
annotations, CLEF eHealth task 2 dataset [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>Discharge summaries are semi-structured reports with the following
appearance:
Admission Date : [ 2014 03 28 ]
D i s c h a r g e Date : [ 2014 04 08 ]
Date o f B i r t h : [ 1930 09 21 ]
Sex : F
S e r v i c e : CARDIOTHORACIC
A l l e r g i e s :
P a t i e n t r e c o r d e d as having No Known A l l e r g i e s to Drugs
Attending : [ Attending I n f o 5 6 5 ]
C h i e f Complaint : Chest pain
Major S u r g i c a l or I n v a s i v e Procedure :
Coronary a r t e r y bypass g r a f t 4 .</p>
        <p>H i s t o r y o f P r e s e n t I l l n e s s :
83 year o l d woman , p a t i e n t o f Dr . [ F i r s t Name4
( NamePattern1 ) ] [ Last Name ( NamePattern1 ) 5 0 0 5 ] ,
Dr . [ F i r s t Name ( S T i t l e ) 5 8 0 4 ] [ Name ( S T i t l e )
2 2 7 5 ] , with i n c r e a s e d SOB with a c t i v i t y , l e f t s h o u l d e r
b l a d e / back pain at r e s t , + MIBI , r e f e r r e d f o r c a r d i a c
cath . This p l e a s a n t 83 year o l d p a t i e n t n o t e s becoming
SOB when walking up h i l l s or i n c l i n e s about one year
ago . This SOB has p r o g r e s s i v e l y worsened and she i s now
SOB when walking [ 01 19 ] c i t y b l o c k ( f l a t s u r f a c e ) .
[ . . . ]
Past Medical H i s t o r y :
a r t h r i t i s ; c a r p a l t u n n e l ; s h i n g l e s r i g h t arm 2 0 0 0 ;
needs r i g h t knee r e p l a c e m e n t ; l e f t knee r e p l a c e m e n t
i n [ 2 0 1 0 ] ; thyroidectomy 1 9 7 8 ; c h o l e c y s t e c t o m y
[ 1 9 8 1 ] ; hysterectomy 2 0 0 1 ; h/o LGIB 2000 2001
a f t e r t a k i n g baby ASA; 81 QOD
[ . . . ]
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Topics</title>
        <p>In this section we describe the creation of the initial English topic set used in
Task 3a, and the translation of this topic set into Czech, French and German to
form a parallel topic corpus for use in Task 3b.</p>
        <p>English Topics The queries used in the task aim to model those used by
laypeople (i.e., patients, their relatives or other representatives) to nd out more
about their disorders, once they have examined a discharge summary.
14 http://mimic.physionet.org</p>
        <p>Topics to be used in this task have been created by experts (each expert
was a registered nurse and clinical documentation researcher) involved in the
CLEF eHealth consortium. This solution has been chosen in place of recruiting
patients because of the issues involved with recruitment and privacy. We believe
that, being on a daily basis in contact with patients receiving treatments and
discharge summaries, nurses are familiar with patients' information needs and
patient pro les.</p>
        <p>
          Topics have been manually created by the experts given discharge summaries,
and the discharge diagnosis. Last year's queries were generated from randomly
selected disorders. Therefore, the disorder was often not central enough in the
discharge summary for it to provide useful IR contextual information [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. This
year, queries were built based on one of the main disorders, identi ed from the
discharge summary, which the patient was hospitalized for. Discharge summaries
are semi-structured documents, and the discharge diagnosis is a eld that can
be found in 85% of the discharge summaries. The discharge diagnosis contains
on average 3 disorders. From these three, the experts selected one which a
patient may have questions on. For discharge summaries which had no discharge
diagnosis, experts selected a main disorder within the discharge summary, which
a patient may have questions on. Using the pairs of disorder and associated
discharge summary, the experts developed a set of patient queries (and criteria
for judging the relevance of documents to the queries, for use in the relevance
assessment task described in the next section). Queries are provided in a
standard TREC format, consisting of a topic title (text of the query), description
(longer description of what the query means), a narrative (expected content of
the relevant documents), and a pro le (brief description of the patient).
        </p>
        <p>The following example outlines a query:
&lt;query&gt;
&lt; t i t l e &gt; thrombocytopenia treatment c o r t i c o s t e r o i d s</p>
        <p>l e n g t h &lt;/ t i t l e &gt;
&lt;desc&gt; How l o n g s h o u l d be the c o r t i c o s t e r o i d s treatment</p>
        <p>to c u r e thrombocytopenia ? &lt;/desc&gt;
&lt;narr&gt; Documents s h o u l d c o n t a i n i n f o r m a t i o n about
t r e a t m e n t s o f thrombocytopenia , and e s p e c i a l l y
c o r t i c o s t e r o i d s . I t s h o u l d d e s c r i b e the treatment ,
i t s d u r a t i o n and how the d i s e a s e i s cured u s i n g i t .
&lt;s c e n a r i o &gt; The p a t i e n t has a s h o r t term d i s e a s e , or
has been h o s p i t a l i s e d a f t e r an a c c i d e n t ( l i t t l e to
no knowledge o f the d i s o r d e r , s h o r t term treatment )
&lt;/ s c e n a r i o &gt;
&lt;p r o f i l e &gt; P r o f e s s i o n a l f e m a l e &lt;/ p r o f i l e &gt;
&lt;/narr&gt;
&lt;/query&gt;</p>
        <p>With this approach, ve training and fty test queries have been generated
for use in the task.</p>
        <p>
          Translated Topics For the purpose of Task 3b, the original topics in
English were manually translated into Czech, German, and French. Based on our
previous experience with manual translation of medical user queries [
          <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
          ] the
translation was performed in three phases: First, the topics were translated from
English to the target languages by medical experts (one translator per language,
not necessarily native speakers but uent in the target languages). Second, the
translations were reviewed by language experts (native speakers or people with a
university degree in that language) and any language-related issues (typos,
grammar, etc.) were resolved. Third, any terminology issues were consulted with the
original translators and resolved together with the language experts.
        </p>
        <p>We asked the translators (and reviewers) to produce translations while
grammatically correct, preserve meaning and use terminology adequate to the
technical level of the original topic descriptions. Unlike the original topics, the resulting
translations do not contain any grammatical errors and typos.
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Relevance Assessment</title>
        <p>
          For this year's task, relevance judgements were collected from professional
assessors (but not medical experts). We used Relevation! [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]15 to manage the
collection of relevance assessments for documents in the assessment pool, where
each document was judged by one assessor.
        </p>
        <p>To form the assessment pool, we selected the top ten documents obtained
from the participants' baseline runs (run 1), their top-two priority runs using
discharge summaries (runs 2 and 3), and their top-two priority runs not using
discharge summaries (runs 5 and 6). This resulted in a pool of 6,800 documents,
in line with the size of the pool for the 2013 task. The relevance assessment
was based on a four point scale. The relevance grades are: (0) irrelevant, (1) on
topic but unreliable, (2) relevant, (3) highly relevant. These relevance grades are
mapped into a binary scale, with grades 0 and 1 corresponding to the binary
grade 0 (irrelevant) and grades 2 and 3 corresponding to the binary grade 1
(relevant). The graded relevance assessment yielded 0: 3,044, 1: 547, 2: 974, 3:
2,235 documents. The binary relevance assessments yielded 0: 3,591 non-relevant
and 1: 3,209 relevant documents. This year's assessment exercise yielded more
relevant documents per topic than last year: 64.18 relevant documents per topic
on average compared to last year's 37.56.</p>
        <p>
          Relevance assessments for the ve training queries were formed based on
pooled sets generated using the Vector Space Model [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and Okapi BM25 [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
Assessments for these ve training queries were conducted by two Finnish nurses.
Each document was assessed by one person. Training queries were distributed
to participants before the test queries were released.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>For this task, the participants were allowed to submit up to seven runs for the
English monolingual retrieval task, Task 3a. These runs comprised, one
mandatory run using no additional information or external resources (run 1), three
runs using the discharge summary and any other external resource (runs 2-4),
and three using external resources but not using the discharge summaries (run
15 http://ielab.github.io/relevation/
5-7). Among each set of additional runs, one had to use only the title and the
description elds of the query. Participants were also asked to rank their runs
2-4 and 5-7 according to their importance. For the cross-language information
retrieval task, Task 3b, participants could submit up to seven runs for each
language (Czech-English, French-English, German-English). These runs had the
same make-up as those in Task 3a.
4.1</p>
      <sec id="sec-4-1">
        <title>Participants</title>
        <p>This year, 91 groups registered for the task, 25 obtained access to the data and
14 submitted run(s) for task 3. The groups are from 11 countries in 4 continents
as listed in Table 1. While only one group from Europe participated last year,
this year the European groups formed the majority.</p>
        <p>Teams submitted in total 62 runs for task 3a in which 11 used discharge
summaries (from teams IRLabDAIICT, SNUMEDINFO, KISTI and Nijmegen).
For task 3b, 24 runs were submitted by two groups.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Evaluation Metrics</title>
        <p>We examined all documents in runs 1, 2, 3, 5 and 6 from Tasks 3a and 3b
up to rank 10 for relevance. The two major evaluation metrics are therefore
metrics at a cut-o of up to 10 documents, i.e. P@5, P@10, NDCG@5, and
NDCG@10. In addition, we considered MAP as an evaluation metric, but we
are aware that MAP is unreliable because only the top ten documents have
been assessed. Nevertheless, we wanted to report a measure covering the full set
of up to 1000 retrieved documents. We also report the number of relevant and
retrieved documents in the top 1000 results as a more recall-oriented measure.</p>
        <p>Performance metrics are computed with the standard trec eval tool16 using
the following commands:
{ -c -M1000 qrels.clef2014.test.bin.txt runName
{ -c -M1000 -m ndcg cut qrels.clef2014.test.graded.txt runName</p>
        <p>We are aware that the performance metrics for other runs might be unreliable
compared to that of runs 1, 2, 3, 5 and 6. However, this situation is common
for IR lab evaluations, where additional experiments on an existing data set
typically do not include re-assessment of documents previously not retrieved or
relevance assessment of additional documents.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Baseline System</title>
        <p>For comparison, we created our own baseline experiments by implementing a
number of information retrieval baselines: tf.idf (baseline.t df), BM25
(baseline.bm25), language modeling with Jelinek-Mercer smoothing (baseline.jm), and
language modeling with Dirichlet smoothing (baseline.dir). These methods do
not incorporate any domain-speci c adaptations. We used the implementations
of the above methods made available in the Indri toolkit17. Indri was also used
to parse the HMTL documents and for stemming (with Krovetz stemming, also
applied to queries). A stop list was applied to the queries but not to the
documents.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Evaluation Results</title>
        <p>The o cial results for all runs submitted to Task 3 (both 3a and 3b) and for our
baseline experiments (highlighted in italics) are shown in Tables 2 and 2, ordered
by decreasing P@10 (Task 3's primary measure). Comparing the participants'
results with respect to P@10 we observe that, for each team, the best e ectiveness
is often achieved when no discharge summaries are considered (runs 5, 6, 7 and
1, which is the teams' baseline); teams KISTI and NIJM are an exception to this
trend. A similar result was found also in the 2013 campaign, with most of the
teams achieving the highest e ectiveness when not using discharge summaries.</p>
        <p>Two teams submitted to the cross-lingual Task 3b: SNUMEDINFO and
CUNI. The results obtained by the SNUMEDINFO team when using the
crosslingual queries demonstrate comparable results to the corresponding
submissions when using English queries: in some cases cross-lingual queries yield even
higher results than the original English queries (e.g. SNUMEDINFO CZ Run.5 vs.
SNUMEDINFO EN Run.5), and these are comparable to the best results obtained
for the original English queries (Task 3a). This is not the case though for team
16 http://trec.nist.gov/trec eval/
17 www.lemurproject.org
CUNI, whose cross-lingual submissions generally yield less e ectiveness than the
corresponding Task 3a submissions.</p>
        <p>The best result in last year's task was obtained by TeamMayo, with a P@10 of
0.5180. This year's best run is obtained by team SNUMEDINFO with a P@10 of
0.7560. Even the baselines have considerably improved on 2014 dataset. Several
changes have been made between the two tasks: the document collection has
been reduced, and the query generation strategy has changed (from a randomly
selected disorder to the main one). One hypothesis to explain the increase could
be the fact that the topics are simpler, in the way that they correspond to
main disorders, that are potentially more frequent and more searched in general.
Further analysis is required to explain this improvement.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Approaches Used</title>
      <p>In this section we describe the approaches used by each team, and summarize
ndings from their analysis. Table 4 provides a condensed view of the techniques
and resources used by each team.</p>
      <p>
        Team CSKU-COMPL [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] used the vector space retrieval model of Lucene as
baseline. As improvement, they proposed a simple pseudo-relevance feedback
method which used the Genomic collection as external resource to perform query
expansion. The expansion terms selection is based on the Rocchio's formula
with dynamic tunable parameter of Pseudo-relevance feedback. Their best run
obtained P@10 of 0.5540.
      </p>
      <p>
        Team CUNI [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] participated in both tasks 3a and 3b, using only the query
titles and the Terrier platform (Hiemstra retrieval model) as their baseline. They
employed various methods for data cleaning and the simplest one, removing only
the HTML tags, had the best results. Their best run for task 3a used suggestions
from the MedlinePlus dictionary to x typos in the queries (P@10 of 0.5360).
They also employed query expansion adding the top ten highest terms from the
top 3 ranked documents, but this did not improve the results. For task 3b, only
one step was included, which was the translation of query titles using Khresmoi
translator system. Their best run here obtained P@10 of 0.4880 for Czech.
Team DEMIR [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] has as baseline the Terrier system. For each query they
predict whether query expansion is likely to improve retrieval performance or
not. Prediction is performed using a Naive Bayes classi er trained on the CLEF
eHealth 2013 test collection and features extracted from the queries and statistics
obtained from the collection. Their best result achieved P@10 of 0.67.
Team ERIAS [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] used the Vector Space Model in Lucene, indexing both
unigrams and bigrams for their baseline. The baseline system uses only the query
title as the query and uses no external resources. Other runs include query
expansion using synonymous terms and descendants from MeSH and the UMLS.
For identifying medical terms in queries, a method has been developed that
focuses on the most speci c terms, i.e. only medical terms not sub-parts of other
medical terms. Their best run obtained a P@10 of 0.5460.
      </p>
      <p>
        Team GRUIM [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] experimented with the use of the UMLS Metathesarus to
explore the e ectiveness of concept-based retrieval techniques. Their baseline
was based on Indri and Language Model with Dirichlet smoothing. They used
Metamap to annotate the documents and extract the medical concept. They also
experiment with query expansion using mutual information to determine related
concepts. Their best run obtained a P@10 of 0.75.
      </p>
      <p>
        Team IRLABDAIICT [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] indexed the document collection using Indri and used
the query likelihood model as their baseline. Other runs compared the Okapi
Model with the query likelihood model. They also experimented with using the
discharge summaries combined with MeSH terminology for query expansion.
Their best run was the baseline, which obtained P@10 of 0.70.
      </p>
      <p>
        Team KISTI [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] proposed a multiple-stage re-ranking method. Their baseline
used Lucene and query-likelihood with Dirichlet smoothing. It focuses on using
various retrieval techniques rather then using external resources and NLP
techniques. The sequential steps used are (i) query expansion with abbreviations,
(ii) query expansion with the discharge summary, (iii) clustering-based
document scoring, (iv) centrality-based document scoring using implicit links among
documents, and (v) pseudo relevance feedback. Their best run obtained a P@10
of 0.74, which applied steps (i), (ii) and (v).
      </p>
      <p>
        Team MIRACL [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] based their submissions on the Terrier retrieval system with
fairly standard settings for tokenization, stop word removal and stemming. Their
only run used a standard Vector Space Model, obtaining a MAP of 0.17 and a
P@10 of 0.55.
      </p>
      <p>
        Team Nijmegen [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] used the Language Modeling retrieval model of the Indri
search engine with Pseudo-Relevance feedback as their baseline. They employed
the Kullback-Leibler divergence for informativeness and phraseness method to
expand the query with terms from the discharge summaries (runs 2 to 4) and
UMLS-thesaurus (runs 5 to 7). The best result was found for run 4, where only
the discharge summaries were used for query expansion (P@10 of 0.6540).
Team RePALI [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] also opted for the Indri system as a baseline (parameters
estimated on the 2013 dataset), and experimented with various methods of
incorporating morpho-syntactic variants, lexical inclusion and hierarchical relations,
and abbreviations. However, results were inconsistent across the query set with
the reasons for this not being clear. Their best run obtained a P@10 of 0.67.
Team SNUMEDINFO [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] submitted to both Tasks 3a and 3b. As baseline,
they used the Indri retrieval system with Dirichlet smoothing language model.
They experimented with query expansion using the Metamap system, in which
candidate expansion keywords were ltered against the discharge summary
associated with the original query. They also experimented with learning to rank
based on random forests. They extracted features such as the \quality feature",
which, by counting how many terms from a pre-compiled list appear in a
document, attempts to estimate the reliability of the medical information presented
in the document. Their best run for Task 3a obtained a P@10 of 0.75. Their
cross-lingual submissions were based on the use of Google Translate, and their
best run here obtained a P@10 of 0.75 for Czech.
      </p>
      <p>
        Team UHU (LABERINTO) [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] used a standard system built on Lucene and
experimented with methods for term boosting and query expansion. They
submitted 4 runs not using the discharge summaries. In run 5, a boosting factor of
1.5 was applied to query terms which appear in UMLS, which increased P@10
from the baseline of 0.56 to 0.58. Query expansion, realized by adding MeSH
descriptors for query terms appearing both in title and description, did not improve
the baseline results.
      </p>
      <p>
        Team UIOWA [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] included all webpage content in their document index, as
opposed to just body text. They used Indri to generate their baseline. The other
approaches they explored performed worse than this baseline (P@10 of 0.69).
They experimented with pseudo relevance feedback and using the Markov
Random Field Model with medical phrase bigrams extracted from MetaMap for
query expansion.
      </p>
      <p>
        Team YORKU [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] has as the core of their approach the use of Learning to
Rank with a total of 231 features from multiple information retrieval models
and di erent parameter settings. The group submitted several runs, in which
they compare binary and graded relevance information, as well as the use of
di erent machine learning algorithms. Their best run obtained a P@10 of 0.60.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this second year of the ShARe/CLEF eHealth2014 evaluation lab Task 3,
there was strong take-up in the community with 14 groups submitting runs to
the task. The challenge of developing retrieval techniques for layperson medical
queries proved di cult.</p>
      <p>Overall, we observed a considerable improvement over 2013 results, both for
the team runs and the baselines. The best run for task 3a was submitted by team
GRIUM, with a P@10 of 0.7560 and a NDCG@10 of 0.7445. The best run for
task 3b was submitted by team SNUMEDINFO on the Czech topics, with P@10
of 0.7551 and NDCG@10 of 0.7011 (their P@10 is slightly higher for Czech topics
than for English ones). The three best teams use language modelling retrieval
methods, perform some query expansion and two of them use UMLS. The best
team for task 3b used Google Translate18 to translate the queries.</p>
      <p>This year, we implemented several state-of-the-art baselines. The highest
performances are achieved using language models with Dirichlet smoothing.</p>
      <p>Four teams submitted runs using the discharge summaries. Two of the
top10 runs (ranked with P@10) use them: SNUMEDINFO and KISTI. Moreover,
all the runs using discharge summaries for these two teams obtain higher results
than their runs without discharge summaries. This is an improvement over 2013,
where no team managed to improve their results with the discharge summaries.
Our new topic generation strategy proved to be more accurate, and discharge
summaries seem to bring useful contextual information to better retrieve
documents.</p>
      <p>
        Given the success of the rst two years of the task, we anticipate even more
interest in next year's campaign. In the third year of this task, we will explore
new topic generation strategies, based on our related research on automatic
generation of queries [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] and analysis of query complexity [37]. Moreover, we
intend to perform more analysis work to better understand the task results and
IR methods to answer laypeople medical information needs.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgement</title>
      <p>Task 3 of the ShARe/CLEFeHealth2013 evaluation lab has been supported in
part by the Khresmoi project, funded by the European Union Seventh
Framework Programme (FP7/2007-2013) under grant agreement no 257528. We
acknowledge the time given to perform the relevance assessment task. We want to
thank the following individuals: Riitta Danielsson-Ojala (University of Turku,
Finland), Sanna Salantera (University of Turku, Finland) for creating the queries
and conducting relevance assessment, and Ondrej Dusek (Charles University),
Brendan Hegarty (Dublin City University), Jaroslava Hlavacova (Charles
University), John Hodmon (Dublin City University), Michal Novak (Charles
University), David Racca (Dublin City University), Rudolf Rosa and Daniel Zeman
(Charles University) for their help on the relevance assessment. We also
acknowledge the time given by Margit Hanbury to check the German translations.
18 http://translate.google.com</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schrek</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leroy</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mowery</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
          </string-name>
          , J.:
          <article-title>Overview of the share/clef ehealth evaluation lab 2014</article-title>
          .
          <source>In: Proceedings of CLEF 2014. Lecture Notes in Computer Science (LNCS)</source>
          , Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Health topics:
          <volume>80</volume>
          %
          <article-title>of internet users look for health information online</article-title>
          .
          <source>Technical report</source>
          , Pew Research Center (
          <year>February 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tong</surname>
            ,
            <given-names>R.M.:</given-names>
          </string-name>
          <article-title>Overview of the TREC 2011 medical records track</article-title>
          .
          <source>In: Proceedings of TREC, NIST</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kalpathy-Cramer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Bedrick</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eggel</surname>
            , I., de Herrera,
            <given-names>A.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The CLEF 2011 medical image retrieval and classi cation tasks</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          (
          <article-title>Cross Language Evaluation Forum)</article-title>
          . (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Muller, H.,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deselaers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caputo</surname>
          </string-name>
          , B., eds.: ImageCLEF |
          <article-title>Experimental Evaluation in Visual Information Retrieval</article-title>
          . Volume
          <volume>32</volume>
          of The Information Retrieval Series. Springer (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H., Salantera,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>ShARe/CLEF eHealth Evaluation Lab 2013, task 3: Information retrieval to address patients' questions when reading clinical reports</article-title>
          .
          <source>In: CLEF online working notes</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horvitz</surname>
          </string-name>
          , E.:
          <article-title>Cyberchondria: Studies of the escalation of medical concerns in web search</article-title>
          .
          <source>Technical report, Microsoft Research</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Boyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gschwandtner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kritz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pletneva</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vargas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Use case de nition including concrete data requirements (D8.2). public deliverable, Deliverable of the Khresmoi EU project (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Suominen</surname>
          </string-name>
          , H., ed.:
          <source>The Proceedings of the CLEFeHealth2012 { the CLEF 2012 Workshop on Cross-Language Evaluation of Methods</source>
          , Applications, and
          <article-title>Resources for eHealth Document Analysis</article-title>
          .
          <source>NICTA</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hersh</surname>
            ,
            <given-names>W.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leone</surname>
            ,
            <given-names>T.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hickam</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          :
          <article-title>OHSUMED: An interactive retrieval evaluation and new large test collection for research</article-title>
          .
          <source>In: Proceedings of SIGIR '94</source>
          . (
          <year>1994</year>
          )
          <volume>192</volume>
          {
          <fpage>201</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Claveau</surname>
          </string-name>
          , V.:
          <article-title>Unsupervised and semi-supervised morphological analysis for information retrieval in the biomedical domain</article-title>
          .
          <source>In: Proceedings of COLING</source>
          . (
          <year>2012</year>
          )
          <volume>629</volume>
          {
          <fpage>645</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bruza</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sitbon</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawley</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An evaluation of corpus-driven measures of medical concept similarity for information retrieval</article-title>
          .
          <source>In: Proceedings of CIKM</source>
          <year>2012</year>
          .
          <article-title>(</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clematide</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rinaldi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kafkas</surname>
          </string-name>
          , S., van
          <string-name>
            <surname>Mulligen</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bui</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellrich</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milward</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poprat</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimeno-Yepes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kors</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Multilingual semantic resources and parallel corpora in the biomedical domain: the clef-er challenge</article-title>
          .
          <source>In: CLEF online working notes</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Leveling</surname>
          </string-name>
          , J.:
          <article-title>Creation of a new evaluation benchmark for information retrieval targeting patient information needs</article-title>
          . In Song, R.,
          <string-name>
            <surname>Webber</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kishida</surname>
          </string-name>
          , K., eds.
          <source>: Proceedings of the 5th International Workshop on Evaluating Information Access (EVIA)</source>
          ,
          <source>a Satellite Workshop of the NTCIR-10 Conference</source>
          , Tokyo/Fukuoka, Japan, National Institute of Informatics/Kijima Printing (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H.:
          <article-title>Khresmoi { multimodal multilingual medical information search</article-title>
          . In:
          <article-title>MIE village of the future</article-title>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mowery</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>South</surname>
            ,
            <given-names>B.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christensen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elhadad</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pradhan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W.W.</given-names>
          </string-name>
          :
          <article-title>Task 2: Share/clef ehealth evaluation lab 2014</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          <year>2014</year>
          .
          <article-title>(</article-title>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Uresova</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dusek</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Multilingual test sets for machine translation of search queries for cross-lingual information retrieval in the medical domain</article-title>
          . In Chair),
          <string-name>
            <given-names>N.C.C.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Loftsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          , S., eds.
          <source>: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          , Reykjavik, Iceland, European Language Resources Association (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dusek</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hlavacova</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marecek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosa</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamchyna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uresova</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Adaptation of machine translation for multilingual information retrieval in the medical domain</article-title>
          .
          <source>Arti cial Intelligence in Medicine</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>Relevation!: An open source system for information retrieval relevance assessment</article-title>
          .
          <source>In: Proceedings of the 37th annual international ACM SIGIR conference on research and development in information retrieval.</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.S.:</given-names>
          </string-name>
          <article-title>A vector space model for automatic indexing</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>18</volume>
          (
          <issue>11</issue>
          ) (
          <year>1975</year>
          )
          <volume>613</volume>
          {
          <fpage>620</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Simple, proven approaches to text retrieval</article-title>
          .
          <source>Technical Report 356</source>
          , University of Cambridge (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Thesprasith</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaruskulchai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Csku gprf-qe for medical topic web retrieval</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Saleh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Cuni at the ShARe/CLEF eHealth Evaluation Lab 2014</article-title>
          .
          <article-title>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</article-title>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Ozturkmenoglu</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alpkocak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kilinc</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Demir at CLEF eHealth:
          <article-title>The effects of selective query expansion to information retrieval</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Drame</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mougin</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diallo</surname>
          </string-name>
          , G.:
          <article-title>Query expansion using external resources for improving information retrieval in the biomedical domain</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>J.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liui</surname>
            ,
            <given-names>X.:</given-names>
          </string-name>
          <article-title>An investigation of the e ectiveness of concept-based approach in medical information retrieval GRIUM @ CLEF2014eHealthTask 3</article-title>
          . In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab. (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Thakkar</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iyer</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majumder</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Team IRLabDAIICT at ShARe/- CLEF eHealth
          <year>2014</year>
          <article-title>Task 3: User-centered Information Retrieval system for Clinical Documents</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>H.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jung</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>A multiple-stage approach to re-ranking clinical documents</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Ksentini</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tmar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gargouri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <source>Miracl at CLEF</source>
          <year>2014</year>
          :
          <article-title>eHealth information retrieval task</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Verberne</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A language-modelling approach to user-centred health information retrieval</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Claveau</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grabar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , ,
          <string-name>
            <surname>Maguer</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          :
          <article-title>RePaLi participation to CLEF eHealth IR challenge 2014: leveraging term variation</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
          </string-name>
          , J.:
          <article-title>Exploring e ective information retrieval technique for the medical web documents: SNUMedinfo at CLEFeHealth2014 Task 3</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Malagon</surname>
            ,
            <given-names>J.M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>na Lopez</surname>
            ,
            <given-names>M.J.M.:</given-names>
          </string-name>
          <article-title>Laberinto at ShARe/CLEF eHealth Evaluation Lab</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhattacharya</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srinivasan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : The University of Iowa at CLEF 2014:
          <article-title>eHealth Task 3</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
          </string-name>
          , J.: York University at CLEF eHealth
          <year>2014</year>
          :
          <article-title>A Learning-to-Rank Approach for Medical Document Retrieval</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Salantera,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Building realistic potential patients queries for medical information retrieval evaluation</article-title>
          .
          <source>In: Proceedings of the LREC workshop on Building and Evaluating Resources for Health and Biomedical Text Processing</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>