<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLEF eHealth Evaluation Lab 2015, Task 2: Retrieving Information About Medical Symptoms</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jo~ao Palotti</string-name>
          <email>p@10</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guido Zuccon</string-name>
          <email>g.zuccon@qut.edu.au</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <email>lorraine.goeuriot@imag.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liadh Kelly</string-name>
          <email>Liadh.Kelly@tcd.ie</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Allan Hanbury</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J.F. Jones</string-name>
          <email>gareth.jones@computing.dcu.ie</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihai Lupu</string-name>
          <email>lupu@ifs.tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavel Pecina</string-name>
          <email>pecina@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University in Prague</institution>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Queensland University of Technology</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Trinity College</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Universite Grenoble Alpes</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Vienna University of Technology</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper details methods, results and analysis of the CLEF 2015 eHealth Evaluation Lab, Task 2. This task investigates the e ectiveness of web search engines in providing access to medical information with the aim of fostering advances in the development of these technologies. The problem considered in this year's task was to retrieve web pages to support information needs of health consumers that are confronted with a sign, symptom or condition and that seek information through a search engine, with the aim to understand which condition they may have. As part of this evaluation exercise, 66 query topics were created by potential users based on images and videos of conditions. Topics were rst created in English and then translated into a number of other languages. A total of 97 runs by 12 di erent teams were submitted for the English query topics; one team submitted 70 runs for the multilingual topics.</p>
      </abstract>
      <kwd-group>
        <kwd>Medical Information Retrieval</kwd>
        <kwd>Health Information Seeking and Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This document reports on the CLEF 2015 eHealth Evaluation Lab, Task 2. The
task investigated the problem of retrieving web pages to support information
needs of health consumers (including their next-of-kin) that are confronted with
a sign, symptom or condition and that use a search engine to seek
understanding about which condition they may have. Task 2 has been developed within
the CLEF 2015 eHealth Evaluation Lab, which aims to foster the development
? In alphabetical order, JP, GZ led Task 2; LG, LK, AH, ML &amp; PP were on the Task
2 organising committee.
of approaches to support patients, their next-of-kin, and clinical sta in
understanding, accessing and authoring health information [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        The use of the Web as source of health-related information is a wide-spread
phenomena. Search engines are commonly used as a means to access health
information available online [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Previous iterations of this task (i.e. the 2013 and 2014
CLEFeHealth Lab Task 3 [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]) aimed at evaluating the e ectiveness of search
engines to support people when searching for information about their conditions,
e.g. to answer queries like \thrombocytopenia treatment corticosteroids length".
These past two evaluation exercises have provided valuable resources and an
evaluation framework for developing and testing new and existing techniques.
The fundamental contribution of these tasks to the improvement of search engine
technology aimed at answering this type of health information need is
demonstrated by the improvements in retrieval e ectiveness provided by the best 2014
system [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] over the best 2013 system [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (using di erent, but comparable, topic
sets).
      </p>
      <p>
        Searching for self-diagnosis information is another important type of health
information seeking activity [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]; this seeking activity has not been considered
in the previous CLEF eHealth tasks, nor in other information retrieval
evaluation campaigns. These information needs often arise before attending a medical
professional (or to help the decision of attending). Previous research has shown
that exposing people with no or scarce medical knowledge to complex medical
language may lead to erroneous self-diagnosis and self-treatment and that access
to medical information on the Web can lead to the escalation of concerns about
common symptoms (e.g., cyberchondria) [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. Research has also shown that
current commercial search engines are yet far from being e ective in answering
such queries [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This type of query is the subject of investigation in this CLEF
2015 eHealth Lab Task 2. We expected these queries to pose a new challenge
to the participating teams; a challenge that, if solved, would lead to signi cant
contributions towards improving how current commercial search engines answer
health queries.
      </p>
      <p>The remainder of this paper is structured as follows: Section 2 details the
task, the document collection, topics, baselines, pooling strategy, and evaluation
metrics; Section 3 presents the participants' approaches, while Section 4 presents
their results; Section 5 concludes the paper.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>The CLEF 2015 eHealth Task 2</title>
      <sec id="sec-2-1">
        <title>The Task</title>
        <p>The goal of the task is to design systems which improve health search, especially
in the case of search for self-diagnosis information. The dataset provided to
participants is comprised of a document collection, topics in various languages,
and the corresponding relevance information. The collection was provided to
participants after signing an agreement, through the PhysioNet website7.
7 http://physionet.org/</p>
        <p>Participating teams were asked to submit up to ten runs for the English
queries, and an additional ten runs for each of the multilingual query languages.
Teams were required to number runs such as that run 1 was a baseline run for
the team; other runs were numbered from 2 to 10, with lower numbers indicating
higher priority for selection of documents to contribute to the assessment pool
(i.e. run 2 was considered of higher priority than run 3).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Document Collection</title>
        <p>
          The document collection provided in the CLEF 2014 eHealth Lab Task 3 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
is also adopted in this year's task. Documents in this collection have been
obtained through a large crawl of health resources on the Web; the collection
contains approximately one million documents and originated from the Khresmoi
project8 [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The crawled domains were predominantly health and medicine
sites, which were certi ed by the HON Foundation as adhering to the
HONcode principles (appr. 60{70% of the collection), as well as other commonly
used health and medicine sites such as Drugbank, Diagnosia and Trip Answers9.
Documents consisted of web pages on a broad range of health topics and were
likely targeted at both the general public and healthcare professionals. They
were made available for download in their raw HTML format along with their
URLs to registered participants.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Topics</title>
        <p>
          Queries were manually built with the following process: images and videos
related to medical symptoms were shown to users, who were then asked which
queries they would issue to a web search engine if they, or their next-of-kin, were
exhibiting such symptoms. Thus, these queries aimed to simulate the situation
of health consumers seeking information to understand symptoms or conditions
they may be a ected by; this is achieved using imaginary or video stimuli. This
methodology for eliciting circumlocutory, self-diagnosis queries was shown to be
e ective by Stanton et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]; Zuccon et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] showed that current commercial
search engines are yet far from being e ective in answering such queries.
        </p>
        <p>
          Following the methodology in [
          <xref ref-type="bibr" rid="ref11 ref9">9, 11</xref>
          ], 23 symptoms or conditions that
manifest with visual or audible signs (e.g. ringworm or croup) were selected to be
presented to users to collect queries. A cohort of 12 volunteer university
students and researchers based in the organisers' institutions generated the queries.
English was the mother-tongue for all volunteers and they had no particular
prior knowledge about the symptoms or conditions, nor they had any speci c
medical background: this cohort was then somehow representative of the average
user of web search engines seeking health advice (although they had a higher
8 Medical Information Analysis and Retrieval, http://www.khresmoi.eu
9 Health on the Net, http://www.healthonnet.org, http://www.hon.ch/HONcode/
Patients-Conduct.html, http://www.drugbank.ca, http://www.diagnosia.com,
and http://www.tripanswers.org
education level than the average level). Each volunteer was given 10 conditions
for which they were asked to generate up to 3 queries per condition (thus each
condition/image pair was presented to more than one assessor10). An example
of images and instructions provided to the volunteers is given in Figure 111.
Imagine you are experiencing the health problem shown below.
        </p>
        <p>Please provide 3 search queries that you would issue to find out what is wrong.
Instructions:
* You must provide 3 distinct search queries.
* The search queries must relate to what you see below.</p>
        <p>A total of 266 possible unique queries were collected; of these, 67 queries
(22 conditions with 3 queries and 1 condition with 1 query) were selected to be
used in this year's task. Queries were selected by randomly picking one query per
condition (we called this the pivot query), and then manually selecting the query
that appeared most similar (called most ) and the one that appeared least similar
(called least ) to the pivot query. Candidates for the most and least queries were
identi ed independently by three organisers and then majority voting was used
to establish which queries should be selected. This set of queries formed the
English query set distributed to participants to collect runs.</p>
        <p>In addition, we developed translations of this query set into Arabic (AR),
Czech (CS), German (DE), Farsi (FA), French (FR), Italian (IT) and Portuguese
(PT); these formed the multilingual query sets which were made available to
participants for submission of multilingual runs. Queries were translated by medical
experts available at the organisers institutions.</p>
        <p>After the query set was released, numbered qtest1-qtest67, one typo was
found in query qtest62, which could compromise the translations. In order to
keep consistency between the English query and all translations made by the
experts, qtest62 was excluded. Thus, the nal query set used in the CLEF 2015
10 With exception of one condition, for which only one query could be generated.
11 Note that additional instructions were given to volunteers at the start and end of
the task, including training and de-brie ng.
eHealth Lab Task 2 for both English and multilingual queries consisted of 66
queries.</p>
        <p>
          An example of one of the query topics generated from the image shown in
Figure 1 is provided in Figure 2. To develop their submissions, participants were
only given the query eld of each query topic, that is, teams were unaware of
the query type (pivot, most, least), the target condition and the image or video
that was shown to assessors to collect queries.
&lt;topics&gt;
...
&lt;top&gt;
&lt;num&gt;qtest.23&lt;/num&gt;
&lt;query&gt;red bloodshot eyes&lt;/query&gt;
&lt;disease&gt;non-ulcerative sterile keratitis&lt;/disease&gt;
&lt;type&gt;most&lt;/type&gt;
&lt;query_index&gt;22&lt;/query_index&gt;
&lt;/top&gt;
...
&lt;/topics&gt;
Relevance assessments were collected by pooling participants' submitted runs
as well as baseline runs (see below for a description of pooling methodology
and baseline runs). Assessment was performed by ve paid medical students
employed at the Medizinische Universitat Graz (Austria); assessors used
Relevation! [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] to visualise and judge documents. For each document, assessors
had access to the query the document was retrieved for, as well as the target
symptom or condition that was used to obtained the query during the query
generation phase.
        </p>
        <p>Target symptoms or conditions were used to provide the relevance criteria
assessors should judge against; for example for query qtest1 { \many red marks on
legs after traveling from US" (the condition used for generating the query was
\Rocky Mountain spotted fever (RMSF)"), the relevance criterion read
\Relevant documents should contain information allowing the user to understand
that they have Rocky Mountain spotted fever (RMSF).". Relevance assessments
were provided on a three point scale: 0, Not Relevant; 1, Somewhat Relevant; 2,
Highly Relevant.</p>
        <p>
          Along with relevance assessments, readability judgements were also collected
for the assessment pool. The notion of readability and understandability of
information is of important concern when retrieving information for health
consumers [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. It has been shown that if the readability of information is accounted
for in the evaluation framework, judgements of relative system e ectiveness can
vary with respect to taking into account (topical) relevance only [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] (this was
the case also when considering the CLEF 2013 and 2014 eHealth Evaluation
Labs).
        </p>
        <p>Readability assessments were collected by asking the assessors whether they
believed a patient would understand the retrieved document. Assessments were
provided on a four point scale: 0, \It is very technical and di cult to read and
understand"; 1, \It is somewhat technical and di cult to read and understand";
2, \It is somewhat easy to read and understand"; 3, \It is very easy to read and
understand".
2.5</p>
      </sec>
      <sec id="sec-2-4">
        <title>Example Topics</title>
        <p>
          A di erent set of 5 queries was released to participants as example queries (called
training ) to help develop their systems (both in English and the other
considered languages). These queries were released together with associated relevance
assessments, obtained by evaluating a pool of 112 documents retrieved by a set
of baseline retrieval systems (TF-IDF, BM25, Language Model with Dirichlet
smoothing as implemented in Terrier [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], with the associated default
parameter values); the pool was formed by sampling the top 10 retrieved documents
for each query. Note that, given the very limited pool and system sample sizes,
these example queries should not be used to evaluate, tune or train systems.
2.6
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Baseline Systems</title>
        <p>The organisers generated baseline runs using BM25, TF-IDF and Language
Model with Dirichlet smoothing, as well as a set of benchmark systems that
ranked documents by estimating both (topical) relevance and readability.
Table 1 shows the 13 baseline systems created, 7 of them took into consideration
some estimation of text readability. No baselines were created for the
multilingual queries.</p>
        <p>
          The rst 6 baselines, named baseline1 -6, were created using either Xapian or
Lucene as retrieval toolkit. We vary the retrieval model used, including BM25
(with parameters k1 = 1; k2 = 0; k3 = 1 and b = 0:5) in baseline1, Vector Space
Model (VSM) with TF-IDF weighting (the default Lucene implementation) in
baseline2, and Language Model (LM), with Dirichlet smoothing with = 2; 000
in baseline3. Our preliminary runs based on the 2014 topics showed that
removing HTML tags from documents in this collection could lead to higher retrieval
e ectiveness when using BM25 and LM. We used the python package
BeautifulSoap (BS4)12 to parse the HTML les and remove HTML tags. Note that it
12 https://pypi.python.org/pypi/beautifulsoup4
does not remove the boilerplate from the HTML (such as headers or navigation
menus), being one of the simplest approaches to clean a HTML page and prepare
it to serve as the input of readability formulas [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] (see below). Baselines 4, 5
and 6 implement the same methods as in baselines 1, 2 and 3, respectively, but
execute a query that has been enhanced by augmenting the original query with
the known target disease names. Note that the target disease names were only
known to the organisers, participants had no access to this information.
        </p>
        <p>
          For the baseline runs that take into account readability estimations, we used
two well-known automatic readability measures: the Dale-Chall measure [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]
and Flesch-Kincaid readability index [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. The python package
ReadabilityCalculator13 was used to compute the readability measures from the cleansed web
documents. We also tested a readability measure based on the frequency of
words in a large collection such as Wikipedia; the intuition behind this measure
is that an easy text would contain a large number of common words with high
frequency in Wikipedia, while a technical and di cult text would have a large
number of rare words, characterised by a low frequency in Wikipedia. In order
to retrieving documents accounting for their readability levels, we rst generate
a readability score Read(d) for each document d in the collection using one of
the three measures above. We then combine the readability score of a document
with its relevance score Rel(d) generated by some retrieval model. Three score
combination methods were considered:
1. Linear combination: Score(d) = Rel(d) + (1:0
        </p>
        <p>is a hyperparameter and 0 1 (in readability1
2. Direct Multiplication: Score(d) = Rel(doc) Read(d)</p>
        <p>Rel(doc)
3. Inverse Logarithm: Score(d) = log(Read(d))
) Read(d), where
is 0.9)</p>
        <p>Table 1 shows the settings of retrieval model, HTML processing, readability
measure and query expansion or score combination method that were considered
to produce the 7 readability baselines used in the task.
2.7</p>
      </sec>
      <sec id="sec-2-6">
        <title>Pooling Methodology</title>
        <p>In Task 2, for each query, the top 10 documents returned in runs 1, 2 and 3
produced by the participants14 were pooled to form the relevance assessment
pool. In addition, the baseline runs developed by the organisers were also pooled
with the same methodology used for participants runs. A pool depth of 10
documents was chosen because this task resembles web-based search, where often
users consider only the rst page of results (that is, the rst 10 results). Thus,
this pooling methodology allowed a full evaluation of the top 10 results for the 3
submissions with top priority for each participating team. The pooling of more
submissions or a deeper pool, although preferable, was ruled out because of the
limited availability of resources for document relevance assessment.
13 https://pypi.python.org/pypi/ReadabilityCalculator/
14 With the exclusion of multilingual submissions, for which runs were not pooled due
to the larger assessment e ort pooling these runs would have required. Note that
only one team submitted multilingual runs.</p>
      </sec>
      <sec id="sec-2-7">
        <title>Multilingual Evaluation: Additional Pooling and Relevance</title>
      </sec>
      <sec id="sec-2-8">
        <title>Assessments</title>
        <p>Because only one team submitted runs for the multilingual queries and only
limited relevance assessment capabilities were available through the paid medical
students that performed the assessment of submissions for the English queries,
multilingual runs were not considered when forming the pools for relevance
assessments. However, additional relevance assessments were sought through the
team that participated in the multilingual task: they were thus asked to perform
a self-assessment of the submissions they produced. A new pool of documents
was sampled with the same pooling methodology used for English runs (see the
previous section). Documents that were already judged by the o cial assessors
were excluded from the pool with the aim to limit the additional relevance
assessment e ort required by the team.</p>
        <p>
          Additional relevance assessments for the multilingual runs were then
performed by a medical doctor (native Czech speaker with uent English)
associated with Team CUNI [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. The assessor was provided with the same
instructions and assessment system that the o cial assessors used. Assessments were
collected and aggregated with those provided by the o cial relevance assessors
to form the multilingual merged qrels. These qrels should be used with caution:
at the moment of writing this paper, it is unknown whether these multilingual
assessments are comparable with those compiled by the original, also medically
trained, assessors. The collection of further assessments from the team to
verify their agreement with the o cial assessors is left for future work. Another
limitation of these additional relevance assessments is that only one system that
considered multilingual queries, that developed by team CUNI, was sampled and
thus it may further bias the assessment of retrieval systems with respect to how
multilingual queries are coped with.
Evaluation was performed in terms of graded and binary assessments. Binary
assessments were formed by transforming the graded assessments such that label
0 was maintained (i.e. irrelevant) and labels 1 and 2 were converted to 1
(relevant). Binary assessments for the readability measures were obtained similarly,
with labels 0 and 1 being converted into 0 (not readable) and labels 2 and 3
being converted into 1 (readable).
        </p>
        <p>
          System evaluation was conducted using precision at 10 (p@10) and
normalised discounted cumulative gain [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] at 10 (nDCG@10) as the primary and
secondary measures, respectively. Precision was computed using the binary
relevance assessments; nDCG was computed using the graded relevance assessments.
These evaluation metrics were computed using trec eval with the following
commands:
./trec eval -c -M1000 qrels.clef2015.test.bin.txt runName
./trec eval -c -M1000 -m ndcg cut qrels.clef2015.test.graded.txt runName
respectively to compute precision and nDCG values.
        </p>
        <p>
          A separate evaluation was conducted using both relevance assessments and
readability assessments following the methods in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. For all runs, Rank Biased
Precision (RBP) [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] was computed along with readability-biased modi cations
of RBP, namely uRBP (using the binary readability assessments) and uRBPgr
(using the graded readability assessments).
        </p>
        <p>
          The RBP parameter which attempts to model user behaviour15 (RBP
persistence parameter) was set to 0.8 for all variations of this measure, following
the ndings of Park and Zhang [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
        <p>To compute uRBP, readability assessments were mapped to binary classes,
with assessments 0 and 1 (indicating low readability) mapped to value 0 and
assessments 2 and 3 (indicating high readability) mapped to value 1. Then,
uRBP (up to rank K) was calculated according to</p>
        <p>K
) X</p>
        <p>k=1
uRBP = (1
k 1r(k)u(k)
(1)
where r(k) is the standard RBP gain function that is 1 if the document at rank k
is relevant and 0 otherwise; u(k) is a similar gain function but for the readability
dimension and is 1 if the document at k is readable (binary class 1), and zero
otherwise (binary class 0).</p>
        <p>To compute uRBPgr, i.e. the graded version of uRBP, each readability label
was mapped to a di erent gain value. Speci cally, label 0 was assigned gain 0
(least readability, no gain), label 1 gain 0.4, label 2 gain 0.8 and label 3 gain 1
(highest readability, full gain). Thus, a document that is somewhat di cult to
read does still generate a gain, which is half the gain generated by a document
15 High values of
users.</p>
        <p>representing persistent users, low values representing impatient
that is somewhat easy to read. These gains are then used to evaluate the function
u(k) in Equation 1 to obtain uRBPgr.</p>
        <p>The readability-biased evaluation was performed using ubire16, which is
publicly available for download.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Participants and Approaches</title>
      <sec id="sec-3-1">
        <title>Participants</title>
        <p>
          This year, 52 groups registered for the task on the web site, 27 got access to the
data and 12 submitted any run for task 2. The groups are from 9 countries in 4
continents as listed in Table 2. 7 out of the 12 participants had never participated
in this task before.
Team CUNI [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] used the Terrier toolkit to produce their submissions. Runs
explored three di erent retrieval models: Bayesian smoothing with Dirichlet prior,
Per- eld normalisation (PL2F) and LGD. Query expansion using the UMLS
metathesaurus was explored by considering terms assigned to the same
concept as synonymous and choosing the terms with the highest inverse
documentfrequency. Blind relevance feedback was also used as contrasting technique.
Finally, they also experimented with linear interpolations of the search results
produced by the above techniques.
16 https://github.com/ielab/ubire
Team ECNU [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] explored query expansion and learning to rank. For query
expansion, Google was queried and the titles and snippets associated with the
top ten web results were selected. Medical terms were then extracted from these
resources by matching them with terms contained in MeSH; the query was then
expanded using those medical terms that appeared more often than a threshold.
As Learning to Rank approach, Team ECNU combined scores and ranks from
BM25, PL2 and BB2 into a six-dimensional vector. The 2013 and 2014 CLEF
eHealth tasks were used to train the system and a Random Forest classi er was
use to calculate the new scores.
        </p>
        <p>Team FDUSGInfo explored query expansion methods that use a range of
knowledge resources to improve the e ectiveness of a statistical Language Model
baseline. The knowledge sources that have been considering for drawing expansion
terms are MeSH and Freebase. Di erent techniques were evaluated to select the
expansion terms, including manual term selection. Team FDUSGInfo,
unfortunately, did not submit their working notes and thus the details of their methods
are unknown.</p>
        <p>
          Team GRIUM [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] explored the use of concept based query expansion. Their
query expansion mechanism exploited Wikipedia articles and UMLS Concept
de nitions and were compared to a baseline method based on Dirichlet
smoothing.
        </p>
        <p>
          Team KISTI [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] focused on re-ranking approaches. Lucene was used for
indexing and initial search, and the baseline used the query likelihood model with
Dirichlet smoothing. They explored three approaches for re-ranking: explicit
semantic analysis (ESA), concept-based document centrality (CBDC), and
clusterbased external expansion model (CBEEM). Their submissions evaluated these
re-ranking approaches as well as their combinations.
        </p>
        <p>
          Team KUCS [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] implemented an adaptive query expansion. Based on the results
returned by a query performance prediction approach, their method selected the
query expansion that is hypothesised to be the most suitable for improving
e ectiveness. An additional process was responsible for re-ranking results based
on readability estimations.
        </p>
        <p>
          Team LIMSI [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ] explored query expansion approaches that exploit external
resources. Their rst approach used MetaMap to identify UMLS concepts from
which to extract medical terms to expand the original queries. Their second
approach used a selected number of Wikipedia articles describing the most common
diseases and conditions along with a selection of MedlinePlus; for each query the
most relevant articles from these corpora are retrieved and their titles used to
expand the original queries, which are in turn used to retrieve relevant documents
from the task collection.
        </p>
        <p>
          Team Miracl [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]'s submissions were based on blind relevance feedback combined
with term selection using their previous work on modeling semantic relations
between words. Their baseline run was based on the traditional Vector Space
Model and the Terrier toolkit. The other runs employed the investigated method
by varying settings of two method parameters: the rst controlling the number of
highly ranked documents from the initial retrieval step and the second controlling
the degree of semantic relationship of the expansion terms.
        </p>
        <p>
          Team HCMUS [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] experimented with two approaches. The rst was based on
concept-based retrieval where only medical terminological expressions in
documents were retained, while other words were ltered-out. The second was based
on query expansion with blind relevance feedback. Common to all their
approaches was the use of Apache Lucene and a bag-of-word baseline based on
Language Modelling with Dirichlet smoothing and standard stemming and
stopword removal.
        </p>
        <p>
          Team UBML [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] investigated the empirical merits of query expansion based
on KL divergence and the Bose-Einstein 1 model for improving a BM25
baseline. The query expansion process selected terms from the local collection or
two external collections. Learning to rank was also investigated along a Markov
Random Fields approach.
        </p>
        <p>
          Team USST [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] used BM25 as a baseline system and explored query expansion
approaches. They investigated pseudo relevance feedback approaches based on
Kullback-Liebler Divergence and Bose-Einstein models.
        </p>
        <p>
          Team YorkU [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] explored BM25 and Divergence from Randomness methods
as provided by the Terrier toolkit, along with the associated relevance feedback
retrieval approaches.
4
4.1
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Findings</title>
      <sec id="sec-4-1">
        <title>Pooling and Coverage of Relevance Assessments</title>
        <p>A total of 8,713 documents were assessed. Of these, 6,741 (77.4%) were assessed
as irrelevant (0), 1,515 (17.4%) as somewhat relevant (1), 457 (5.2%) as highly
relevant (2). For readability assessments, the recorded distribution was: 1,145
(13.1%) documents assessed as di cult (0), 1,568 (18.0%) as somewhat di cult
(1), 2,769 (31.8%) as somewhat easy (2), and 3,231 (37.1%) as easy (3).</p>
        <p>Table 3 details the coverage of the relevance assessments with respect to the
participant submissions, averaged over the whole query set. While in theory all
runs 1-3 should have full coverage (100%), in practice a small portion of
documents included in the relevance assessment pool were left unjudged because the
documents were not in the collection (participants provided an invalid document
identi er) or the page failed to render in the relevance assessment toolkit (for
example because the page contained redirect scripts or other scripts that were
not executed within Relevation17). Overall, the mean coverage for runs 1-3 was
above 99%, with only run 3 from team KUCS being sensibly below this value.
This suggests that the retrieval e ectiveness for runs 1-3 can be reliably
measured. The coverage beyond submissions 3 is lower but always above 90% (and
the mean above 95%); this suggest that the evaluation of runs 4-10 in terms of
precision at 10 may be underestimated of an average maximum of 0.05 points
over the whole query set.</p>
        <p>Table 4 details the coverage of relevance assessment for the multilingual runs.
As mentioned in Section 2.8, due to limited relevance assessment availability, only
the English runs were considered when forming the pool for relevance assessment.
The coverage of these relevance assessments with respect to the top 10 documents
ranked by each participants' submissions is shown in the columns marked as Eng.
in Table 4. An additional document pool, made using only documents in
runs13 of multilingual submissions, was created to further increase the coverage of
multilingual submissions; the coverage of the union of the original assessments
and these additional ones (referred to as merged ) is shown in the columns marked
as Merged in Table 4 for the multilingual runs. The merged set of relevance
assessments was enough to provide a fairly high coverage for all runs, including
those not in the pool (i.e., runs beyond number 3), with a minimal coverage of
97%; this is likely because only one team submitted runs for the multilingual
challenge, thus providing only minimal variation in terms of top retrieval results.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Evaluation Results and Findings</title>
        <p>Table 5 reports the evaluation of the participants submissions and the organisers
baselines based on P@10 and nDCG@10 for English queries. The evaluation
based on RBP and the readability measures is reported in Table 6.</p>
        <p>Most of the approaches developed by team ECNU obtain signi cantly higher
values of P@10 and nDCG@10 compared to the other participants,
demonstrat17 Note that before the relevance assessment exercise started, we removed the majority
of scripts from the pooled pages to avoid this problem.
ing about 40% increase in e ectiveness in their best run compared to the
runnerup team (KISTI). The best submission developed by the organisers and based
on both relevance and readability estimates has been proved di cult to
outperform by most teams (only 4 out of 12 teams obtained higher e ectiveness).
The pooling methodology does not appear to have signi cantly in uenced the
evaluation of non-pooled submissions, as demonstrated by the fact that the best
runs of some teams are not those that were fully pooled (e.g. team KISTI, team
CUNI, team GRIUM).</p>
        <p>There are no large di erences between system rankings produced using P@10
or nDCG@10 as evaluation measure (Kendall = 0.88). This is unlike when
readability is also considered in the evaluation (the Kendall between system
rankings obtained with P@10 or uRBP is 0.76). In this latter case, while ECNU's
submissions are con rmed to be the most e ective, there are large variations in
system rankings when compared to those obtained considering relevance
judgements only. In particular, runs from team KISTI, which in the relevance-based
evaluation were ranked among the top 20 runs, are not performing as well when
considering also readability, with their top run (KISTI EN RUN.7) being ranked
only 37th according to uRBP.</p>
        <p>The following considerations could be drawn when comparing the di erent
methods employed by the participating teams. Query expansion is found to
often improve results. In particular, team ECNU obtained the highest e ectiveness
among the systems that took part in this task; this was achieved when query
expansion terms are mined from Google search results returned for the original
queries (ECNU EN Run.3). This approach indeed obtained higher e ectiveness
compared to learning-to-rank alternatives (ECNU EN Run.10). The results of
team UBML show that query expansion using the Bose-Einstein model 1 and
the local collection works better than other query expansion methods and
external collections. Team USST also found that query expansion was e ective
to improve results, however they found that the Bose-Einstein models did not
provide improvements over their baseline, while the Kullback-Liebler Divergence
based query expansion provided minor improvements. Health-speci c query
expansion methods based on the UMLS were shown to be e ective above common
baselines and other considered query expansion methods by Team LIMSI and
GRIUM (this form of query expansion was the only one that delivered higher
e ectiveness than their baseline).Team KISTI found that the combination of
concept-based document centrality (CBDC) and cluster- based external
expansion model (CBEEM) improved the results best. Few teams did not observe
improvements over their baselines; this was the case for teams KUCS, Miracl,
FDUSGInfo and HCMUS.</p>
        <p>Tables 7 and 8 report the evaluation of the multilingual submissions based
on P@10 and nDCG@10; results are reported with respect to both the original
qrels (obtained by sampling English runs only) and the additional qrels (obtained
by sampling also multilingual runs, but using a di erent set of assessors); see
Section 2.8 for details about the di erence between these relevance assessments.
Only one team (CUNI) participated in the multilingual task; they also submitted
to the English-based task and thus it is possible to discuss the e ectiveness of
their retrieval system when answering multilingual queries compared to that
achieved when answering English queries.</p>
        <p>The evaluation based on the original qrels allows us to compare multilingual
runs directly with English runs. Note that the original relevance assessments
exhibit a level of coverage for the multilingual runs that is similar to those obtained
for English submissions numbered 4-10. The evaluation based on the additional
qrels (merged) allows analysis of the multilingual runs using the same pooling
method used for English runs; thus submissions 1-3 for the multilingual runs
can be directly compared to the corresponding English ones, at the net of
differences in expertise, sensibility and systematic errors between the paid medical
assessors and the volunteer, student self-assessor used to gather judgements for
the multilingual runs.</p>
        <p>When only multilingual submissions are considered, it can be observed that
there is not a language in which CUNI's system is more e ective: e.g.
submissions that considered Italian queries are among the best performing with original
assessments and are the best performing with the additional assessments, but
di erences in e ectiveness among top runs for di erent languages are not
statistically signi cant. However, it can be observed that none of CUNI's submissions
that addressed queries expressed in not European languages (Farsi and Arabic)
are among the top ranked systems, regardless of the type of relevance
assessments.</p>
        <p>The use of the additional relevance assessments naturally translates in
observing increased retrieval e ectiveness across all multilingual runs (because some of
the documents in the top 10 ranks that were not assessed, and thus irrelevant, in
the original assessments may have been marked as relevant in the additional
assessments). However, a noteworthy observation is that the majority of the most
e ective runs according to the additional assessments are those that were not
fully sampled to form the relevance assessment pools (i.e. runs 4-10, as opposed
to the pooled runs 1-3).</p>
        <p>
          When the submissions of team CUNI are compared across English and
multilingual queries, it is possible to observe that the best multilingual runs do not
outperform English runs (unlike when the same comparison was instructed in
the 2014 task [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]), regardless of the type of relevance assessments. This result
does not come as unexpected and it indicates that the translation from a foreign
language to English as part of the retrieval process does degrade the quality of
queries (in terms of retrieval e ectiveness), suggesting that more work is needed
to bridge the gap in e ectiveness between English and multilingual queries when
these are used to retrieve English content.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This paper has described methods, results and analysis of the CLEF 2015 eHealth
Evaluation Lab, Task 2. The task considered the problem of retrieving web pages
for people seeking health information regarding unknown conditions or
symptoms. 12 teams participated in the task; the results have shown that query
expansion plays an important role in improving search e ectiveness. The best
results were achieved by a query expansion method that mined the top results
from the Google search engine. Despite the improvements over the organisers'
baselines achieved by some teams, further work is needed to sensibly improve
search in this context, as only about half of the top 10 results retrieved by the
best system were found to be relevant.</p>
      <p>As a by-product of this evaluation exercise, the task contributes to the
research community a collection with associated assessments and evaluation
framework (including readability evaluation) that can be used to evaluate the e
ectiveness of retrieval methods for health information seeking on the web. Queries,
assessments and participants runs are publicly available at http://github.com/
CLEFeHealth/CLEFeHealth2015Task2.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgement</title>
      <p>This task has been supported in part by the European Union Seventh
Framework Programme (FP7/2007-2013) under grant agreement no257528
(KHRESMOI), by Horizon 2020 program (H2020-ICT-2014-1) under grant agreement
no 644753 (KCONNECT), by the Austrian Science Fund (FWF) project no
I1094-N23 (MUCKE), and by the Czech Science Foundation (grant number
P103/12/G084). We acknowledge the time of the people involved in the
translation and relevance assessment tasks, in special we want to thank Dr. Johannes
Bernhardt-Melischnig (Medizinische Universitat Graz) for coordinating the
recruitment and management of the paid medical students that participated in
the relevance assessment exercise.</p>
      <p>R Run Name</p>
      <p>Run Name
1 CUNI DE Run10 0.2985
2 CUNI DE Run7 0.2970
3 CUNI FR Run10 0.2833
4 CUNI FR Run7 0.2773
5 CUNI IT Run10 0.2758
6 CUNI IT Run1 0.2652
7 CUNI IT Run4 0.2621
8 CUNI PT Run6 0.2530
9 CUNI PT Run8 0.2515
10 CUNI DE Run8 0.2500
10 CUNI FR Run9 0.2500
12 CUNI FR Run8 0.2455
12 CUNI IT Run6 0.2455
14 CUNI DE Run9 0.2409
14 CUNI PT Run10 0.2409
16 CUNI IT Run2 0.2394
17 CUNI IT Run3 0.2348
17 CUNI PT Run7 0.2348
19 CUNI IT Run8 0.2333
20 CUNI CS Run10 0.2303
20 CUNI FA Run10 0.2303
20 CUNI PT Run1 0.2303
20 CUNI PT Run5 0.2303
24 CUNI PT Run4 0.2288
25 CUNI AR Run10 0.2273
25 CUNI FA Run4 0.2273
25 CUNI IT Run9 0.2273
28 CUNI FA Run1 0.2258
29 CUNI CS Run7 0.2242
30 CUNI FA Run3 0.2227
30 CUNI FA Run5 0.2227
32 CUNI AR Run5 0.2197
32 CUNI AR Run6 0.2197
34 CUNI FA Run2 0.2182
34 CUNI FA Run8 0.2182
1 CUNI IT Run10 0.3727
1 CUNI IT Run4 0.3727
3 CUNI IT Run1 0.3712
4 CUNI FR Run10 0.3682
4 CUNI FR Run7 0.3682
6 CUNI IT Run6 0.3606
7 CUNI PT Run2 0.3576
8 CUNI DE Run10 0.3561
9 CUNI DE Run7 0.3545
10 CUNI IT Run8 0.3515
11 CUNI PT Run1 0.3500
12 CUNI PT Run4 0.3485
13 CUNI IT Run2 0.3424
14 CUNI IT Run3 0.3394
15 CUNI PT Run10 0.3379
16 CUNI PT Run3 0.3364
17 CUNI FA Run10 0.3333
17 CUNI FR Run9 0.3333
17 CUNI PT Run6 0.3333
20 CUNI CS Run1 0.3318
20 CUNI CS Run7 0.3318
20 CUNI IT Run5 0.3318
20 CUNI IT Run7 0.3318
20 CUNI PT Run8 0.3318
25 CUNI CS Run10 0.3288
26 CUNI FA Run4 0.3273
26 CUNI FR Run3 0.3273
28 CUNI FA Run1 0.3258
29 CUNI FA Run3 0.3242
30 CUNI FA Run2 0.3227
31 CUNI FR Run4 0.3182
31 CUNI FR Run8 0.3182
31 CUNI PT Run5 0.3182
34 CUNI FR Run1 0.3121
35 CUNI IT Run9 0.3106
0.2498
0.2539
0.2661
0.2259
0.2343
0.2206
0.2669
0.2672
0.2613
0.2556
0.2569
0.2544
0.2397
0.2255
0.2426
0.2392
0.2493
0.2446
0.2122
0.2504
0.2256
0.2070
0.2327
0.2255
0.2403
0.2058
0.2101
0.2039
0.2178
0.2237
0.2211
0.2261
0.1984
0.1954
0.1811</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the clef ehealth evaluation lab 2015</article-title>
          .
          <source>In: CLEF 2015 - 6th Conference and Labs of the Evaluation Forum, Lecture Notes in Computer Science (LNCS)</source>
          , Springer (
          <year>September 2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Health topics:
          <volume>80</volume>
          %
          <article-title>of internet users look for health information online</article-title>
          .
          <source>Pew Internet &amp; American Life Project</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H., Salantera,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2013</year>
          ,
          <article-title>Task 3: Information retrieval to address patients' questions when reading clinical reports</article-title>
          .
          <source>In: Online Working Notes of CLEF, CLEF</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gareth</surname>
            <given-names>J.F.</given-names>
          </string-name>
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>H.M.</given-names>
          </string-name>
          :
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2014</year>
          ,
          <article-title>Task 3: User-centred health information retrieval</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop: Online Working Notes, She eld,
          <source>UK</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>J.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liui</surname>
            ,
            <given-names>X.:</given-names>
          </string-name>
          <article-title>An investigation of the e ectiveness of concept-based approach in medical information retrieval grium@ clef2014ehealthtask 3</article-title>
          .
          <source>Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>James</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carterette</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , H.:
          <article-title>Using discharge summaries to improve information retrieval in clinical domain</article-title>
          .
          <source>Proceedings of the ShARe/-CLEF eHealth Evaluation Lab</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Benigeri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pluye</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Shortcomings of health information on the internet</article-title>
          .
          <source>Health promotion international 18(4)</source>
          (
          <year>2003</year>
          )
          <volume>381</volume>
          {
          <fpage>386</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horvitz</surname>
          </string-name>
          , E.:
          <article-title>Cyberchondria: studies of the escalation of medical concerns in web search</article-title>
          .
          <source>ACM TOIS 27(4)</source>
          (
          <year>2009</year>
          )
          <fpage>23</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
          </string-name>
          , J.:
          <article-title>Diagnose this if you can: On the e ectiveness of search engines in nding medical self-diagnosis information</article-title>
          .
          <source>In: Advances in Information Retrieval</source>
          . Springer (
          <year>2015</year>
          )
          <volume>562</volume>
          {
          <fpage>567</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H.:
          <article-title>Khresmoi { multimodal multilingual medical information search</article-title>
          . In:
          <article-title>MIE village of the future</article-title>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Stanton</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ieong</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mishra</surname>
          </string-name>
          , N.:
          <article-title>Circumlocution in diagnostic medical queries</article-title>
          .
          <source>In: Proceedings of the 37th international ACM SIGIR conference on Research &amp; development in information retrieval</source>
          ,
          <source>ACM</source>
          (
          <year>2014</year>
          )
          <volume>133</volume>
          {
          <fpage>142</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>Relevation!: an open source system for information retrieval relevance assessment</article-title>
          .
          <source>In: Proceedings of the 37th international ACM SIGIR conference on Research &amp; development in information retrieval</source>
          ,
          <source>ACM</source>
          (
          <year>2014</year>
          )
          <volume>1243</volume>
          {
          <fpage>1244</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Walsh</surname>
            ,
            <given-names>T.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volsko</surname>
            ,
            <given-names>T.A.</given-names>
          </string-name>
          :
          <article-title>Readability assessment of internet-based consumer health information</article-title>
          .
          <source>Respiratory</source>
          care
          <volume>53</volume>
          (
          <issue>10</issue>
          ) (
          <year>2008</year>
          )
          <volume>1310</volume>
          {
          <fpage>1315</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Integrating understandability in the evaluation of consumer health search engines</article-title>
          .
          <source>In: Medical Information Retrieval Workshop at SIGIR</source>
          <year>2014</year>
          .
          <article-title>(</article-title>
          <year>2014</year>
          )
          <fpage>32</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macdonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plachouras</surname>
          </string-name>
          , V.:
          <article-title>Research directions in terrier</article-title>
          .
          <source>Novatica UPGRADE Special Issue on Web Information Access</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.:</given-names>
          </string-name>
          <article-title>The in uence of pre-processing on the estimation of readability of web documents</article-title>
          .
          <source>In: Proceedings of the 24th ACM International Conference on Conference on Information and Knowledge Management (CIKM)</source>
          .
          <article-title>(</article-title>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Kincaid</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fishburne</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chissom</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Derivation of New Readability Formulas for Navy Enlisted Personnel</article-title>
          .
          <source>Technical report</source>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kincaid</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fishburne</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chissom</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Derivation of New Readability Formulas for Navy Enlisted Personnel</article-title>
          .
          <source>Technical report</source>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Jarvelin,
          <string-name>
            <surname>K.</surname>
          </string-name>
          , Kekalainen, J.:
          <article-title>Cumulated gain-based evaluation of IR techniques</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          <volume>20</volume>
          (
          <issue>4</issue>
          ) (
          <year>2002</year>
          )
          <volume>422</volume>
          {
          <fpage>446</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. Mo at, A.,
          <string-name>
            <surname>Zobel</surname>
          </string-name>
          , J.:
          <article-title>Rank-biased precision for measurement of retrieval e ectiveness</article-title>
          .
          <source>ACM Transactions on Information Systems (TOIS) 27(1)</source>
          (
          <year>2008</year>
          )
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Y.:
          <article-title>On the distribution of user persistence for rank-biased precision</article-title>
          .
          <source>In: Proceedings of the 12th Australasian document computing symposium</source>
          . (
          <year>2007</year>
          )
          <volume>17</volume>
          {
          <fpage>24</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Saleh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bibyna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>CUNI at the CLEF 2015 eHealth Lab Task 2</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>ECNU at 2015 eHealth Task 2: User-centred Health Information Retrieval</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>received</surname>
          </string-name>
          , N.:
          <string-name>
            <surname>Missing</surname>
          </string-name>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nie</surname>
          </string-name>
          , J.Y.:
          <article-title>Bridging Layperson's Queries with Medical Concepts - GRIUM@CLEF2015 eHealth Task 2</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>H.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jung</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
          </string-name>
          , K.Y.:
          <article-title>KISTI at CLEF eHealth 2015 Task 2</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Thesprasith</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaruskulchai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Task 2a: Team KU-CS: Query Coherence Analysis for PRF and Genomics Expansion</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>D'hondt</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Grau</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : LIMSI @
          <article-title>CLEF eHealth 2015 - task 2</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Ksentini</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tmar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boughanem</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gargouri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <source>Miracl at Clef</source>
          <year>2015</year>
          :
          <article-title>UserCentred Health Information Retrieval Task</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Huynh</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
          </string-name>
          , T.T.,
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          :
          <article-title>TeamHCMUS: A Concept-Based Information Retrieval Approach for Web Medical Documents</article-title>
          .
          <source>In: Proceedings of the ShARe/- CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Thuma</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mosweunyane</surname>
          </string-name>
          , G.:
          <article-title>UBML participation to CLEF eHealth IR challenge 2015: Task 2</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Employing query expansion models to help patients diagnose themselves</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Ghoddousi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>J.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          : York University at CLEF eHealth 2015:
          <article-title>Medical Document Retrieval</article-title>
          .
          <source>In: Proceedings of the ShARe/CLEF eHealth Evaluation Lab</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>