<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The IR Task at the CLEF eHealth Evaluation Lab 2016: User-centred Health Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guido Zuccon</string-name>
          <email>g.zuccon@qut.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joao Palotti</string-name>
          <email>palotti@ifs.tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liadh Kelly</string-name>
          <email>liadh.kelly@tcd.ie</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihai Lupu</string-name>
          <email>lupu@ifs.tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavel Pecina</string-name>
          <email>pecina@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Mu¨ller</string-name>
          <email>henning.mueller@hevs.ch</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julie Budaher</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anthony Deacon</string-name>
          <email>aj.deacon@qut.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Queensland University of Technology</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Trinity College Dublin</institution>
          ,
          <addr-line>Irland</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universit ́e Grenoble Alpes</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Applied Sciences Western Switzerland</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Vienna University of Technology</institution>
          ,
          <addr-line>Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper details the collection, systems and evaluation methods used in the IR Task of the CLEF 2016 eHealth Evaluation Lab. This task investigates the e↵ectiveness of web search engines in providing access to medical information for common people that have no or little medical knowledge. The task aims to foster advances in the development of search technologies for consumer health search by providing resources and evaluation methods to test and validate search systems. The problem considered in this year's task was to retrieve web pages to support the information needs of health consumers that are faced by a medical condition and that want to seek relevant health information online through a search engine. As part of the evaluation exercise, we gathered 300 queries users posed with respect to 50 search task scenarios. The scenarios were developed from real cases of people seeking health information through posting requests of help on a web forum. The presence of query variations for a single scenario helped us capturing the variable quality at which queries are posed. Queries were created in English and then translated into other languages. A total of 49 runs by 10 di↵erent teams were submitted for the English query topics; 2 teams submitted 29 runs for the multilingual topics.</p>
      </abstract>
      <kwd-group>
        <kwd>Evaluation</kwd>
        <kwd>Health Search</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This document reports on the CLEF 2016 eHealth Evaluation Lab, IR Task
(task 3). The task investigated the problem of retrieving web pages to support
information needs of health consumers (including their next-of-kin) that are
confronted with a health problem or medical condition and that use a search engine
to seek better understanding about their health. This task has been developed
within the CLEF 2016 eHealth Evaluation Lab, which aims to foster the
development of approaches to support patients, their next-of-kin, and clinical sta↵ in
understanding, accessing and authoring health information [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        The use of the Web as source of health-related information is a wide-spread
practice among health consumers [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and search engines are commonly used as a
means to access health information available online [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Previous iterations of this
task (i.e. the 2013 and 2014 CLEFeHealth Lab Task 3 [
        <xref ref-type="bibr" rid="ref8 ref9">8,9</xref>
        ]) aimed at evaluating
the e↵ectiveness of search engines to support people when searching for
information about their conditions, e.g. to answer queries like “thrombocytopenia
treatment corticosteroids length”. These two evaluation exercises have provided
valuable resources and an evaluation framework for developing and testing new
and existing techniques. The fundamental contribution of these tasks to the
improvement of search engine technology aimed at answering this type of health
information need is demonstrated by the improvements in retrieval e↵ectiveness
provided by the best 2014 system [27] over the best 2013 system [30] (using
di↵erent, but comparable, topic sets). The 2015 task has instead focused on
supporting consumers searching for self-diagnosis information [23], an important type
of health information seeking activity [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This year’s task expands on the 2015
task, by considering not only self-diagnosis information needs, but also needs
related to treatment and management of health conditions. Previous research has
shown that exposing people with no or scarce medical knowledge to complex
medical language may lead to erroneous self-diagnosis and self-treatment and
that access to medical information on the Web can lead to the escalation of
concerns about common symptoms (e.g., cyberchondria) [
        <xref ref-type="bibr" rid="ref3">3,29</xref>
        ]. Research has also
shown that current commercial search engines are still far from being e↵ective
in answering such unclear and underspecified queries [33].
      </p>
      <p>The remainder of this paper is structured as follows: Section 2 details the
sub-tasks we considered this year; Section 3 described the query set and the
methodology used to create it; Section 5 details the baselines created by the
organisers as a benchmark for participants; Section 6 describes participants
submissions; Section 7 details the methods used to create the assessment pools and
relevance criteria; Section 8 lists the evaluation metrics used for this Task; finally,
Section 9 concludes this overview paper.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Tasks</title>
      <sec id="sec-2-1">
        <title>Sub-Task 1: Ad-hoc Search</title>
        <p>Queries for this task are generated by mining health web forums to identify
example information needs, as detailed in section 3. Every query is treated as
independent and participants are asked to generate retrieval runs in answer to
such queries, as in a common ad-hoc search task. This task extends the evaluation
framework used in 2015 (which considered, along with topical relevance, also
the readability of the retrieved documents) to consider further dimensions of
relevance such as the reliability of the retrieved information.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Sub-Task 2: Query Variations</title>
        <p>
          This task explores query variations for each single information need. Previous
research has shown that di↵erent users tend to issue di↵erent queries for the
same information need and that the use of query variations for evaluation of
IR systems leads to as much variability as system variations [
          <xref ref-type="bibr" rid="ref1 ref2">1,2,23</xref>
          ]. This was
the case also in this year’s task. Note that we explored query variations also in
the 2015 IR task [23], and we found that for the same image showing a health
condition, di↵erent query creators issued very di↵erent queries: they di↵er not
only in terms of the keywords contained in the query, but also with respect to
their retrieval e↵ectiveness.
        </p>
        <p>Di↵erent query variations are generated for the same information need
(extracted from a web forum entry, as explained in section 3), thus capturing the
variability intrinsic in how people search when they have the same information
need. Participants were asked to exploit query variations when building their
systems: participants were told which queries related to the same information
need and they were required to produce one set of results to be used as answer
for all query variations of an information need. This task aims to foster research
into building systems that are robust to query variations, for example, through
considering the fusion of ranked lists produced in answer to each single query
variation.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Sub-Task 3: Multilingual Ad-hoc Search</title>
        <p>The multilingual task extends the Ad-hoc Search task by providing a
translation of the queries from English into Czech, French, Hungarian, German, Polish,
Spanish and Swedish. The goal of this sub-task is to support research in
multilingual information retrieval, developing techniques to support users that can
express their information need well in their native language and can read the
results in English.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Query Set</title>
      <p>We considered real health information needs expressed by the general public
through posts published in public health web forums. Forum posts were
extracted from the AskDocs section of Reddit7. This section allows users to post a
description of a medical case or ask a medical question seeking medical
information such as diagnosis, or details regarding treatments. Users can also interact
through comments. We selected posts that were descriptive, clear and
understandable. Posts with information regarding the author or patient (in case the
7 https://www.reddit.com/r/AskDocs/
post author sought help for another person), such as demographics (age, gender),
medical history and current medical condition, were preferred.</p>
      <p>In order to collect query variants that could be compared, we also selected
posts where a main and single information need could be identified. These
constraints guarantee as much as possible getting queries on the same aspects of
the post.</p>
      <p>The comments were also taken into account in the selection. Any user can
add a comment to a post, and all users are labeled according to their medical
expertise8. We mainly selected posts with comments, including some from
labeled users. The post were manually selected by a student, and a total of 50
posts were used for query creation.</p>
      <p>Each of the selected forum posts were presented to 6 query creators with
di↵erent medical expertise: these included 3 medical experts (final year medical
students undertaking rotations in hospitals) and 3 lay users with no prior medical
knowledge.</p>
      <p>All queries were preprocessed to correct for spelling mistakes; this was done
using the Linux program aspell. This was however manually supervised so that
spelling correction was performed only when appropriate, as for example not to
change drug names. We explicitly did not remove punctuation marks from the
queries, e.g., participants could take advantage of the quotation marks used by
the query creators to indicate proximity terms or other features.</p>
      <p>A total of 300 queries were created. Queries were numbered using the
following convention: the first 3 digits of a query id identify a post number (information
need), while the last 3 digits of a query id identify each individual query creator.
Expert query creators used the identifier 1, 2 and 3 and laypeople query creator
used the identifiers 4, 5 and 6. In Figure 2 we show variants 1, 2 (both generated
by laypeople) and 4 (generated by an expert) created for post number 103 (posts
started from number 101), shown in Figure 1.</p>
      <p>For the query variations element of the task (sub-task 2), participants were
told which queries were related to the same information need, to allow them to
produce one set of results to be used as answer for all query variations of an
information need.</p>
      <p>For the multilingual element of the challenge (sub-task 3), Czech, French,
Hungarian, German, Polish, Spanish and Swedish translations of the queries
were provided. Queries were translated by medical experts hired through a
professional translation company.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Dataset</title>
      <p>
        Previous IR tasks in the CLEF eHealth Lab have used the Khresmoi
collection [
        <xref ref-type="bibr" rid="ref10 ref12">12,10</xref>
        ], a collection of about 1 million health web pages. This year we set
a new challenge to the participants by using the ClueWeb12-B139, a collection
8 To be labeled as a medical expert, users have to send Reddit a proof such as a
student ID, or a diploma.
9 http://lemurproject.org/clueweb12/
of more than 52 million web pages. As opposed to the Khresmoi collection, the
crawl in ClueWeb12-B13 is not limited to certified Health On the Net websites
and known health portals, but it is a higher-fidelity representation of a common
Internet crawl, making the dataset more in line with the content current web
search engines index and retrieve.
      </p>
      <p>
        For participants who did not have access to the ClueWeb dataset, Carnegie
Mellon University granted the organisers permission to make the dataset
available through cloud computing instances10 provided by Microsoft Azure. The
Azure instances that were made available to participants for the IR challenge
included (1) the Clueweb12-B13 dataset, (2) standard indexes built with the
Terrier11 [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and the Indri12 [28] toolkits, (3) additional resources such as a spam
list [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Page Rank scores, anchor texts [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], urls, etc. made available through
the ClueWeb12 website.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Baselines</title>
      <p>We generated 55 runs, from which 19 were for Sub-Task 1 and 36 for Sub-Task
2, based on common baseline models and simple approaches for fusing query
variations. In this section we describe the baseline runs.
10 The organisers are thankful to Carnegie Mellon University, and in particular to
Jamie Callan and Christina Melucci, for their support in obtaining the permission
to redistribute ClueWeb 12. The organisers are also thankful to Microsoft Azure
who provided the Azure cloud computing infrastructure that was made available to
participants through the Microsoft Azure for Research Award CRM:0518649.
11 http://terrier.org/
12 http://www.lemurproject.org/indri.php
&lt;queries&gt;
...
&lt;query&gt;
&lt;id&gt; 103001 &lt;/id&gt;
&lt;title&gt;headaches relieved by blood donation&lt;/title&gt;
&lt;/query&gt;
&lt;query&gt;
&lt;id&gt; 103002 &lt;/id&gt;
&lt;title&gt;high iron headache&lt;/title&gt;
&lt;/query&gt;
...
&lt;query&gt;
&lt;id&gt; 103004 &lt;/id&gt;
&lt;title&gt;headaches caused by too much blood or</p>
      <p>"high blood pressure"&lt;/title&gt;
&lt;/query&gt;
...
&lt;/queries&gt;</p>
      <p>For both systems, we created the runs using and not using the default
pseudorelevance feedback (PRF) of each toolkit. When using PRF, we added to the
original query the top 10 terms of the top 3 documents. All these baseline runs
were created using the Terrier and Indri instances made available to participants
in the Azure platform.</p>
      <p>Additionally, we created a set of baseline runs that take into account the
reliability and understandability of information.</p>
      <p>
        Five reliability baselines were created based on the Spam rankings distributed
with ClueWeb1213[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. For a given run, we removed all documents that had a
spam score smaller than a given threshold th. We used the BM25 baseline run
of Terrier, and 5 di↵erent values for th (50, 60, 70, 80 and 90).
      </p>
      <p>
        Two understandability baselines were created using readability formulae. We
created runs based on CLI (Coleman-Liau Index) and GFI (Gunning Fox Index)
scores [
        <xref ref-type="bibr" rid="ref11 ref4">4,11</xref>
        ], which are a proxy for the number of years of the school required
to read the text being evaluated. These two readability formulae were chosen
13 http://www.mansci.uwaterloo.ca/~msmucker/cw12spam/
because they showed to be robust across di↵erent methods for HTML
preprocessing [24]. We followed one of the methods suggested in [24], in which the
HTML documents are preprocessed using Justext14[26], the main text is
extracted, periods at the end of sentences are added whenever they are necessary
(e.g., in presence of line breaks), and then readability scores are calculated. Given
the initial score S for a document and its readability score R, the final score for
each document is the combination of score obtained as S ⇥ 1.0/R.
5.2
      </p>
      <sec id="sec-5-1">
        <title>Baselines for Sub-Task 2</title>
        <p>
          We explored three ways to combine query variations:
– Concatenation: we concatenated the text of each query variation into a single
query.
– Reciprocal Ranking Fusion [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]: we fuse the ranks of each query variations
using the reciprocal ranking fusion approach, i.e.,
        </p>
        <p>RRF Score(d) = X
r2 R</p>
        <p>
          1
k + r(d)
,
where D is set the documents to be ranked, R is the set of document rankings
retrieved for each query variation by the same retrieval model, r(d) is the
rank of document d, and k is a constant set to 60, as in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
– Rank Biased Precision Fusion: similarly to the reciprocal ranking fusion,
we fuse the documents retrieved for each query variation with the Ranking
Biased Precision (RBP) formula [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ],
        </p>
        <p>RBP Score(d) =</p>
        <p>p) ⇥ (p)r(d) 1,
X(1
r2 R
where p is the free parameter of the RBP model used to estimate user
persistence. Here we set p = 0.80.</p>
        <p>For each of the three methods described above, we created a run based on
each of three baselines for Terrier and Indri, with and without pseudo-relevance
feedback. A total of 36 baseline runs were created for sub-task 2 (combination
of 3 fusion approaches, 2 toolkits, 3 models, with and without PRF).
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Participant Submissions</title>
      <p>The number of registered participants for CLEF eHealth IR Task was 58; of
these, 10 submitted at least one run for any of the sub-tasks, as shown in Table 1.
Each team could submit up to 3 runs for Sub-Tasks 1 and 2, and up to 3 runs
for each language of Sub-Task 3.</p>
      <p>We include below a summary of the approach of each team, self described by
them when their runs were submitted.
14 https://pypi.python.org/pypi/jusText
Team Name University
CUNI: The CUNI team participated in all the subtasks but their main focus
was put on the multilingual search in sub-task 3. The monolingual runs in
sub-tasks 1 and 2 are mainly intended for comparison with the multilingual
runs in sub-task 3. In sub-task 1 (Ad-Hoc Search), Run 1 employs the Terrier
implementation of Dirichlet-smoothed language model with the µ parameter
tuned on the data from previous CLEF eHealth tasks; Run 2 uses the Terrier
vector space TF/IDF model. In sub-task 2 (Query variants), all query variants
of one information need are searched for by the retrieval system
(Dirichletsmoothed language model in Run 1 and vector space TF/IDF model in Run2)
and the resulting lists of ranked documents are merged and reranked by
document scores to produce one ranked list of documents for each information
need. In sub-task 3 (Multilingual search), all the non-English queries
(including variants) are translated into English using their own statistical machine
translation systems adapted to translate search queries in the medical domain.
For each non-English query, 15 translation variants (hypotheses) are obtained.
Their multilingual Run 1 employs the single best translation for each query
as provided by the translation systems. In their multilingual Run 2, the top
15 translation hypotheses are reranked using a discriminative regression model
employing a) features provided by the translation system and b) various kinds
of features extracted from the document collection, external resources (UMLS
Unified Medical Language System, Wikipedia), or the translations themselves.
Run 3 employs the same reranking method applied to the translation system
features only.</p>
      <p>ECNU: The ECNU team proposes a Web-based query expansion model and
a combination method to better understand and satisfy the task. They use as
baseline the Terrier implementation of the BM25 model. The other runs for
subtask 1, 2 and 3 explore Google search and MeSH to do query expansion. BM25,
DFR BM25, BB2, the PL2 models of Terrier and TF IDF, the BM25 models
of Indri were used and combined. For the sub-task 3 runs, Google Translator
was used to translate the queries from Czech, French, Polish and Swedish to
English before applying the same methods of sub-task 1.</p>
      <p>GUIR: GUIR studies the use of medical terms for query reformulation.
Synonyms and hypernyms from UMLS are used to generate reformulations of
the queries; Terrier with Divergence from Randomness is used for retrieving
and scoring documents. For sub-task 1, results obtained from the reformulated
queries are used with the Borda rank aggregation algorithm. For sub-task 2,
for each topic, results obtained are merged by any reformulated query in the
topic using the Borda rank aggregation algorithm.</p>
      <p>Infolab: Team InfoLab analyses the performance of several query expansion
strategies using di↵erent methods to select the terms to be added to the original
query. One of the methods uses the similarity between Wikipedia articles, found
through an analysis of incoming and outgoing links, for term selection. The
other method applies the Latent Dirichlet Allocation to Wikipedia articles to
extract topics each containing a set of words that are used for term selection.
In the end, readability metrics were used to re-rank the documents retrieved
using the expanded queries.</p>
      <p>KDEIR: KDEIR submitted runs for sub-tasks 1 and 2. In both sub-tasks the
Waterloo spam score was used to filter out the spammiest documents, and the
link structure present in the remaining documents was explored on top of their
language model baseline.</p>
      <p>KISTI: KISTI attempts two approaches using word vectors learned by Word2Vec
based on medical Wikipedia pages. At first, initial documents are obtained
using a search engine. Based on the documents, pseudo-relevance feedback (PRF)
is applied with two di↵erent usage of the word vectors. In the first approach,
PRF is performed with new relevance scores using the word vectors, while it
is performed with a new query expanded using the word vectors in the second
approach.</p>
      <p>MayoNLPTeam: Mayo explores a Part-of-Speech (POS) based query term
weighting approach which assigns di↵erent weights to the query terms according
to their POS categories. The weights are learned by defining an objective
function based on the mean average precision. They apply the proposed approach
with the optimal weights obtained from the TREC 2011 and 2012 Medical
Records Tracks into the Query Likelihood model (Run 2) and Markov Random
Field (MRF) models (Run 3). The conventional Query Likelihood model was
implemented as the baseline (Run 1).</p>
      <p>MRIM: MRIM’s objective is to investigate the e↵ectiveness of the word
embedding for query expansion on consumer health search, as well as the e↵ect of
the learning resource for learning on the results. Their system uses the Terrier
index provided by the organizers. As a retrieval model the Dirichlet language
model is used with default settings. Query expansion is applied on two training
sets using word embedding sources. Word2vec is used for word embedding.
ub-botswana: In this participation, the e↵ectiveness of three retrieval
strategies is evaluated. In particular, PL2 is deployed with a Boolean Fallback score
modifier as baseline system. If any of the retrieved documents contains all
undecorated query terms (i.e. query terms without any operators), then
documents are removed from the result set that do not contain all undecorated
query terms With this score modifier. Otherwise, nothing is done. In another
approach, the collection enrichment approach is employed, where the original
query is expanded with additional terms from an external collection (collection
not being searched). To deliver an e↵ective ranking, the first two rankers are
combined using data fusion techniques.</p>
      <p>WHUIRgroup: WHUIR uses Indri to conduct the experiments. CHV is used
to expand queries and propose a learning-to-rank algorithm to re-rank the
result.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Assessments</title>
      <p>
        A Pool of 25,000 documents was created using the RBP-based Method A
(Summing contributions) by Mo↵at et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], in which documents are weighted
according to their overall contribution to the e↵ectiveness evaluation as provided
by the RBP formula (with p=0.8, following Park and Zhang [25]). This stategy
was chosen because it was shown that it should be preferred over traditional
fixed-depth or stratified pooling when deciding upon the pooling strategy to be
used to evaluate systems under fixed assessment budget constraints [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. A total
of 100 runs were used (all baselines + all participant runs for Sub-Tasks 1 and
2) to form the assessment pools.
      </p>
      <p>Assessment was performed by paid final year medical students who had access
to queries, documents, and relevance criteria drafted by a junior medical doctor.
The relevance criteria were drafted considering the entirety of the forum posts
used to create the queries, a link to the forum posts was also provided to the
assessors.</p>
      <p>
        Relevance assessments were provided with respect to the grades Highly
relevant, Somewhat relevant and Not Relevant. Readability/understandability and
reliability/trustworthiness judgments were also collected for the documents in
the assessment pool. These judgements were collected using a integer value
between 0 and 100 (lower values meant harder to understand document / low
reliability) provided by judges through a slider tool; these judgements were used
to evaluate systems across di↵erent dimensions of relevance [32,31]. All
assessments were collected through a purpously costumised version of the Relevation
toolkit [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
8
      </p>
    </sec>
    <sec id="sec-8">
      <title>Evaluation Metrics</title>
      <p>
        System evaluation was conducted using precision at 10 (p@10) and normalised
discounted cumulative gain [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] at 10 (nDCG@10) as the primary and secondary
measures, respectively. Precision was computed using the binary relevance
assessments by collapsing Highly relevant and Somewhat relevant assessments into
the Relevant class, while nDCG was computed using the graded relevance
assessments.
      </p>
      <p>
        A separate evaluation was conducted using the multidimensional relevance
assessments (topical relevance, understandability and trustworthiness)
following the methods in [31]. For all runs, Rank biased precision (RBP) [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], with a
persistence parameter p = 0.80 (see [25]), was computed along with the
multidimensional modifications of RBP, namely uRBP (using binary understandability
assessments), uRBPgr (using graded understandability assessments), u+tRBP
(using binary understandability and trustworthiness assessments).
      </p>
      <p>Precision and nDCG were computed using trec eval15 along with RBP,
while the multidimensional evaluation was performed using ubire16 [31].
9</p>
    </sec>
    <sec id="sec-9">
      <title>Conclusions</title>
      <p>This paper describes methods, results and analysis of the CLEF 2016 eHealth
Evaluation Lab, IR Task. The task considers the problem of retrieving web pages
for people seeking health information regarding medical conditions, treatments
and suggestions. The task was divided into 3 sub-tasks including ad-hoc search,
query variations, and multilingual ad-hoc search. Ten teams participated in the
task; relevance assessment is underway and assessments along with the
participants results will be released at the CLEF 2016 conference (and will be available
at the task’s GitHub repository).</p>
      <p>
        As a by-product of this evaluation exercise, the task makes available to the
research community a collection with associated assessments and evaluation
framework (including readability and reliability evaluation) that can be used
to evaluate the e↵ectiveness of retrieval methods for health information seeking
on the web (e.g. [
        <xref ref-type="bibr" rid="ref21">21,22</xref>
        ]).
      </p>
      <p>Baseline runs, participant runs and results, assessments, topics and query
variations are available online at the GitHub repository for this Task: https:
//github.com/CLEFeHealth/CLEFeHealth2016Task3.
10</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>This work has received funding from the European Union’s Horizon 2020
research and innovation programme under grant agreement No 644753
(KConnect), and from the Austrian Science Fund (FWF) projects P25905-N23
(ADmIRE) and I1094-N23 (MUCKE). We also would like to thank Microsoft Azure
grant (CRM:0518649), ESF for the support for financial relevance assessments
and query creation, and the many assessor for their hard work.
15 http://trec.nist.gov/trec_eval/trec_eval_latest.tar.gz
16 https://github.com/ielab/ubire
22. J. Palotti, G. Zuccon, J. Bernhardt, A. Hanbury, and L. Goeuriot. Assessors
Agreement: A Case Study across Assessor Type, Payment Levels, Query Variations and
Relevance Dimensions. In Experimental IR Meets Multilinguality, Multimodality,
and Interaction: 7th International Conference of the CLEF Association, CLEF’16
Proceedings. Springer International Publishing, 2016.
23. J. Palotti, G. Zuccon, L. Goeuriot, L. Kelly, A. Hanburyn, G. J. Jones, M. Lupu,
and P. Pecina. CLEF eHealth Evaluation Lab 2015, Task 2: Retrieving Information
about Medical Symptoms. In CLEF 2015 Online Working Notes. CEUR-WS, 2015.
24. J. Palotti, G. Zuccon, and A. Hanbury. The influence of pre-processing on the
estimation of readability of web documents. In Proceedings of the 24th ACM
International on Conference on Information and Knowledge Management, CIKM ’15,
pages 1763–1766, New York, NY, USA, 2015. ACM.
25. L. Park and Y. Zhang. On the distribution of user persistence for rank-biased
precision. In Proceedings of the 12th Australasian document computing symposium,
pages 17–24, 2007.
26. J. Pomik´alek. Removing boilerplate and duplicate content from web corpora. PhD
thesis, Masaryk university, Faculty of informatics, Brno, Czech Republic, 2011.
27. W. Shen, J.-Y. Nie, X. Liu, and X. Liui. An investigation of the
e↵ectiveness of concept-based approach in medical information retrieval GRIUM@
CLEF2014eHealthTask 3. In Proceedings of the CLEF eHealth Evaluation Lab,
2014.
28. T. Strohman, D. Metzler, H. Turtle, and W. B. Croft. Indri: A language
modelbased search engine for complex queries. In Proceedings of the International
Conference on Intelligent Analysis, volume 2, pages 2–6. Citeseer, 2005.
29. R. W. White and E. Horvitz. Cyberchondria: studies of the escalation of medical
concerns in web search. ACM TOIS, 27(4):23, 2009.
30. D. Zhu, S. T.-I. Wu, J. J. Masanz, B. Carterette, and H. Liu. Using discharge
summaries to improve information retrieval in clinical domain. In Proceedings of
the CLEF eHealth Evaluation Lab, 2013.
31. G. Zuccon. Understandability biased evaluation for information retrieval. In
Advances in Information Retrieval, pages 280–292, 2016.
32. G. Zuccon and B. Koopman. Integrating understandability in the evaluation of
consumer health search engines. In Medical Information Retrieval Workshop at
SIGIR 2014, page 32, 2014.
33. G. Zuccon, B. Koopman, and J. Palotti. Diagnose this if you can: On the
e↵ectiveness of search engines in finding medical self-diagnosis information. In Advances
in Information Retrieval, pages 562–567. Springer, 2015.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          .
          <article-title>Query Side Evaluation: An Empirical Analysis of E↵ectiveness and E↵ort</article-title>
          .
          <source>In Proc. of SIGIR</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>P.</given-names>
            <surname>Bailey</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Mo↵at,
          <string-name>
            <given-names>F.</given-names>
            <surname>Scholer</surname>
          </string-name>
          , and
          <string-name>
            <surname>P. Thomas.</surname>
          </string-name>
          <article-title>User Variability and IR System Evaluation</article-title>
          .
          <source>In Proc. of SIGIR</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Benigeri</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Pluye</surname>
          </string-name>
          .
          <article-title>Shortcomings of health information on the internet</article-title>
          .
          <source>Health promotion international</source>
          ,
          <volume>18</volume>
          (
          <issue>4</issue>
          ):
          <fpage>381</fpage>
          -
          <lpage>386</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Coleman</surname>
          </string-name>
          and
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Liau</surname>
          </string-name>
          .
          <article-title>A computer readability formula designed for machine scoring</article-title>
          .
          <source>Journal of Applied Psychology</source>
          ,
          <volume>60</volume>
          :
          <fpage>283</fpage>
          -
          <lpage>284</lpage>
          ,
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Cormack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L. A.</given-names>
            <surname>Clarke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Buettcher</surname>
          </string-name>
          .
          <article-title>Reciprocal rank fusion outperforms condorcet and individual rank learning methods</article-title>
          .
          <source>In Proceedings of the 32Nd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '09</source>
          , pages
          <fpage>758</fpage>
          -
          <lpage>759</lpage>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Cormack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Smucker</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Clarke</surname>
          </string-name>
          .
          <article-title>Ecient and e↵ective spam filtering and re-ranking for large web datasets</article-title>
          .
          <source>Information retrieval</source>
          ,
          <volume>14</volume>
          (
          <issue>5</issue>
          ):
          <fpage>441</fpage>
          -
          <lpage>465</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S.</given-names>
            <surname>Fox</surname>
          </string-name>
          . Health topics:
          <volume>80</volume>
          %
          <article-title>of internet users look for health information online</article-title>
          .
          <source>Pew Internet &amp; American Life Project</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leveling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          , H. Mu¨ller, S. Salantera,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon.</surname>
          </string-name>
          <article-title>Share/clef ehealth evaluation lab 2013, task 3: Information retrieval to address patients' questions when reading clinical reports</article-title>
          .
          <source>CLEF 2013 Online Working Notes</source>
          ,
          <volume>8138</volume>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pecina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H. M. Gareth J.F.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2014</year>
          ,
          <article-title>Task 3: User-centred health information retrieval</article-title>
          .
          <source>In CLEF 2014 Evaluation Labs and Workshop:</source>
          Online Working Notes, Sheeld, UK,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Kelly</surname>
            , G. Zuccon, and
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Palotti</surname>
          </string-name>
          .
          <article-title>Building evaluation datasets for consumer-oriented information retrieval</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), Paris, France, may
          <year>2016</year>
          .
          <article-title>European Language Resources Association (ELRA).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>R.</given-names>
            <surname>Gunning</surname>
          </string-name>
          .
          <article-title>The Technique of Clear Writing</article-title>
          .
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          ,
          <year>1952</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          .
          <article-title>Medical information retrieval: an instance of domain-specific search</article-title>
          .
          <source>In Proceedings of SIGIR 2012</source>
          , pages
          <fpage>1191</fpage>
          -
          <lpage>1192</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          and
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Hau↵</article-title>
          . Mirex:
          <article-title>Mapreduce information retrieval experiments</article-title>
          .
          <source>arXiv preprint arXiv:1004.4489</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>K. Ja</surname>
          </string-name>
          <article-title>¨rvelin and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Keka</surname>
          </string-name>
          <article-title>¨la¨inen. Cumulated gain-based evaluation of IR techniques</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ):
          <fpage>422</fpage>
          -
          <lpage>446</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. L.
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Palotti</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          .
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2016</article-title>
          . In Information Access Evaluation. Multilinguality, Multimodality, and Visualization. Springer Berlin Heidelberg,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon.</surname>
          </string-name>
          <article-title>Relevation! an open source system for information retrieval relevance assessment</article-title>
          .
          <source>arXiv preprint</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>A.</given-names>
            <surname>Lipani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mihai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          .
          <article-title>The impact of fixed-cost pooling strategies on test collection bias</article-title>
          .
          <source>In Proceedings of the 2016 International Conference on The Theory of Information Retrieval</source>
          , ICTIR '
          <fpage>16</fpage>
          , New York, NY, USA,
          <year>2016</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>C. Macdonald</surname>
            ,
            <given-names>R. McCreadie</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Ounis.</surname>
          </string-name>
          <article-title>From puppy to maturity: Experiences in developing terrier</article-title>
          .
          <source>Proc. of OSIR at SIGIR</source>
          , pages
          <fpage>60</fpage>
          -
          <lpage>63</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>D. McDaid</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Park</surname>
          </string-name>
          .
          <article-title>Online health: Untangling the web. evidence from the bupa health pulse 2010 international healthcare survey</article-title>
          .
          <source>Technical report</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. A. Mo↵at and J.
          <string-name>
            <surname>Zobel</surname>
          </string-name>
          .
          <article-title>Rank-biased precision for measurement of retrieval e↵ectiveness</article-title>
          .
          <source>ACM Trans. Inf</source>
          . Syst.,
          <volume>27</volume>
          (
          <issue>1</issue>
          ):2:
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          :
          <fpage>27</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>J. Palotti</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <article-title>Zuccon, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          .
          <article-title>Ranking health web pages with relevance and understandability</article-title>
          .
          <source>In Proceedings of the 39th international ACM SIGIR conference on Research and development in information retrieval</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>