<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Consumer Health Search at CLEF eHealth 2021</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <email>lorraine.goeuriot@univ-grenoble-alpes.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hanna Suominen</string-name>
          <email>hanna.suominen@anu.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gabriella Pasi</string-name>
          <email>gabriella.pasi@unimib.it</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elias Bassani</string-name>
          <email>elias.assani@unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Brew-Sam</string-name>
          <email>nicola.brew-sam@anu.edu.au</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gabriela González-Sáez</string-name>
          <email>gabriela-nicole.gonzalez-saez@univ-grenoble-alpes.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liadh Kelly</string-name>
          <email>liadh.kelly@mu.ie</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philippe Mulhem</string-name>
          <email>philippe.mulhem@univ-grenoble-alpes.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandaru Seneviratne</string-name>
          <email>sandaru.seneviratne@anu.edu.au</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rishabh Upadhyay</string-name>
          <email>r.upadhyay@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Viviani</string-name>
          <email>marco.viviani@unimib.it</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chenchen Xu</string-name>
          <email>chenchen.xu@anu.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Consorzio per il Trasferimento Tecnologico - C2T</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Data61/Commonwealth Scientific and Industrial Research Organisation</institution>
          ,
          <addr-line>Canberra, ACT</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Maynooth University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The Australian National University</institution>
          ,
          <addr-line>Canberra, ACT</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Université Grenoble Alpes</institution>
          ,
          <addr-line>CNRS, Grenoble INP, LIG, F-38000 Grenoble</addr-line>
          <country country="FR">France</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>DISCo</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University of Turku</institution>
          ,
          <addr-line>Turku</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper details materials, methods, results, and analyses of the Consumer Health Search Task of the CLEF eHealth 2021 Evaluation Lab. This task investigates the efectiveness of information retrieval (IR) approaches in providing access to medical information to laypeople. For this a TREC-style evaluation methodology was applied: a shared collection of documents and queries is distributed, participants' runs received, relevance assessments generated, and participants' submissions evaluated. The task generated a new representative web corpus including web pages acquired from a 2021 CommonCrawl and social media content from Twitter and Reddit, along with a new collection of 55 manually generated layperson medical queries and their respective credibility, understandability, and topicality assessments for returned documents. This year's task focused on three subtask: (i) ad-hoc IR, (ii) weakly supervised IR, and (iii) document credibility prediction. In total, 15 runs were submitted to the three subtasks: eight addressed the ad-hoc IR task, three the weakly supervised IR challenge, and 4 the document credibility prediction challenge. As in previous years, the organizers have made data and tools associated with the task available for future research and development.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Dimensions of Relevance</kwd>
        <kwd>eHealth</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Health Records</kwd>
        <kwd>Medical Informatics</kwd>
        <kwd>Information Storage and Retrieval</kwd>
        <kwd>Self-Diagnosis</kwd>
        <kwd>Test-set Generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In today’s information overloaded society, using a web search engine to find information
related to health and medicine that is credible, easy to understand, and relevant to a given
information need is increasingly dificult, thereby hindering patient and public involvement in
healthcare [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These problems are referred to as credibility, understandability, and topicality
dimensions of relevance in information retrieval (IR) [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5 ref6">2, 3, 4, 5, 6</xref>
        ]. The CLEF eHealth lab (https://
clefehealth.imag.fr/) of the Conference and Labs of the Evaluation Forum (CLEF, formerly known
as Cross-Language Evaluation Forum, http://www.clef-initiative.eu/) is a research initiative
that aims at providing datasets and gathering researchers working on information extraction,
management and retrieval tasks in the medical comain. Since its establishment in 2012, CLEF
eHealth has included eight IR tasks on CHS with more more than twenty subtasks in total
[
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref14 ref15 ref16 ref17 ref7 ref8 ref9">7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17</xref>
        ]. Its usage scenario is to ease and support laypeople,
policymakers, and healthcare professionals in understanding, accessing, and authoring eHealth
information in a multilingual setting.
      </p>
      <p>In 2021, CLEF eHealth initiative has organised a CHS task with the following three IR subtasks:
1. Adhoc IR,
2. Weakly Supervised IR, and
3. Document Credibility Prediction.</p>
      <p>The task has challenged researchers, scientists, engineers, analysts, and graduate students to
develop better IR systems to support creating web-based search tools that return webpages that
are better suited to laypeople’s information needs from perspectives of
1. information topicality (i.e., how relevant are the contents to the search topic),
2. information understandability (i.e., how easily can a layperson understand the
contents), and
3. information credibility (i.e., should the contents be trusted).</p>
      <p>
        As a continuation of the previous CLEF eHealth IR tasks that ran in 2013–2018, and 2020 [
        <xref ref-type="bibr" rid="ref18 ref19 ref2 ref20 ref21 ref22 ref23 ref24">2,
18, 19, 20, 21, 22, 23, 24</xref>
        ], the 2021 CHS task has embraced the Text REtrieval Conference (TREC)
-style evaluation process. Namely, it has designed, developed, and deployed a shared collection of
documents and queries, called for the contribution of runs from participants, and conducted the
subsequent formation of relevance assessments and evaluation of the participants’ submissions.
      </p>
      <p>The main contributions of the CLEF eHealth 2021 task on CHS are as follows:
1. generating a novel representative web corpus,
2. collecting layperson medical queries,
3. attracting new submissions from participants,
4. contributing to IR evaluation metrics relevant to the three dimensions of document
relevance, and
5. evaluating IR systems (i.e., runs submitted by the task participants or from organizers’
baseline systems) on newly conducted assessments.</p>
      <p>The remainder of this paper is structured as follows: First, Section 2 details the task. Then,
Section 3 introduces the document collection, topics, baselines, pooling strategy, and evaluation
metrics. After this, Section 4 presents the participants and their approaches while Section 5
addresses their results. Finally, Section 6 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Description of the Tasks</title>
      <p>In this section, we provide a description of the three subtasks ofered in this year’s CHS task,
namely: Subtask 1 on ad-hoc IR; Subtask 2 on weakly supervised IR; and Subtask 3 on document
credibility prediction.</p>
      <sec id="sec-2-1">
        <title>2.1. Subtask 1: Adhoc Information Retrieval</title>
        <p>Similar to previous years of the CHS task, this was a standard ad-hoc IR task. A document
collection was provided to task participants along with realistic layperson medical information
need use cases, both described in Section 3. The purpose of the task was to evaluate IR
systems’ abilities to provide users with credible, understandable, and topical documents. As
such, participating teams submitted their runs, which were pooled together with baseline runs
and manual relevance assessments conducted, also described in Section 3. These systems’
performance was assessed on multiple dimensions of relevance — credibility, understandability,
and topicality. Evaluation metrics employed for this task are described in Section 3.6.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Subtask 2: Weakly Supervised Information Retrieval</title>
        <p>This task aimed to evaluate the ability of machine learning-based ad-hoc IR models, trained
with weak supervision, to retrieve relevant documents in the health domain. In order to train
neural models to address the search task, a large collection of real-world health related queries
extracted from commercial search engine query logs and synthetic (weak) relevance scores
computed with a competitive IR system were considered. The submissions were evaluated
against the same test set as Subtask 1 submissions, in order to allow a full comparison of
traditional vs neural approaches. Details on the methods and materials associated with this task
are described in Section 3.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Subtask 3: Document Credibility Prediction</title>
        <p>The purpose of this task was the automatic assessment of the credibility of information that is
disseminated online, through the Web and social media. Using the dataset related to Subtask 1,
this task aimed at comparing approaches estimating documents credibility. The ground truth
for this classification task is based on the credibility labels assigned to documents during the
relevance assessment process. Evaluation of the runs includes classical classification metrics
and investigates the ability of the participants systems to perform well on both web documents
and social media content.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Materials and Methods</title>
      <p>In this section, we will describe the materials and methods used in the CHS task of the CLEF
eHealth evaluation lab 2021. After introducing our new document collection and novel topics,
we will describe our baseline systems and pooling methodology. Finally, we will address our
relevance assessments and evaluation metrics.</p>
      <sec id="sec-3-1">
        <title>3.1. Document Collection</title>
        <p>The 2021 CLEF eHealth Consumer Health Search document collection consisted of two separate
crawls of documents: web documents acquired from the CommonCrawl and social media
documents composed by Reddit and Twitter submissions.</p>
        <p>
          First, we acquired web pages from the CommonCrawl. We extracted an initial list of websites
from the 2018 CHS task of CLEF eHealth. This list was built by submitting a set of medical
queries to the Microsoft Bing Application Programming Interfaces (through the Azure Cognitive
Services) repeatedly over a period of a few weeks, and acquiring the uniform resource Locators
(URL) of the retrieved results [
          <xref ref-type="bibr" rid="ref13 ref22">22, 13</xref>
          ]. We included the domains of the acquired URLs in the
2021 list, except some domains that were excluded for decency reasons. After this, similarly
to 2018, we augmented the list by including a number of known reliable and unreliable health
websites, domains, and social media contents of ranging reliability levels. Finally, we further
extended the list by including websites that were highly relevant for the task queries to finalize
our list of 600 domains. In summary, we introduced thirteen new domains in 2021 compared
to the 2018 collection, and then the newly crawled domains from the latest CommonCrawl
2021-04 (https://commoncrawl.org/2021/02/january-2021-crawl-archive-now-available/).
        </p>
        <p>Second, we complemented the collection with social media documents from Reddit and
Twitter. As the first step, we selected a list of 150 health topics related to various health
conditions. Then, we generated search queries manually from those topics and submitted them
to Reddit to retrieve posts and comments. After this, we applied the same process on Twitter to
get related tweets from the platform.</p>
        <p>Please note that for the purposes of the 2021 CHS task, we defined a social media document
as a text obtained by a single interaction. This implied that
• on Reddit, a document is composed by a post, one comment of the post, and associated
meta-information and
• on Twitter, a document is a single tweet with its associated meta-information.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Topics</title>
        <p>The set of topics of CLEF eHealth IR task aimed at being representative of laypersons’ medical
information needs in various scenarios. This year, the set of topics was collected from two
sources as follows:
• The first part was based on our insights from consulting laypeople with lived experience
of multiple sclerosis (MS) or diabetes to motivate, validate, and refine the search
scenarios; we, as experts in IR and CHS tasks, captured these layperson-informed insights as
the scenarios.
• The second part was based on use cases from Reddit health forums; we extracted and
manually selected a list of topics from Google trends to best fit each use case.</p>
        <p>As a result, we had a new set of 55 queries in English for authoring realistic search scenarios.
To describe the scenarios, we enriched each query manually by labels in English either to
characterize the search intent (for manually created queries) or to capture the submission text
(for social medial queries) (Table 3.2).</p>
        <sec id="sec-3-2-1">
          <title>Scenario</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>Query Narrative 8 57</title>
          <p>126
best apps daily activ- I’m a 15 year old with diabetes. I’m planning to join the school
ity exercise diabetes hiking club. What are the best apps to track my daily activity
and exercises?
multiple sclerosis I read that MS develops with several stages and includes phases
stages phases of relapse. I want to know more about how this disease develops
over time.
birth control sup- So I’ve been on birth control taking only active pills for months
pression antral now and I went in to a doctor’s ofice to inquire about egg
freezfollicle count ing. She seemed optimistic that my ultrasound would be very
reassuring but instead she came away from the ultrasound deeply
concerned. I did an AMH test and it came back 0.3 which was
consistent with what was seen in the ultrasound. I’m terrified.</p>
          <p>Like...suicidal terrified. I don’t have any underlying conditions
and while I’m not young I’m not old either. I don’t smoke and I’m
not terribly overweight. I don’t have insulin resistance and my
reproductive organs looks good other than my ovaries. I’d been
taking the combo pill to skip periods for months now. I don’t even
take the brown pills. When I first heard that you can skip
periods and only have 2-3 a year when taking brown pills I started on
them. The first time I tried to have a period on the brown pills
nothing happened. No period. This was 2 years ago. I freaked out
but they said I just respond strongly to the pills. So I went of of
the pills completely and had 2-3 normal periods starting about 2
months later. Could the birth control be suppressing my antral
follicle count and amh this much? Please help.</p>
          <p>Subtasks 1, 2, and 3 used these 55 queries with 5 released for training and 50 reserved for
testing; the test topics contained a balanced sample of the manually constructed and automatically
extracted search scenarios.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Baseline Systems</title>
        <p>
          With respect to both Subtask 1 and Subtask 2, we provided six baseline systems. These systems
applied the Okapi BM25 (BM for Best Matching), Dirichlet Language Model (DirichletLM),
and Term Frequency times Inverse Document Frequency (TF× IDF) algorithms with default
parameters. Each of them was implemented without and with pseudo relevance feedback (RF)
using default parameters (DFR Bo1 model [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] on three documents, selecting ten terms). This
resulted in the 3 × 2 = 6 baseline systems. The systems are implemented using Terrier version
5.4 [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] as following:
terrier batchretrieve -t topics.txt -w [TF_IDF|DirichletLM|BM25] [-q|].
        </p>
        <p>Regarding Subtask 3, where it is necessary to evaluate the efectiveness of the approaches with
respect to assessing the credibility of information, it was decided to approach the problem as a
binary classification (identification of credible versus non-credible information). For this reason,
simple baselines were developed based on supervised classifiers that act on a set of features that
can be extracted from the documents under consideration. Dealing with documents that are
both Web pages and social content, it was necessary to consider, in baseline development, only
those features that are directly extractable from the text of the documents. For social content, it
would also be possible to consider other metadata related to the social network of the authors of
the posts, but for consistency with the evaluation of Web pages, such metadata have not been
considered.</p>
        <p>
          The employed supervised classifiers are based on Support Vector Machines (SVM), Random
Forests (RF), and Logistic Regression (LR), i.e., the machine learning techniques that have proven
to be more efective for binary credibility assessment in the literature [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. Such baselines act
on the following linguistic features: the TF× IDF text representation, the word embedding text
representation obtained by Word2vec pre-trained on Google News,1 and, given the health-related
scenario considered, the word embedding representation obtained by Word2vec pre-trained on
both PubMed and Medical Information Mart for Intensive Care III (MIMIC-III).2 Python and
the scikit-learn library [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ] have been employed for implementing the machine learning
algorithms and for producing the TF× IDF text representation. Furthermore, we have considered
threshold tuning to get the optimal cut-of for the runs, based on the Youden’s index [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Pooling Methodology</title>
        <p>
          Similar to the 2016, 2017, 2018, and 2020 pools, we created the pool using the rank-biased
precision (RBP)-based Method A (Summing contributions) [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] in which documents are weighted
according to their overall contribution to the efectiveness evaluation as provided by the RBP
formula (with  = 0.8, following a study published in 2007 on RBP [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]). This strategy, called
RBPA, has been proven more eficient than traditional fixed-depth or stratified pooling to
evaluate systems under fixed assessment budget constraints [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ], as it was the case for this
task. All participants’ runs were considered on the document’s pool, along with six baselines
provided by the organizers. In order to guarantee the judgements of the documents of the
participants’ runs, half of the pool was composed by their documents and half from documents
of the baselines’ runs which resulted in 250 documents per query in the pool.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Relevance Assessment</title>
        <p>The credibility, understandability, and topicality assessments were performed by 26 volunteers
(19 women and 7 men) in May–June 2021. Of these assessors, 16 were from Australia, 4 from
Italy, 3 from France, 2 from Ireland, and 1 from Finland. We recruited, trained, and supervised
them by using bespoke written materials from April to June 2021. The recruitment took place
via email and on social media, using both our existing contacts and snowballing.</p>
        <p>We implemented these assessments online by expanding and customising the Relevation!
tool for relevance assessments [33] to capture the three dimensions of document relevance,
and their scale (see [17, Figures 4–6] for illustrations of the online assessment environment).
1https://github.com/mmihaltz/word2vec-GoogleNews-vectors
2https://github.com/ncbi-nlp/BioSentVec
We associated every query with 250 documents to be assessed with respect to their credibility,
understandability, and topicality. Initially, we allocated each assessor with 2 queries for their
assessment and then revised these allocations based on their individual needs and availability.
In the end, every assessor completed 1–4 queries.</p>
        <p>Ethical approval (2021/013) was obtained for all aspects of this assessment study involving
human participants as assessors from the Human Research Ethics Committee of the Australian
National University (ANU). Each study participant provided informed consent.</p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6. Evaluation Metrics</title>
        <p>We considered evaluation measures that allowed to evaluate both:
1. the efectiveness of the systems with respect to the ranking produced by taking into
account the three criteria considered, and
2. the accuracy of the (binary) classification of documents with respect to credibility
(Subtask 3).</p>
        <p>
          Specifically, with regard to the first aspect, this included the following performance
evaluation measures: Mean Average Precision (MAP), preference-based BPref metric, normalized
Discounted Cumulative Gain (nDCG), understandability-based variant of Rank Biased Precision
(uRBP), and credibility-ranked Rank Biased Precision (cRBP) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>With respect to the second aspect, we referred to classical measures to assess the goodness of
a classifier, such as Accuracy and the Area under the Receiver Operating Characteristic (ROC)
Curve (AUC) [34]. In particular, since Subtask 3 is also devoted to assess credibility in relation
to information needs (topics), also the average of a Topic-based Credibility Precision w.r.t. each
topic, namely CP(), has been considered.</p>
        <p>
          A brief explanation of the three measures that need more detail, namely uRBP, cRBP, and
CP() is provided below.
3.6.1. Understandability-Ranked Biased Precision
The uRBP measure evaluates IR systems by taking into account both topicality and
understandability dimensions of relevance. In particular, the function for calculating uRBP was [35]:

uRBP = (1 −  ) ∑︁  − 1(@)· (@),
=1
(1)
where (@) is the relevance of the document  at position , (@) is the understandability
value of the document  at position , and the persistent parameter  models the user desire to
examine every answer, which was set to 0.50, 0.80, and 0.95 to obtain three version of uRBP,
according to diferent user behaviors.
3.6.2. Credibility-Ranked Biased Precision
In CLEF eHealth 2020, we have adapted uRBP to credibility, obtaining the so-called
credibilityranked biased precision (cRBP) measure [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. In this case the function for calculating cRBP was
the same used to calculate uRBP, replacing (@) by the credibility value of the document 
at position , (@):
        </p>
        <p>cRBP = (1 −  ) ∑︁  − 1(@)· (@).</p>
        <p>=1
(2)
As in uRBP, the parameter  was set to three values, from an impatient user (0.50) to more
persistent users (0.80 and 0.95).
3.6.3. Topic-based Credibility Precision
The precision in retrieving credible documents can be calculated over the top- documents in
the ranking, for each topic (query) , as follows:</p>
        <p>CP() =
#____()
#___()
.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Participants and Approaches</title>
      <p>In 2021, 43 teams registered for the task on the web site and two teams submitted runs to the
subtasks. We provided the registered participants with the crawler code and the domain list of
the crawl for both the web documents and social media documents. They also had access to
indexes built by organizers from the document collection on demand. Participants’ submissions
were due by 8 May 2021.</p>
      <p>Of the four run submissions, two were to Subtask 1 on Adhoc IR, one was to Subtask 2 on
Weakly Supervised IR, and one was to Subtask 3 on Document Credibility Prediction. The teams
were from two countries (i.e., China and Italy) in two continents (i.e., Asia and Europe). In
Subtask 1, the submissions were by
1. a 4-member team from the School of Computer Science, Zhongyuan University of
Technology (ZUT) in Zhengzhou, China [36] and
2. a 2-member team from the Information Management Systems (IMS) Research Group,</p>
      <p>University of Padova (UniPd), Padova, Italy [37].</p>
      <p>In Subtasks 2 and 3, the submissions were by the leader of this IMS UniPd team [37] — a regular
participant in our previous CHS tasks.</p>
      <p>The two teams submitted the following types of document ranking approaches as runs to
the Adhoc IR subtask (Table 4): Team ZUT used a learning-to-rank approach in all its four
submitted runs, but with four diferent machine learning algorithms to train their models [ 36].
In contrast, Team UniPd founded their four submissions on a renown Python framework for IR
called PyTerrier as variants of Reciprocal Rank Fusion with the provided Terrier index [37].</p>
      <p>The UniPd team submitted closely related approaches to the other two subtasks (Table 4).
Namely, they also submitted their aforementioned four approaches to the Weekly Supervised
IR subtask and two of them and another two of their variants to the Document Credibility
Prediction subtask [37].
LM
MR
RFs
RB
BM25, QLM, &amp; DFR</p>
      <sec id="sec-4-1">
        <title>BM25, QLM, &amp; DFR using the</title>
      </sec>
      <sec id="sec-4-2">
        <title>RM3 relevance language model for pseudo RF</title>
      </sec>
      <sec id="sec-4-3">
        <title>BM25, QLM, &amp; DFR on manual variants of the query</title>
      </sec>
      <sec id="sec-4-4">
        <title>BM25, QLM, &amp; DFR on manual</title>
        <p>variants using RM3 pseudo RF</p>
      </sec>
      <sec id="sec-4-5">
        <title>UniPd’s runs a &amp; b above (i.e., the ones without manual variants of the query) merged with min-max normalization</title>
      </sec>
      <sec id="sec-4-6">
        <title>UniPd’s runs c &amp; d above (i.e., the ones with manual variants of the query) merged with min-max normalization</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In 2021, the CLEF eHealth CHS task generated a new representative Web corpus including
Web pages acquired from a 2021 CommonCrawl, and social media content from Twitter and
Reddit. A new collection of 55 manually generated layperson medical queries was also created,
along with their respective credibility, understandability, and topicality assessments for returned
documents. In total, 15 runs were submitted to the three subtasks on adhoc IR, weakly supervised
IR, and document credibility prediction, respectively.</p>
      <sec id="sec-5-1">
        <title>5.1. Coverage of Relevance Assessments</title>
        <p>A total of 12, 500 assessments were made on 11, 357 documents: 7, 400 Web documents, and
3, 957 social media documents. Figure 1 shows the number of social media and Web documents
that were assessed for each query. The bottom part shows which queries were created from
discussion with patients (Expert queries), and queries created from discussions on social media
(Social media queries). We can see that Web documents took a bigger part of the pool of
documents. Nevertheless, the proportion of social media documents was bigger for queries
based on social media. A special case is presented in queries 22, 63, and 116, where the pool
# of documents</p>
        <sec id="sec-5-1-1">
          <title>Topical</title>
        </sec>
        <sec id="sec-5-1-2">
          <title>Understandable</title>
        </sec>
        <sec id="sec-5-1-3">
          <title>Credible Highly 925</title>
          <p>Web</p>
        </sec>
        <sec id="sec-5-1-4">
          <title>Somewhat</title>
          <p>2, 540
3, 014
3, 123</p>
          <p>Not
4, 259
1, 202
753
of documents was composed by Web documents only. This means that the submitted runs
retrieved only Web documents.</p>
          <p>While the distribution of the assessments on topicality and understandability dimensions
were similar on social media and Web documents, the credibility assessments presented a big
diference (Table 5.1), with only 12 documents (less than 1% of social media assessed documents)
were assessed as highly credible in the social media set, in contrast to the 4, 552 (54% of Web
assessed documents) highly credible Web documents. Finally, an important part of social media
documents were assessed as not credible. Figure 2 shows the relationship between the three
dimensions of relevance; its diagonal exemplifies the distribution of each dimension. Each plot
exposes the number of documents evaluated as “highly”, “somewhat”, or “not”, with respect to
two dimensions of relevance (one in each axis). By example, in social media assessments the
Figure 2 shows for the not credible documents that a fraction is not topical relevant, a set of
documents is somewhat topical relevant, and a small part of the not credible is anyways highly
topically relevant.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Subtask 1: Adhoc Information Retrieval</title>
        <p>In this section we present the results for the Adhoc IR subtask where the systems were evaluated
in diferent dimensions of relevance. For topicality, we evaluated the systems using MAP, BPref,
and NDCG@10 performances metrics. For readability, we made use of uRBP that considers
topicality and understandability of the documents. Finally, in order to include credibility in our
evaluation, we measured systems’ cRBP performance that considers topicality and credibility of
the document.</p>
        <p>Table 5.2 presents the ranking of participant systems and organizers baselines, with respect
to three metrics of topical relevance: MAP, BPref, and NDCG@10. The team achieving the
highest results was UniPd. Their top run, original_rm3_rrf used Reciprocal Rank Fusion
with BM25, QLM, DFR approaches using pseudo relevance feedback with 10 documents and 10
terms (query weight 0.5). It achieved 0.43 MAP and 0.51 BPref. The best system of ZUT team
was their run3, which used learning to rank techniques with a model trained using Random
Forests. ZUT run3 was the best systems of the team on the three metrics with 0.4 MAP, 0.47
BPref, and 0.66 NDCG@10. For BPref and NDCG@10, the organizers baseline using TF× IDF
with query expansion obtained higher results.</p>
        <p>Table 5.2 compares the results of our baseline systems with respect to last year’s Adhoc IR
task results. In the case of MAP and BPref metrics, all the systems had better performance on
2021 test collection; for NDCG@10, DirichletLM with and without relevance feedback had a
better performance on 2020 test collection. Despite this clear improvement from 2020 to 2021,
the ranking of systems has changed in all the metrics. Namely, in 2020, the best MAP and BPref
performance was achieve by DirichletLM, and TF× IDF with respect to NDCG@10 metric. In
contrast, in this year’s task, the best baseline MAP, Bpref, and NDCG@10 was achieved by
TF× IDF with relevance feedback.</p>
        <p>Rank</p>
        <p>Team - run
1
2
3
4
5
6
7
8
9
10
11
12
13
14</p>
        <p>UniPd
original_rm3_rrf
UniPd
simplified_rm3_rrf
ZUT
run3_clef2021_task2
Baseline
terrier_TF× IDF_qe
Baseline
terrier_BM25_qe
ZUT
run4_clef2021_task2
Baseline
terrier_DirichletLM
Baseline
terrier_TF× IDF
Baseline
terrier_BM25
ZUT
run2_clef2021_task2
ZUT
run1_clef2021_task2
Baseline
terrier_DirichletLM_qe
UniPd
original_rrf
UniPd
simplified_rrf
0.431
0.431</p>
        <p>The results for all participants for readability and credibility evaluation is shown in Table 5.2.
For readability evaluation the best run was the baseline system TF× IDF with relevance feedback.
The best participant system was ZUT’s run3 that is the same system that presented the best
performance in terms of topicality as well. For credibility evaluation, the best run was the
baseline system DirichletLM while the best participant system was UniPd’s original_rm3_rrf —
the best system of the team in topicality evaluation.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Subtask 2: Weakly Supervised Information Retrieval</title>
        <p>The purpose of the task was to evaluate systems based on machine learning, trained on the
weakly supervised dataset. UniPD submitted their runs to Subtask 1 for this subtask, which are
not trained on the weakly supervised dataset. Their submission provides a kind of baseline for
a system that does not use any information. Since no other team submitted runs to this task,
we cannot provide any evaluation.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Subtask 3: Document Credibility Prediction</title>
        <p>The UniPd team, to assess document credibility, reused the runs computed in Subtask 1 and
grouped them in order to produce a single score for each document. Their simple hypothesis
was that documents that have a higher score across diferent search engines are also more
credible. Their approach did not consider any additional information about the provenance of
the document.
5.4.1. Credibility Assessment as a Binary Classification Problem
The results illustrated in this section were obtained by performing a binary credibility assessment
using the CLEF 2020 eHealth dataset for training the baselines and testing on a subset of the CLEF
2021 eHealth data, that is, those that were employed for both runs subtask1_ims_original
and subtask1_ims_simplified submitted by the UniPd team. In this case, the topics against
which the documents were retrieved are not taken into account. The results of classifying
documents according to their credibility using both simple baselines (illustrated in Section 3.3)
and considered runs are illustrated in Table 7.</p>
        <p>The results obtained from both baselines and runs were evaluated by means of classical
measures in the context of document classification, that is, the AUC and Accuracy (as discussed
in Section 3.6).</p>
        <sec id="sec-5-4-1">
          <title>Model</title>
        </sec>
        <sec id="sec-5-4-2">
          <title>Baseline SVM_orig_tf_idf</title>
        </sec>
        <sec id="sec-5-4-3">
          <title>Baseline RF_orig_tf_idf</title>
        </sec>
        <sec id="sec-5-4-4">
          <title>Baseline LR_orig_tf_idf</title>
        </sec>
        <sec id="sec-5-4-5">
          <title>Baseline SVM_orig_w2v_google</title>
        </sec>
        <sec id="sec-5-4-6">
          <title>Baseline RF_orig_w2v_google</title>
        </sec>
        <sec id="sec-5-4-7">
          <title>Baseline LR_orig_w2v_google</title>
        </sec>
        <sec id="sec-5-4-8">
          <title>Baseline SVM_orig_w2v_bio</title>
        </sec>
        <sec id="sec-5-4-9">
          <title>Baseline RF_orig_w2v_bio</title>
        </sec>
        <sec id="sec-5-4-10">
          <title>Baseline LR_orig_w2v_bio</title>
        </sec>
        <sec id="sec-5-4-11">
          <title>Baseline SVM_simp_tf_idf</title>
        </sec>
        <sec id="sec-5-4-12">
          <title>Baseline RF_simp_tf_idf</title>
        </sec>
        <sec id="sec-5-4-13">
          <title>Baseline LR_simp_tf_idf</title>
        </sec>
        <sec id="sec-5-4-14">
          <title>Baseline SVM_simp_w2v_google</title>
        </sec>
        <sec id="sec-5-4-15">
          <title>Baseline RF_simp_w2v_google</title>
        </sec>
        <sec id="sec-5-4-16">
          <title>Baseline LR_simp_w2v_google</title>
        </sec>
        <sec id="sec-5-4-17">
          <title>Baseline SVM_simp_w2v_bio</title>
        </sec>
        <sec id="sec-5-4-18">
          <title>Baseline RF_simp_w2v_bio</title>
        </sec>
        <sec id="sec-5-4-19">
          <title>Baseline LR_simp_w2v_bio</title>
        </sec>
        <sec id="sec-5-4-20">
          <title>Run subtask1_ims_original</title>
        </sec>
        <sec id="sec-5-4-21">
          <title>Run subtask1_ims_simplified</title>
          <p>As it can be observed from the table, having considered only linguistic features, the
classification results obtained were not particularly significant to the considered problem, neither for the
runs nor for the baselines considered. The idea discussed by the members of the UniPd group,
that documents having a higher score across diferent search engines are also more credible,
did not appear to be supported by these results, but nevertheless merited future investigation,
in particular in relation to the results reported in the next section. Regarding the baselines,
it was possible to observe that, when supervised classifiers employ the word embedding text
representation obtained by Word2vec pre-trained on biomedical datasets, such as Pubmed
and MIMIC-III, we obtained best results as compared to their counterparts employing text
representation features based on TF× IDF and Word2vec pre-trained on the general-purpose
Google News dataset. However, it must be considered that, in general, the Accuracy values in
particular could be influenced by the fact that the dataset considered was strongly unbalanced,
being made up of about 20% of documents labeled as credible and 80% labeled as non-credible.
Probably, this should also raise the need to deepen an analysis about the way in which human
assessors evaluate the documents proposed to them.
5.4.2. Topic-based Credibility Assessment
Regarding the problem of assessing the credibility of the documents retrieved with respect
to the topics considered, we provide below the results relative to the runs sent by UniPD,
namely subtask2_ims_original and subtask2_ims_simplified. In this case, for each
topic considered, the precision in retrieving credible documents with respect to the topic was
evaluated based on the value of CP() calculated as shown in Section 3.6.3.</p>
          <p>In Table 8, we report the average CP() values with respect to the set of topics considered.
Specifically, the top-100 and top-200 documents retrieved with respect to each topic were taken
into account in computing the CP() metric.</p>
        </sec>
        <sec id="sec-5-4-22">
          <title>Model</title>
        </sec>
        <sec id="sec-5-4-23">
          <title>Run subtask2_ims_original</title>
        </sec>
        <sec id="sec-5-4-24">
          <title>Run subtask2_ims_simplified</title>
        </sec>
        <sec id="sec-5-4-25">
          <title>Run subtask2_ims_original</title>
        </sec>
        <sec id="sec-5-4-26">
          <title>Run subtask2_ims_simplified</title>
          <p>As it can be observed from the table, with respect to the evaluations carried out in the previous
case related to binary credibility assessment (i.e., Section 5.4.1), it seems that the methods applied
in the two submitted runs that consider credibility associated with documents retrieved with
respect to specific topics, actually managed to find credible information in the first top- 
positions, in particular as  decreases. In general, the most efective method seems to be the
one implemented in the run subtask2_ims_original; however, from specific observations
with respect to the results submitted by the participants (which are omitted in this paper),
we can state that the second method, the one implemented in subtask2_ims_simplified,
proved to be particularly eficient for some topics (even more than the method implemented
in subtask2_ims_original) and much less so with respect to others, leading to a lower
average CP() value. It is therefore worth investigating this aspect more in the future.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>This paper has described methods, results and analysis of the Consumer Health Search (CHS)
challenge at the CLEF eHealth 2021 Evaluation Lab. The task considered the problem of lay
people searching for medical information related to their health condition on the web. In
particular, across web pages and social media content. The task included three subtasks on ad
hoc IR, weakly supervised IR, and document credibility prediction. 15 runs were submitted to
these tasks.</p>
      <p>The CLEF eHealth CHS challenge, first ran at the inaugural CLEF eHealth lab in 2013, and
has ofered since then TREC-style IR challenges with medical datasets consisting of large
medical web and lay person query collections. As a by-product of this evaluation exercise,
the task contributes to the research community a collection with associated assessments and
evaluation framework that can be used to evaluate the efectiveness of retrieval methods for
health information seeking on the web. Queries, assessments, and participants’ runs for the
2021 CHS challenge are publicly available at https://github.com/CLEFeHealth/CHS-2021 and
previous years’ CHS collections are available at https://github.com/CLEFeHealth/.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>Thanks: The CHS task of the CLEF eHealth 2021 evaluation lab has been supported in part
by the CLEF Initiative. It has also been supported in part by the Our Health in Our Hands
(OHIOH) initiative of the Australian National University (ANU), as well as the ANU School of
Computing, ANU Research School of Population Health, and Data61/Commonwealth Scientific
and Industrial Research Organisation. OHIOH is a strategic initiative of the ANU which aims
to transform health care by developing new personalised health technologies and solutions
in collaboration with patients, clinicians, and health care providers. Moreover, the task has
been supported in part by the bi-lateral Kodicare (Knowledge Delta based improvement and
continuous evaluation of retrieval engines) project funded by the French ANR
(ANR-19-CE230029) and Austrian FWF. Finally, the task has been supported in part by the EU Horizon 2020
Research and Innovation Programme under the Marie Skłodowska-Curie Grant Agreement No
860721 – DoSSIER: “Domain Specific Systems for Information Extraction and Retrieval”. We are
also thankful to the people involved in the query creation and relevance assessment exercises.
Last but not least, we gratefully acknowledge the participating teams’ hard work. We thank
them for their submissions and interest in the task.</p>
      <p>Author Contribution Statement: With equal contribution, Task 2 was led by LG, GP, and
HS, and organized by EB, NB-S, GG-S, LK, PM, GP, SS, HS, RU, MV, and CX.
[33] B. Koopman, G. Zuccon, Relevation!: an open source system for information retrieval
relevance assessment, in: Proceedings of the 37th International ACM SIGIR Conference
on Research &amp; Development in Information Retrieval, ACM, 2014, pp. 1243–1244.
[34] H. Suominen, S. Pyysalo, M. Hiissa, F. Ginter, S. Liu, D. Marghescu, T. Pahikkala, B. Back,
H. Karsten, T. Salakoski, Performance evaluation measures for text mining, in: M. Song,
Y. Wu (Eds.), Handbook of Research on Text and Web Mining Technologies, IGI Global,
Hershey, Pennsylvania, USA, 2008, pp. 724–747.
[35] G. Zuccon, Understandability biased evaluation for information retrieval, in: Advances in</p>
      <p>Information Retrieval, 2016, pp. 280–292.
[36] H. Yang, X. Liu, B. Zheng, G. Yang, Learning to rank for Consumer Health Search, in:
CLEF 2021 Evaluation Labs and Workshop: Online Working Notes, CEUR-WS, September
2021.
[37] G. Di Nunzio, F. Vezzani, IMS-UNIPD @ CLEF eHealth Task 2: Reciprocal Ranking Fusion
in CHS, in: CLEF 2021 Evaluation Labs and Workshop: Online Working Notes, CEUR-WS,
September 2021.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Soroya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farooq</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Mahmood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Isoaho</surname>
          </string-name>
          , S. e Zara,
          <article-title>From information seeking to information avoidance: Understanding the health information behavior during a global health crisis</article-title>
          ,
          <source>Information Processing and Management</source>
          <volume>58</volume>
          (
          <year>2021</year>
          )
          <fpage>102440</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leveling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Salantera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          , G. Zuccon,
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2013</year>
          ,
          <article-title>Task 3: Information retrieval to address patients' questions when reading clinical reports</article-title>
          ,
          <source>CLEF 2013 Online Working Notes</source>
          <volume>8138</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          , L. Goeuriot,
          <article-title>Scholarly influence of the Conference and Labs of the Evaluation Forum eHealth Initiative: Review and bibliometric study of the 2012 to 2017 outcomes</article-title>
          ,
          <source>JMIR Research Protocols</source>
          <volume>7</volume>
          (
          <year>2018</year>
          )
          <article-title>e10961</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <article-title>Consumer health search on the web: Study of web page understandability and its integration in ranking algorithms</article-title>
          ,
          <source>J Med Internet Res</source>
          <volume>21</volume>
          (
          <year>2019</year>
          )
          <article-title>e10986</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Goeuriot,</surname>
          </string-name>
          <article-title>The scholarly impact and strategic intent of CLEF eHealth Labs from 2012 to 2017</article-title>
          , in: N.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Peters (Eds.),
          <source>Information Retrieval Evaluation in a Changing World: Lessons Learned from 20 Years of CLEF</source>
          , Springer International Publishing, Cham,
          <year>2019</year>
          , pp.
          <fpage>333</fpage>
          -
          <lpage>363</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          , G. Pasi,
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Saez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Xu, Overview of the CLEF eHealth 2020 task 2: consumer health search with ad hoc and spoken queries</article-title>
          , in: Working Notes of Conference and
          <article-title>Labs of the Evaluation (CLEF) Forum</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          (Ed.),
          <source>The Proceedings of the CLEFeHealth2012 - the CLEF 2012 Workshop on Cross-Language Evaluation of Methods</source>
          , Applications, and
          <article-title>Resources for eHealth Document Analysis</article-title>
          ,
          <string-name>
            <surname>NICTA</surname>
          </string-name>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Salanterä</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Velupillai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. W.</given-names>
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Savova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Elhadad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pradhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>South</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Mowery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leveling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Martinez</surname>
          </string-name>
          , G. Zuccon,
          <source>Overview of the ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2013</year>
          , in: Information Access Evaluation. Multilinguality, Multimodality, and Visualization, Springer Berlin Heidelberg,
          <year>2013</year>
          , pp.
          <fpage>212</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schreck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Leroy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Mowery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Velupillai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Martinez</surname>
          </string-name>
          , G. Zuccon,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <source>Overview of the ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2014</year>
          , in: Information Access Evaluation. Multilinguality, Multimodality, and Visualization, Springer Berlin Heidelberg,
          <year>2014</year>
          , pp.
          <fpage>172</fpage>
          -
          <lpage>191</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hanlen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Névéol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Grouin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          , G. Zuccon,
          <source>Overview of the CLEF eHealth Evaluation Lab</source>
          <year>2015</year>
          , in: Information Access Evaluation. Multilinguality, Multimodality, and Visualization, Springer Berlin Heidelberg,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Névéol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          , G. Zuccon,
          <source>Overview of the CLEF eHealth Evaluation Lab</source>
          <year>2016</year>
          , in: International
          <source>Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , Springer Berlin Heidelberg,
          <year>2016</year>
          , pp.
          <fpage>255</fpage>
          -
          <lpage>266</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Névéol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Robert</surname>
          </string-name>
          , E. Kanoulas,
          <string-name>
            <given-names>R.</given-names>
            <surname>Spijker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          , G. Zuccon,
          <article-title>CLEF 2017 eHealth Evaluation Lab overview</article-title>
          ,
          <source>in: International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , Springer Berlin Heidelberg,
          <year>2017</year>
          , pp.
          <fpage>291</fpage>
          -
          <lpage>303</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Névéol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ramadier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Robert</surname>
          </string-name>
          , E. Kanoulas,
          <string-name>
            <given-names>R.</given-names>
            <surname>Spijker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          , Jimmy,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          , G. Zuccon,
          <source>Overview of the CLEF eHealth Evaluation Lab</source>
          <year>2018</year>
          , in: International
          <source>Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , Springer Berlin Heidelberg,
          <year>2018</year>
          , pp.
          <fpage>286</fpage>
          -
          <lpage>301</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kanoulas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Spijker</surname>
          </string-name>
          , G. Zuccon,
          <string-name>
            <given-names>H.</given-names>
            <surname>Scells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <source>Overview of the CLEF eHealth Evaluation Lab</source>
          <year>2019</year>
          , in: F. Crestani,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Savoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rauber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. Heinatz</given-names>
            <surname>Bürki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , Springer International Publishing, Cham,
          <year>2019</year>
          , pp.
          <fpage>322</fpage>
          -
          <lpage>339</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miranda-Escalada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          , G. Pasi,
          <string-name>
            <given-names>G. Gonzalez</given-names>
            <surname>Saez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Xu, Overview of the CLEF eHealth evaluation lab 2020</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickhof</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Névéol</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , Springer International Publishing, Cham,
          <year>2020</year>
          , pp.
          <fpage>255</fpage>
          -
          <lpage>271</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Alemany</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Brew-Sam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Cotik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Filippo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Saez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Luque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mulhem</surname>
          </string-name>
          , G. Pasi,
          <string-name>
            <given-names>R.</given-names>
            <surname>Roller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Seneviratne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vivaldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          , C. Xu,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <article-title>eHealth 2021 Evaluation Lab</article-title>
          ,
          <source>in: Advances in Information Retrieval - 43st European Conference on IR Research</source>
          , Springer, Heidelberg, Germany,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Alonso</given-names>
            <surname>Alemany</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Bassani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Brew-Sam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Cotik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Filippo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gonzalez-Saez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Luque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mulhem</surname>
          </string-name>
          , G. Pasi,
          <string-name>
            <given-names>R.</given-names>
            <surname>Roller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Seneviratne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Upadhyay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vivaldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Xu, Overview of the CLEF eHealth evaluation lab 2021</article-title>
          ,
          <source>in: CLEF 2021 - 11th Conference and Labs of the Evaluation Forum, Lecture Notes in Computer Science (LNCS)</source>
          , Springer, Heidelberg, Germany,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pecina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <string-name>
            <surname>H. M. Gareth J.F. Jones</surname>
          </string-name>
          ,
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2014</year>
          ,
          <article-title>Task 3: User-centred health information retrieval, in: CLEF 2014 Evaluation Labs</article-title>
          and Workshop: Online Working Notes, Shefield, UK,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanburyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          , P. Pecina,
          <source>CLEF eHealth Evaluation Lab</source>
          <year>2015</year>
          ,
          <article-title>Task 2: Retrieving Information about Medical Symptoms</article-title>
          , in: CLEF 2015 Online Working Notes, CEUR-WS,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pecina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Budaher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Deacon</surname>
          </string-name>
          ,
          <source>The IR Task at the CLEF eHealth Evaluation Lab</source>
          <year>2016</year>
          :
          <article-title>User-centred Health Information Retrieval, in: CLEF 2016 Evaluation Labs</article-title>
          and Workshop: Online Working Notes, CEUR-WS,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          , G. Zuccon, Jimmy,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pecina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <article-title>CLEF 2017 Task Overview: The IR Task at the eHealth Evaluation Lab</article-title>
          , in: Working Notes of Conference and
          <article-title>Labs of the Evaluation (CLEF) Forum</article-title>
          , CEUR Workshop Proceedings,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22] . Jimmy, G. Zuccon,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <article-title>Overview of the CLEF 2018 consumer health search task</article-title>
          , in: Working Notes of Conference and
          <article-title>Labs of the Evaluation (CLEF) Forum</article-title>
          , CEUR Workshop Proceedings,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leveling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <article-title>An Analysis of Evaluation Campaigns in ad-hoc Medical Information Retrieval: CLEF eHealth 2013 and</article-title>
          2014, Springer Information Retrieval Journal (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          , G. Pasi,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Gonzales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Xu, Overview of the CLEF eHealth 2020 task 2: Consumer health search with ad hoc and spoken queries</article-title>
          , in: Working Notes of Conference and
          <article-title>Labs of the Evaluation (CLEF) Forum</article-title>
          , CEUR Workshop Proceedings,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>G.</given-names>
            <surname>Amati</surname>
          </string-name>
          ,
          <article-title>Probabilistic Models for Information Retrieval Based on Divergence from Randomness</article-title>
          ,
          <source>Ph.D. thesis</source>
          , Glasgow University, Glasgow, the UK,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>I.</given-names>
            <surname>Ounis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lioma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          , V. Plachouras,
          <article-title>Research directions in terrier: a search engine for advanced retrieval on the web</article-title>
          ,
          <source>CEPIS Upgrade Journal</source>
          <volume>8</volume>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          , G. Pasi,
          <article-title>Credibility in social media: opinions, news, and health information-a survey</article-title>
          ,
          <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery</source>
          <volume>7</volume>
          (
          <year>2017</year>
          )
          <article-title>e1209</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>L.</given-names>
            <surname>Buitinck</surname>
          </string-name>
          , G. Louppe,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Niculae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Grobler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Layton</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. VanderPlas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Holt</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Varoquaux, API design for machine learning software: experiences from the scikit-learn project</article-title>
          ,
          <source>in: ECML PKDD Workshop: Languages for Data Mining and Machine Learning</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>108</fpage>
          -
          <lpage>122</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Youden</surname>
          </string-name>
          ,
          <article-title>Index for rating diagnostic tests</article-title>
          ,
          <source>Cancer</source>
          <volume>3</volume>
          (
          <year>1950</year>
          )
          <fpage>32</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mofat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zobel</surname>
          </string-name>
          ,
          <article-title>Rank-biased precision for measurement of retrieval efectiveness</article-title>
          ,
          <source>ACM Transactions on Information Systems</source>
          <volume>27</volume>
          (
          <year>2008</year>
          ) 2:
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          :
          <fpage>27</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Zhang,</surname>
          </string-name>
          <article-title>On the distribution of user persistence for rank-biased precision</article-title>
          ,
          <source>in: Proceedings of the 12th Australasian document computing symposium</source>
          ,
          <year>2007</year>
          , pp.
          <fpage>17</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lipani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <article-title>Fixed-cost pooling strategies based on ir evaluation measures</article-title>
          ,
          <source>in: European Conference on Information Retrieval</source>
          , Springer,
          <year>2017</year>
          , pp.
          <fpage>357</fpage>
          -
          <lpage>368</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>