<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CUNI team: CLEF eHealth Consumer Health Search Task 2018</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shadi Saleh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavel Pecina</string-name>
          <email>pecinag@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University Faculty of Mathematics and Physics Institute of Formal and Applied Linguistics</institution>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present our participation in CLEF Consumer Health Search Task 2018, mainly, its monolingual and multilingual subtasks: IRTask1 and IRTask4. In IRTask1, we use language-model based retrieval model, vector-space model and Kullback-Leiber divergence query expansion mechanism to build our runs. In IRTask4, we submitted 4 runs for each language of Czech, French and German. We follow query-translation approach in which we employ a Statistical Machine Translation (SMT) system to get a ranked list of translation hypotheses in English. We use this list for two systems: the rst one uses 1-best-list translation to construct queries, and the second one uses a hypotheses reranker to select the best translation (in terms of retrieval performance) to construct queries. We also present our term reranking model for query expansion, in which we deploy feature set from di erent resources (the document collection, Wikipedia articles, translation hypotheses). These features are used to train a logistic regression model that can predict the performance when a candidate term is added to a base query.</p>
      </abstract>
      <kwd-group>
        <kwd>Multilingual information retrieval</kwd>
        <kwd>statistical machine translation</kwd>
        <kwd>hypotheses reranking</kwd>
        <kwd>term reranking</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Internet searches for medical topics had been increasing recently, and have
gotten the attention of information retrieval researchers. Fox [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] reported that about
80% of Internet users in the United States look for medical information online.
The main challenge in the medical information retrieval systems that people
with di erent experience express their information need in di erent way [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
Laypeople express their medical information need using non-medical terms, while
medical experts tend to use advanced medical terms, thus, information retrieval
systems need to be stable for such di erent query variations. The signi cant
increasing of non-English digital content on the World Wide Web has been followed
by an increase in looking for this information by internet users. Grefenstette and
Nioche [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] presented an estimation of language size in 1996, late 1999 and early
2000 for documents captured from the internet. Their study showed that the
English content has grown 800%, German 1500%, and Spanish 1800% in the
same period. Furthermore, users started to look for information needs that is
represented in documents which are not available in their native languages.
      </p>
      <p>
        The system that searches for information in a language di erent from the
one of user is called Cross-Lingual (multilingual) Information Retrieval (CLIR)
system. It enables users to write queries (information need) represented in a
language (lang. A), and returns results from a document collection written in
a di erent language (lang. B). Usually, the baseline system in CLIR is to take
1-best-list translations which are returned by a statistical machine translation
(SMT) system and perform the retrieval as shown in the CLEF eHealth
Information Retrieval tasks before [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Nikoulina et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] presented an approach
to develop Cross-lingual information retrieval (CLIR) system which is based on
reranking the hypotheses given from the SMT system. Saleh and Pecina [20]
considered Nikoulina's work as a starting point and expanded it by adding a
rich set of features for training. They presented approach covered translating
queries from Czech, French and German into English and rerank the alternative
translations to predict the hypothesis that gives better CLIR performance.
      </p>
      <p>In this paper, we describe our participation at the CLEF 2018 eHealth
consumer health search task [23]. We focus in our participation in the multilingual
IR Task. We present our machine learning model which reranks the alternative
translations given by the machine translation system for better IR results. We
also present our new approach to expand translated queries using our machine
learning model.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Task Description</title>
      <p>
        CLEF eHealth Consumer Health Search Task 2018 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is similar to the IR tasks
in the previous years (2013{2017). The participants this year are required to
retrieve relevant web pages from the provided document collection in response to
users' queries. These queries represent information need in the medical domain.
The IR task consists of IRTask1 which is a standard ad-hoc monolingual search
task. IRTask2 is a similar task of the personalised search task in 2017 [
        <xref ref-type="bibr" rid="ref16 ref7">16, 7</xref>
        ], the
retrieved documents are personalised to match user expertise (how likely the user
is able to understand the content of the retrieved documents). IRTask 3 contains
query variations for the same information need, and the participants have to
design a search system that is steady when the same information need is expressed
in di erent query variations. In the multilingual ad-hoc search task (IRTask4 ),
the monolingual English queries were translated by experts into Czech, French
and German, and the participants are asked to design a search system to retrieve
relevant documents to these queries from the English document collection.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Document Collection</title>
        <p>Document collection in the CLEF 2018 consumer health search task is created
using CommonCrawl platform 1. First, the query set (described in Section 2.2)
is submitted to Microsoft Bing APIs, and a list of domains is extracted from the
top retrieved results. This list is extended by adding reliable health websites,
at the end clefehealth2018 B (which we use in this work) contained 1; 653 sites,
after excluding non-medical websites such as news websites. After preparing the
domain list, these domains are crawled and provided as an indexed collection
to the participants. Two indexes are provided, in the rst one, documents are
stemmed and a stop-word list is used, while no preprocessing is done in the
second index. The collection contains 5; 560; 074 documents, the stemmed
index contains 14; 213; 903 vocabularies, while the non-stemmed index contains
15; 298; 904 ones.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Queries</title>
        <p>
          The query set this year includes 50 English queries. This set is a subset of 150
medical queries that were created from HON and TRIP query logs within the
Khresmoi project [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Table 1 shows the average number of terms in the 50 test
queries in all languages. Although the average number of terms in the English
queries is 5:64, there are queries that are much longer (e.g. query 199001 ), as
shown in Table 2. Queries might contain typos since they are constructed from
real query logs, as shown in query 175001, which contains Emugel instead of
Emulgel.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The training data</title>
      <p>
        The data that we use to train our systems was presented by the CLEF eHealth
2014 Task 3 - Information Retrieval [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and CLEF eHealth 2015 Task 2:
UserCentred Health Information Retrieval [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. It is almost identical to the
collection used in CLEFeHealth 2013 Task 3 - User-Centred Health Information
Retrieval, which contained a few additional documents which were excluded from
the 2014/2015 collection due to license issues. The document collection includes
a total of 1,104,298 web pages in HTML, automatically crawled from various
English medical websites such as Genetics Home Reference, ClinicalTrial.gov
and Diagnosia. To clean the HTML pages in the collection, we follow the work
of Saleh and Pecina [19]. The queries have also been adopted from the CLEF
eHealth series and include all the test queries from the IR task of 2013 (50
queries), 2014 (50 queries), and 2015 (66 queries). We joined them to create a
more representative and balanced sample for IR experiments. The set of all 166
queries was split into 100 queries for training and 66 queries for testing. The
two sets are strati ed in terms of distribution of the year of origin, number of
relevant/not-relevant documents, and query length (number of words).
4
4.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <sec id="sec-4-1">
        <title>Translation system</title>
        <p>
          For the multilingual task (IRTask4 ), we follow the query translation approach,
in which a query is translated into the collection language (English), then the
retrieval is conducted. Query translation approach reduces the task into
monolingual task (both queries and documents are expressed in the same language).
We use Khresmoi statistical machine translation (SMT) system [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], for
language pairs: Czech-English, French-English and German-English, to translate
the queries into English. Khresmoi SMT system was trained to translate queries,
and tuned on parallel and monolingual data taken from the medical domain
resources like Wikipedia, UMLS concept descriptions and UMLS metathesaurus.
Such domain speci c data made Khresmoi perform better when translating
sentences in the medical domain like the queries in our case. Generally, feature
weights in SMT systems are tuned toward BLEU [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], a method for automatic
evaluation of SMT systems correlates with human judgments. It is not
necessary to have correlation between the quality of general SMT system and the
quality of CLIR performance [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]; therefore Khresmoi SMT system was tuned
using MERT [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] towards PER (position-independent word error rate), because
it does not penalise word reorder; which is not important for the performance
of IR systems.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Hypotheses reranking</title>
        <p>
          Khresmoi SMT system produces a list of ranked translations in the target
language, for each sentence in the source language, this list is called n-best-list.
However, this n-best-list is ranked based on the translation quality rather than
the retrieval performance. Saleh and Pecina [20] presented an approach to rerank
an n-best-list and predict a translation that gives the best retrieval performance
in terms of P@10. The reranker is a generalized linear regression model that
uses a set of features which can be divided according to their sources into: 1)
The SMT system: This includes features that are derived from the verbose
output of the Khresmoi SMT system (e.g. phrase translation model, the
target language model, the reordering model and word penalty). 2) Document
collection: This includes IDF scores and features that are based on the
blindrelevance feedback approach. 2) External resources: Resources like Wikipedia
articles and UMLS metathesaurus [22] are employed to create a rich set of
features for each query hypothesis. 3) Retrieval status value (RSV): RSV is the
score of the retrieval scoring function when constructing a query from a
translation hypothesis. It helps to involve more information from the collection in the
reranking process by assigning to each hypothesis the score from the retrieval
function. This feature is based on the work of Nottelman et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], where they
investigated the correlation between RSV and relevance probability. To train the
model, we join the training and test sets that we presented in Section 3 in one
set, then calculate feature values from each language, and merge them from all
seven languages in one training set. The test set is the CLEF eHealth 2018 query
set in Czech, French and German.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Query Expansion</title>
        <p>
          Query expansion is a process that reformulates user's initial queries as an
attempt to represent more information to improve retrieval performance
eventually. In this section, we present our approach to reformulate user's query in the
CLIR task using machine learning model, based on the presented work of Saleh
and Pecina [21]. This approach is based on expanding a query by adding
candidate terms from an existing pool. This is done by reranking candidate terms
using machine learning model towards better IR performance and adding the
top ranked terms to the original query. To create a pool of candidate terms for
each query, we use two main resources:
{ Translation hypotheses: This pool is built by merging n-best-list
translations for each query, after ltering stopwords and terms that already
appeared in the 1-best-list translation.
{ Wikipedia titles: First, we index English Wikipedia articles (titles and
abstracts without any preprocessing) using Terrier [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] and its implementation
of Dirichlet language model as an IR model, then we conduct retrieval for
each query's 1-best-list translation from this index, then the top 10 ranked
Wikipedia articles are selected and their titles are added to the pool.
To train the model, we use the training data that we presented in Section 3,
while for testing, we use the provided queries from the CLEF eHealth 2018 IR
task in Czech, French and German languages. After building a pool of candidate
terms, we generate the following features for each term:
{ IDF The inverse document frequency which is calculated from the relevant
document collection.
{ Translation pool frequency This feature represents how many times a
term appeared in the translation pool. When a term appears in multiple
hypotheses, this means that the probability of being a relevant translation
to one of the terms in the original query is high.
{ Wikipedia frequency The frequency of a term in the top 10 retrieved
Wikipedia articles. Retrieval is conducted using the 1-best-list translation
for the query that we want to expand with the candidate term.
{ Retrieval Status Value di erence To calculate this feature, we conduct
two retrievals, the rst one using the original query (1-best-list translation
), and the second one using the original query expanded with the candidate
term, then we take the score of the highest ranked document in each
retrieval and calculate the di erence between them. This feature tells us the
contribution of the candidate term to the retrieval status value.
{ Similarity To calculate the similarity between a candidate term tm and
the query terms, we use a trained model of word2vec embeddings on 25
millions articles from PubMed 2. First, we get the word embeddings for each
term in the original query and we sum these embeddings to get a vector
that represents the entire query. Then we take the embeddings for tm, and
calculate the cosine similarity between the query vector and tm vector.
{ Co-occurrence frequency The co-occurrences of a candidate term tm and
the query terms ti 2 Q indicates how likely tm is related to the original
query Q. We sum up the co-occurrence frequency for each term in query Q
and the candidate term tm in all documents dj in the collection C, as shown
in the Equation 1.
        </p>
        <p>co(tm; Q) =</p>
        <p>X
dj2C;ti2Q
tf (dj ; ti)tf (dj ; tm)
(1)
{ Term frequency First, we perform retrieval from the collection using a
query that is constructed from the 1-best-list translation, then we calculate
the term frequency of a candidate term tm in the top 10 ranked documents
from the retrieval result.
{ Medical term count This feature represents how many times a term
appeared in the UMLS lexicon, as an attempt to give more weight to the
medical terms.</p>
        <p>Our goal is to design a model that can predict the performance of the retrieval
when expanding a query with a term from the terms pool, and add terms that
can improve the performance. To train the model, we perform the following
steps:
{ Generate a pool of candidate terms for each query in the training and test
set.
2 https://www.ncbi.nlm.nih.gov/CBBresearch/Wilbur/IRET/DATASET/
{ Add one term from the pool to the query that we want to expand (1-best-list
translation) and perform the retrieval using the baseline system (Dirichlet
model)
{ Calculate the feature values for each term as we described above.
{ For training queries, we evaluate the performance for each expanded query
considering P @10 as a main metric, P @10 being the objective function for
our model.
{ Merge training queries from the 7 languages together to enrich the training
set with more instances.
{ After preparing the training set, we normalise feature values using standard
scaling by removing the mean and scaling them to have unit variance. This
is done independently on each feature, then we use the scaler coe cient to
standardise the test set. Scaling is important since the range of the feature
values varies widely.</p>
        <p>The term reranker is a generalised linear regression model which predicts P @10
value for each term when expanding the original query with, we choose the term
that has the highest predicted value of P @10.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Systems</title>
      <p>We submit runs for the monolingual task (IRTask1) and the multilingual task
(IRTask4), as we present in the following sections.
5.1</p>
      <sec id="sec-5-1">
        <title>Monolingual system</title>
        <p>
          In the monolingual task, we submit four runs:
{ Run 1 In this system, we use the Terrier's index that is provided by the
organisers without applying any data preprocessing. Terrier's implementation
of Dirichlet smoothing language model is used as the retrieval model with
its default parameters.
{ Run 2 This system also uses the same retrieval model as in Run 1, while
as an index, we use Terrier's index that uses Porter-stemming method and
English stop-word list.
{ Run 3 This system uses Terrier's implementation of TF-IDF model, for the
purpose of comparing between a vector-space model and an LM model (the
one that is used in Run 1), we use the same index as in Run 1.
{ Run 4 In this run, we use Terrier's implementation of Kullback-Leiber
divergence (KLD) [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] for query expansion, with number of top documents is
set to 10 and number of terms for expansion is set to 3. These 3 terms are
selected as following: rst, an initial retrieval is done using the base query
and the top 10 documents are chosen as pseudo-relevant documents. Then
each term in these documents is scored as shown in Equation 2, where Pr(t)
is the probability of term t in the pseudo-relevant documents (these
documents are treated as a bag-of-words), and Pc(t) is probability of term t in
the document collection c. Finally the top 3 scored terms are added to the
base query and a nal retrieval is done using the new expanded query.
Pr(t)
        </p>
        <p>Pc(t)
Score(t) = Pr(t) log
(2)
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Cross-lingual system</title>
        <p>{ Run 1 In this run, we translate the queries in the source languages into
English and get 1-best-list translations. Retrieval is conducted using Dirichlet
model, and non-stemmed index. The same retrieval settings are used in the
following runs.
{ Run 2 This run uses hypotheses reranking approach, in which each query is
translated into English and from the 15-best-list translations, the 1-best-list
(in terms of IR quality) translation is selected for the retrieval as described
in Section 4.2
{ Run 3 First we translate the queries into English and the 1-best-list that is
produced by the SMT system is chosen as a base query, then this query is
expanded by one term using the term reranking approach that is presented
in Section 4.3
{ Run 4 This run is similar to Run1, the only di erence is that Google
Translate 3 is used to translate the queries into English.
Table 3 shows the percent of similar documents that are retrieved (among the
highest 10 ranked ones) by di erent runs. It is clear from the table that
different approaches tend to retrieve di erent documents, for example, run 3 uses
query expansion based approach. Query expansion means that a query will be
expanded by more terms to include more information, leading to retrieve di
erent documents, that is the reason why this run has the lowest similarity to the
other runs. Both of run 1 and run 4 use 1-best-list translation from two di erent
machine translation systems (Khresmoi and Google Translate respectively) to
3 translate.google.com
construct the queries. This explains why these two systems share similar
documents more than all other systems. Run 2 uses hypotheses reranking approach to
select best translation to be used for retrieval, while run1 uses 1-best-list
translation as it is selected from the SMT system to construct queries. According
to further analysis we performed between the di erence between the retrieved
documents by these two runs, we found that 23 queries (out of 50) have 100%
similarity of the top 10 retrieved documents, and this correlates with what was
shown by Saleh and Pecina [20], that an SMT system fails in 50% of the cases
to select the best translation to perform the best performance for the retrieval.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>We presented our participation in CLEF eHealth Consumer Health Search Task
2018 (monolingual and multilingual subtasks). Four runs were submitted to the
monolingual task, two runs use a language-model IR with Dirichlet smoothing,
they di er in the used index (one uses a stemmed index and one uses an
index without stemming). As for the multilingual task, we submitted four runs
for each language of Czech, French and German. The rst one uses 1-best-list
translation from a statistical machine translation system, the second run uses
hypotheses translations reranking, the third run is an implementation of query
expansion using term reranker model, while the last run uses Google Translate
to translate the provided queries into English. Our results analysis shows that
similar approaches tend to share more similar retrieved documents than di erent
approaches.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments References</title>
      <p>This research was supported by the Czech Science Foundation (grant n. P103/12/G084).
19. Saleh, S., Pecina, P.: CUNI at the ShARe/CLEF eHealth Evaluation Lab 2014. In:
Working Notes of CLEF 2015 - Conference and Labs of the Evaluation forum. vol.
1180, pp. 226{235. She eld, UK (2014)
20. Saleh, S., Pecina, P.: Reranking hypotheses of machine-translated queries for
crosslingual information retrieval. In: Experimental IR Meets Multilinguality,
Multimodality, and Interaction. The 7th International Conference of the CLEF
Association, CLEF 2016. pp. 54{66. Springer, Evora, Portugal (2016)
21. Saleh, S., Pecina, P.: Task3 patient-centred information retrieval: Team CUNI. In:
Working Notes of CLEF 2017 - Conference and Labs of the Evaluation Forum,
CEUR Workshop Proceedings. vol. 1866. Dublin, Ireland (2017)
22. Schuyler, P.L., Hole, W.T., Tuttle, M.S., Sherertz, D.D.: The UMLS
Metathesaurus: representing di erent views of biomedical concepts. Bulletin of the Medical
Library Association 81(2), 217 (1993)
23. Suominen, H., Kelly, L., Goeuriot, L., Kanoulas, E., Azzopardi, L., Spijker, R., Li,
D., Neveol, A., Ramadier, L., Robert, A., Palotti, J., Jimmy, Zuccon, G.: Overview
of the CLEF eHealth evaluation lab 2018. In: CLEF 2018 - 8th Conference and
Labs of the Evaluation Forum. Lecture Notes in Computer Science LNC, Springer,
Avignon, France (2018)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amati</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carpineto</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romano</surname>
          </string-name>
          , G.:
          <article-title>Query di culty, robustness, and selective application of query expansion</article-title>
          .
          <source>In: European conference on information retrieval</source>
          . pp.
          <volume>127</volume>
          {
          <fpage>137</fpage>
          . Springer (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Dusek</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hlavacova</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosa</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , et al.:
          <article-title>Machine translation of medical texts in the Khresmoi project</article-title>
          .
          <source>In: Proceedings of the Ninth Workshop on Statistical Machine Translation</source>
          . pp.
          <volume>221</volume>
          {
          <fpage>228</fpage>
          . ACL, Baltimore, USA (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Health Topics:
          <volume>80</volume>
          %
          <article-title>of internet users look for health information online</article-title>
          .
          <source>Tech. rep.</source>
          , Pew Research Center (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamon</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robertson</surname>
          </string-name>
          , J.:
          <source>D7</source>
          .
          <article-title>2 Meta-analysis of the rst phase of empirical and user-centered evaluations</article-title>
          .
          <source>Tech. rep. (August</source>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mueller</surname>
          </string-name>
          , H.:
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2014</year>
          ,
          <article-title>Task 3: Usercentred health information retrieval</article-title>
          .
          <source>In: Proceedings of CLEF 2014</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>22</fpage>
          . Springer, She eld,
          <source>UK</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nevaol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the CLEF eHealth evaluation lab 2015</article-title>
          .
          <source>In: The 6th Conference and Labs of the Evaluation Forum</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>15</fpage>
          . Springer, Berlin, Germany (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nvol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spijker</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>CLEF 2017 eHealth evaluation lab overview</article-title>
          .
          <source>In: CLEF 2017 - 8th Conference and Labs of the Evaluation Forum, Lecture Notes in Computer Science (LNCS)</source>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Grefenstette</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nioche</surname>
          </string-name>
          , J.:
          <article-title>Estimation of english and non-english language use on the www</article-title>
          .
          <source>In: Content-Based Multimedia Information Access - Volume 1</source>
          . pp.
          <volume>237</volume>
          {
          <fpage>246</fpage>
          . RIAO,
          <article-title>Centre de hautes etudes internationales d'informatique documentaire</article-title>
          , Paris, France (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jimmy</surname>
            , Zuccon,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Overview of the CLEF 2018 consumer health search task</article-title>
          . In:
          <article-title>CLEF 2018 Evaluation Labs</article-title>
          and Workshop: Online Working Notes. CEUR-WS, Avignon, France (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Nikoulina</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovachev</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagos</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Adaptation of statistical machine translation model for cross-lingual information retrieval in a service context</article-title>
          .
          <source>In: Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics</source>
          . pp.
          <volume>109</volume>
          {
          <fpage>119</fpage>
          .
          <string-name>
            <surname>Avignon</surname>
          </string-name>
          , France (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Nottelmann</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuhr</surname>
          </string-name>
          , N.:
          <article-title>From retrieval status values to probabilities of relevance for advanced IR applications</article-title>
          .
          <source>Information retrieval 6</source>
          , 363{
          <fpage>388</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Och</surname>
            ,
            <given-names>F.J.:</given-names>
          </string-name>
          <article-title>Minimum error rate training in statistical machine translation</article-title>
          .
          <source>In: Proceedings of the 41st Annual Meeting on Association for Computational Linguistics - Volume 1</source>
          . pp.
          <volume>160</volume>
          {
          <fpage>167</fpage>
          .
          <string-name>
            <surname>Sapporo</surname>
          </string-name>
          ,
          <string-name>
            <surname>Japan</surname>
          </string-name>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amati</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plachouras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macdonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Terrier: A high performance and scalable information retrieval platform</article-title>
          .
          <source>In: Proceedings of Workshop on Open Source Information Retrieval</source>
          . pp.
          <volume>18</volume>
          {
          <fpage>25</fpage>
          . ACM, Seattle, WA, USA (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Jr.</surname>
            ,
            <given-names>C.E.K.</given-names>
          </string-name>
          :
          <article-title>How users search and what they search for in the medical domain - understanding laypeople and experts through query logs</article-title>
          .
          <source>Inf. Retr. Journal</source>
          <volume>19</volume>
          (
          <issue>1-2</issue>
          ),
          <volume>189</volume>
          {
          <fpage>224</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lupu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <source>CLEF eHealth Evaluation Lab</source>
          <year>2015</year>
          ,
          <article-title>Task 2: Retrieving information about medical symptoms</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>22</fpage>
          .
          <string-name>
            <surname>Spriner</surname>
          </string-name>
          , Berlin, Germany (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimmy</surname>
            , Pecina,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lupu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>CLEF 2017 task overview: The IR Task at the eHealth evaluation lab</article-title>
          . In: Working Notes of Conference and
          <article-title>Labs of the Evaluation (CLEF) Forum</article-title>
          . CEURWS, Dublin, Ireland (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Papineni</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roukos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>W.J.:</given-names>
          </string-name>
          <article-title>BLEU: A method for automatic evaluation of machine translation</article-title>
          .
          <source>In: Proceedings of the 40th annual meeting on Association for Computational Linguistics</source>
          . pp.
          <volume>311</volume>
          {
          <fpage>318</fpage>
          . Philadelphia, USA (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Pecina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dusek</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hlavarova</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.</given-names>
          </string-name>
          , et al.:
          <article-title>Adaptation of machine translation for multilingual information retrieval in the medical domain</article-title>
          .
          <source>Arti cial Intelligence in Medicine</source>
          <volume>61</volume>
          (
          <issue>3</issue>
          ),
          <volume>165</volume>
          {
          <fpage>185</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>