<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Integrating Understandability in the Evaluation of Consumer Health Search Engines</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guido Zuccon</string-name>
          <email>g.zuccon@qut.edu.au</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bevan Koopman</string-name>
          <email>bevan.koopman@csiro.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Copyright is held by the author/owner(s).</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Australian e-Health Research Centre, CSIRO</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>MedIR July 11</institution>
          ,
          <addr-line>2014, Gold Coast, Australia., ACM SIGIR.</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Queensland University of Technology</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <fpage>32</fpage>
      <lpage>35</lpage>
      <abstract>
        <p>In this paper we propose a method that integrates the notion of understandability, as a factor of document relevance, into the evaluation of information retrieval systems for consumer health search. We consider the gain-discount evaluation framework (RBP, nDCG, ERR) and propose two understandability-based variants (uRBP) of rank biased precision, characterised by an estimation of understandability based on document readability and by different models of how readability influences user understanding of document content. The proposed uRBP measures are empirically contrasted to RBP by comparing system rankings obtained with each measure. The findings suggest that considering understandability along with topicality in the evaluation of information retrieval systems lead to different claims about systems effectiveness than considering topicality alone.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Searching for health advice on the Web is an increasingly
common practice. A recent research has found that in 2012
about 58% of US adults (72% of all US Internet users – 66%
in 2011) have consulted the Internet for health advice [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ];
of these, 77% have used search engines like Google, Bing,
or Yahoo! to gather health information, while only 13%
have started their health information seeking activities from
specialised sites such as WebMD. It is, therefore, crucial to
create and evaluate information retrieval (IR) systems that
specifically support consumers searching for health advise
on the Web. In this paper, we focus on the evaluation of IR
systems for consumer health search.
      </p>
      <p>
        Previous studies within health informatics have
investigated online consumer health information beyond
topicality to specific health topics; in particular, with respect to
the understandability and reliability of such information.
For example, Wiener and Wiener-Pla [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] have investigated
the readability (measured by the SMOG reading index [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ])
of Web pages concerning pregnancy and the periodontium
as retrieved by Google, Bing and Yahoo!. Walsh and
Volsko [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] have shown that most online information sampled
from five US consumer health organisations and related to
the top 5 medical related causes of death in US is
presented at a readability level (measured by the SMOG, FOG
and Flesch-Kincaid reading indexes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) that exceeds that of
the average US citizen (7th grade level). Ahmed et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
have highlighted the variability in readability (measured by
the Flesch Reading Ease and the Flesch-Kincaid reading
index [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) and quality of concussion information accessed
through Google searches. The understandability and
reliability of online health information has been considered as a
critical issue for supporting online consumer health search
because (1) consumers may not benefit from health
information that is not provided in an understandable way; and
(2) the provision of unreliable, misleading or false
information on a health topic, e.g., a medical condition or treatment,
may led to negative health outcomes. This previous research
suggests that topicality should not be considered as the only
relevance factor for assessing the effectiveness of IR systems
for consumer health search: other factors, such as
understandability and reliability, should also be included in the
evaluation framework.
      </p>
      <p>
        Research on the user perception of document relevance
has shown that users’ relevance assessments are affected by
a number of factors beyond topicality, although topicality
has been found to be the essential relevance criteria. For
example, Xu and Chen proposed and validated a five-factor
model of relevance which consists of novelty, reliability,
understandability, scope, along with topicality [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Their
empirical findings highlight the importance of
understandability, reliability and novelty along with topicality in the
relevance judgements they collected. Nevertheless, typical
evaluation of IR systems commonly considers only relevance
assessments in terms of topicality1; this is also the case when
evaluating systems for consumer health search, for example,
within CLEF eHealth 2013 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In this paper, we aim to
close this gap in the evaluation of IR systems and focus on
integrating understandability along with topicality for the
evaluation of consumer health search engines. The
integration of other factors influencing relevance, such as reliability,
are left for future work.
      </p>
      <p>
        The integration of understandability within the
evaluation methodology is achieved by extending the general
gaindiscount framework synthesised by Carterette [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]; this
framework encompasses the widely-used nDCG, RBP and ERR.
The result is a series of understandability-biased evaluation
measures. Specifically, we examine one such measure, the
understandability-based rank biased precision (uRBP) – a
variant of rank biased precision (RBP) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]; variants of nDCG
1With the recent exception of novelty and diversity, e.g., [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
and ERR may also be derived within our framework.
      </p>
      <p>
        The proposed evaluation measure is further instantiated
by considering specific estimations of understandability based
on readability measures computed for each retrieved
document. While understandability encompasses other aspects
in addition to text readability (e.g., prior knowledge), the
use of readability measures is a good first approximation for
understandability. This choice is also supported by prior
work in health informatics regarding understandability of
consumer health information (e.g., see [
        <xref ref-type="bibr" rid="ref1 ref10 ref11">11, 10, 1</xref>
        ]).
      </p>
      <p>
        The impact of the proposed framework and the specific
resultant measures on the evaluation of IR systems is
investigated in the context of the consumer health search task of
CLEF eHealth 2013 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]; empirical findings show that systems
that are most effective according to uRBP are not
necessarily as effective when considering topicality alone (i.e. RBP).
      </p>
    </sec>
    <sec id="sec-2">
      <title>UNDERSTANDABILITY-BASED EVALU</title>
    </sec>
    <sec id="sec-3">
      <title>ATION</title>
    </sec>
    <sec id="sec-4">
      <title>The gain-discount framework</title>
      <p>
        We tackle the problem of jointly evaluating topicality and
understandability for measuring IR system effectiveness within
the gain-discount framework synthesised by Carterette [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Within this framework, the effectiveness of a system,
conveyed by a ranked list of documents, is measured by the
evaluation measure M , defined as:
(1)
      </p>
      <p>
        N k=1
where g(k) and d(k) are respectively the gain and discount
function computed for the document at rank k,2 K is the
depth of assessment at which the measure is evaluated, and
1/N is a (optional) normalisation factor, which serves to
bound the value of the sum into the range [
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ] (see [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]).
      </p>
      <p>
        Different measures developed within the gain-discount
framework are characterised by different instantiations of its
components. For example, the discount function in RBP is
modelled by d(k) = βk−1, where β ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] reflects user
behaviour (high values representing persistent users, low values
representing impatient users); while in nDCG the discount
function is given by d(k) = 1/(log2(1 + k)) and in ERR by
d(k) = 1/k. Similarly, instantiations of gain functions differ
depending upon the considered measure. In RBP, the gain
function is binary-valued (i.e., g(k) = 1 if the document at
rank k is relevant, g(k) = 0 otherwise); while for nDCG
g(k) = 2r(k) − 1 and for ERR g(k) = (2r(k) − 1)/2rmax (with
r(k) being the relevance grade of the document at rank k).
      </p>
      <p>
        Without loss of generality, we can express the gain
provided by a document at rank k as a function of its probability
of relevance; for simplicity we shall write g(k) = f (P (R|k)),
where P (R|k) is the probability of relevance given the
document at rank k. Note that a similar form has been used for
the definition of the gain function for time-biased evaluation
measures [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The specific instantiations of g(k) in measures
like RBP, nDCG and ERR can be seen as the application of
different functions f (.) to estimations of P (R|k).
      </p>
      <p>Traditional TREC-style relevance assessors are instructed
to consider topicality as the only (explicit) factor influencing
2For simplicity of notation, in the following we override k to
represent the rank position k, or the document at rank k: the context
of use will determine the meaning of k.
relevance, thus P (R|k) = P (T |k), i.e., the probability that
the document at k is topically relevant (to a query).
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Integrating understandability</title>
      <p>
        As discussed by previous work, e.g. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], relevance is
influenced by many factors; topicality being only one of them
– although the most important. To integrate
understandability into the gain-discount framework, we model P (R|k)
as the joint P (T , U |k), i.e. the probability of relevance of a
document (at rank k) is estimated using the joint probability
of the document being topical and understandable.
      </p>
      <p>To compute the joint probability we assume that
topicality and understandability are compositional events and their
probabilities independent, i.e., P (T , U |k) = P (T |k)P (U |k).
This is a strong assumption and its limitations are briefly
discussed in Section 4. Following this assumption, the gain
function in the gain-discount framework is expressed as:
g(k) = f (P (R|k)) = f P (T |k)P (U |k)
(2)</p>
      <p>Different evaluation measures that may be developed within
this framework would instantiate f P (T |k)P (U |k) in
different ways. In the following we will propose two RBP-based
instantiations; other instantiations are left for future work.
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Estimating understandability</title>
      <p>In the traditional TREC settings, assessments about the
topicality of a document to a query are collected through
manual annotation of query-document pairs from assessors
(i.e., binary or graded relevance assessments3); these are
then turned into estimations of P (T |k). This process may
be mimicked to collect understandability assessments; in this
paper however we do not explore this possibility. Instead, we
explore the possibility of computing understandability as a
property of a document and integrate this in the evaluation
process, along with standard relevance assessments. To this
aim, readability is used as a proxy for understandability.
(The limitations of this choice are briefly noted in Section 4.)
Its use is however justifiable because readability is one of the
aspects that influence the understanding of text.</p>
      <p>
        To estimate readability (and thus understandability), we
employ established general readability measures as those
used in [
        <xref ref-type="bibr" rid="ref1 ref10 ref11">1, 10, 11</xref>
        ], e.g., SMOG, FOG and Flesch-Kincaid
reading indexes. These measures consider the surface level
of the text contained in Web pages, that is, wording and
syntax of sentences. In this framework, the presence of long
sentences, words containing many syllables and unpopular
words, are all indicators of difficult text to read [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In this
paper, we use the FOG measure to estimate the readability
of a text; the FOG reading level is computed as
F OG(d) = 0.4 ∗ (avgslen(d) + phw(d))
(3)
where avgslen(d) is the average length of sentences in a
document d and phw(d) is the percentage of hard words
(i.e., words with more than two syllables) in d.
      </p>
      <p>
        The use of such general readability measures to assess the
readability of documents concerning health information has
been questioned [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] as these do not seem to adequately
correlate with human judgments for documents in this
domain [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Nevertheless, the adoption of standard
readabil3Recall that although called “relevance assessments,” in
TRECstyle assessments, annotators are usually instructed to consider
only the topicality of a document to a query, isolating this factor
from others influencing relevance in real settings.
      </p>
      <p>cno
0.15 lite</p>
      <p>
        frrcyscoeoo
0.10 iilts
ysedbaoned
fr
0.05 it
ity measures in this paper is a first step towards
demonstrating the use of the proposed understandability biased
measures and analyse how system rankings would change
accordingly. In addition, their usage is partially supported
by previous work within health informatics on assessing the
readability of online health advice [
        <xref ref-type="bibr" rid="ref1 ref10 ref11">1, 10, 11</xref>
        ].
2.4
      </p>
    </sec>
    <sec id="sec-7">
      <title>Modelling P(U|k)</title>
      <p>Given the readability score for a document at rank k,
P (U |k) needs to be estimated; this is achieved by
considering user models that encode different ways in which a user
is affected by document readability.</p>
      <p>
        We first consider a user model P1(U |k) where a user is
characterised by a readability threshold th and every
document that has a readability score below th is considered
certainly understandable, i.e., P1(U |k) = 1; while documents
with readability above th are considered not understandable,
i.e. P1(U |k) = 0. This is a (Heaviside) step function centred
in th; this function is depicted in Figure 1 (P1(U |k)) with
th = 20, along with the FOG readability score distribution
for documents from CLEF e-Health 2013 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The use of a
step function to model P (U |k) is akin to the gain function
in RBP (also a step function). The understandability-based
RBP for user model one is then given by:
      </p>
      <p>
        To understand how accounting for understandability
influences the evaluation of IR systems tailored to searching
health advice on the Web, we consider the runs submitted to
the CLEF eHealth 2013 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which specifically aimed at
evaluating systems for this task. Our empirical experiments and
subsequent analysis specifically focus on the changes in
system rankings obtained when evaluating with standard
mea
      </p>
      <p>
        K sures (RBP) and understandability-based measures (uRBP1
uRBP1 = (1 − β) βk−1r(k)u1(k) (4) and uRBP2). System rankings are compared using Kendall
k=1 rank correlation (τ) and AP correlation [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] (τAP ), which
where, for simplicity of notation, u1(k) indicates the value weights higher rank changes that affect top systems. We do
of P1(U |k) and r(k) is the (topical) relevance assessment of not experiment with different values of β in RBP, and set
document k (alternatively, the value of P(T|k)); thus g(k) = β = .95 across RBP and uRBP.
f (P (T |k)P (U |k)) = P (T |k)P (U |k) = r(k)u1(k). The document collection used in CLEF eHealth 2013 has
      </p>
      <p>A second user model (P2(U |k)) is proposed, where the been retired due to removal of duplicates and copyrighted
probability estimation is similar to a step function, but smoothed documents; we thus use the CLEF eHealth 2014 collection
in the surroundings of the threshed value; this provides a (which is a subset of the CLEF eHealth 2013 collection) to
more realistic transition between readable and not-readable allow reproducibility of the reported results and the 2013
content: qrels for relevance assessment. For each document in the
1 arctan F OG(k) − th collection, the FOG readability scores (Equation 3) were
P2(U|k) ∝ 2 − π (5) cCoLmEpFuteeHde a–tlthhe2s0c1o3reqrdeilsstrisibsuhtoiwonn ifnorFaiglludreoc1u.mTehnrtese
itnhrtehseholds on the FOG readability values were explored for the
computation of the two alternative formulations of uRBP:
th = 10, 15, 20; documents with a FOG score below 10
should be near-universally understandable, while documents
with FOG scores above 15 and 20 increasingly restrict the
audience able to understand the text.</p>
    </sec>
    <sec id="sec-8">
      <title>3.2 Results and analysis</title>
      <p>where arctan is the arctangent trigonometric function and
F OG(k) is the FOG readability score of document at rank k;
other readability scores could be used instead of FOG. The
distribution of P2(U |k) values is shown in Figure 1.
Equation 5 is not a proper probability distribution, but this can
be obtained by normalising Equation 5 by its integral
between [min F OG(k) , max F OG(k) ]; however Equation 5
is rank equivalent to such distribution, not changing the
effect on the uRBP variant. These settings lead to the
formulation of a second understandability-based RBP, uRBP2,
based on the second user model, by simply substituting
u2(k) = P2(U |k) to u1(k) in Equation 4.</p>
      <p>Note that in both understandability-based measures (as
well as in the original RBP) the contribution of an irrelevant
document is zero, irrespective of its P (U |k). The
contribution (to the gain) of a relevant document with readability
score above th is 1 for RBP, 0 for uRBP1 and less than 0.5
for uRBP2 (for uRBP2 the score will quickly tend to 0 the
more the readability score is above the threshold value).</p>
      <p>Finally, note that it is possible to design other user
models representing how readability inuflences document
understandability; the challenge is to determine which model
better represents the relationship between readability and
document understanding.</p>
      <p>Figure 2 reports RBP vs. uRBP of IR systems
participating to CLEF eHealth 2013 for the two user models proposed
in Section 2.4 and for the three readability thresholds
considered in the experiments. Similarly, Table 1 reports the
values of Kendall rank correlation (τ) and AP correlation
(τAP ) between system rankings obtained with RBP and the
two versions of uRBP.</p>
      <p>Higher correlation between systems rankings obtained with
RBP and uRBP is observed for higher values of th,
irrespectively of uRBP version (see Table 1). This is expected as
the higher the threshold, the more documents will be
characterised by a P (U |k) = 1 (or ≈ 1 for uRBP2), thus reducing
uRBP to RBP. The fact that in general uRBP2 is correlated
with RBP more than uRBP1 is to RBP highlights the effect
of smoothing obtained by the arctan function; specifically,
the increase of readability scores for which P (U |k) is not zero</p>
      <p>RBP RBP
Figure 2: RBP vs. uRBP of CLEF eHealth 2013
systems (left: uRBP1; right: uRBP2) for varying values of
threshold on the readability scores (th = 10,15,20).</p>
      <p>RBP vs.
uRBP1
RBP vs.
uRBP2
th = 10
τ = .1277
τAP = −.0255</p>
      <p>τ = .5887
τAP = .2877
beyond th narrows the scope for ranking differences between
systems effectiveness. These observations are confirmed in
Figure 2, where only few changes in the rank of systems are
shown for th = 20 (× in Figure 2), while more changes are
found for th = 10 (◦) and th = 15 (+). Note that the small
differences in the absolute values of effectiveness recorded
by uRBP with th = 10 should not be interpreted as a lack
of discriminative power. When th = 10 only 1.4% of the
documents in the CLEF eHealth 2013 qrels are relevant and
readable, thus contributing to uRBP.</p>
      <p>Figure 2 demonstrates the importance of considering
understandability along with topicality in the evaluation of
systems for the considered task. The system ranked highest
according to RBP (MEDINFO.1.3.noadd) is second to a
number of systems according to uRBP if user understandability
of up to FOG level 15 is wanted. Specifically, the
highest uRBP1 for th = 10 is achieved by UTHealth_CCB.1.3.
noadd, which is ranked 28th according to RBP, and for
th = 15 by teamAEHRC.6.3, which is ranked 19th
according to RBP and achieves the highest uRBP2 for th = 10, 15.</p>
    </sec>
    <sec id="sec-9">
      <title>LIMITATIONS AND CONCLUSIONS</title>
      <p>In this paper, we have investigated how understandability
can be integrated in the gain-discount framework for
evaluating IR systems. The approach studied here is general
and can be adopted to other factors of relevance, such as
reliability. Information reliability plays an important role in
consumer health advice search; its integration will be
studied in future work.</p>
      <p>
        In the proposed approach, the relevance (P(R|k)) was
modelled as the joint probability P (T , U |k). This joint
probability was assumed to be independent and the two events
to be compositional, thus allowing to derive P (T , U |k) =
P (T |k)P (U |k) and to treat topicality and understandability
separately. This is a strong assumption and it is not
necessarily true; alternatives are under investigation, e.g. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The approach was demonstrated by deriving
understandability-based variants of RBP; other measures can also be
extended, e.g., nDCG and ERR. Note, however, that
nDCGstyle versions would require normalising the gain function by
the ideal gain, which in turns requires finding the optimal
ranking based on two criteria, relevance score and
understandability, instead of one as in the standard nDCG.</p>
      <p>
        Xu and Chen [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] have noted that factors of relevance
influence relevance assessments in different proportions, e.g.,
in their study, topicality was found to be more influential
than understandability. The specific uRBP measures
studied here did not consider this aspect; however weighting of
different factors could be accomplished through a different
f (.) function for converting P (T , U |k) into gain values.
      </p>
      <p>
        In this paper, we have used readability as a proxy for
understandability, but this is only one aspect that influences
understandability [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]; future work may explore other
factors, e.g., users’ prior knowledge, as well as the presence of
images that further explain the textual information.
Furthermore, readability was estimated using general, surface
level readability measures. Previous work has shown that
these measures are often not suitable to evaluate the
readability of health information. For example, Yan et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
claim that people experience the highest readability
difficulties at word level rather than at sentence level; they further
propose a new metric based on concept-based readability,
specifically instantiated in the health domain. A number of
alternative approaches that measure text readability beyond
the surface characteristics of text have been proposed.
Future work will investigate their use to estimate P(U|k), along
with actual readability assessments collected from users.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>O. H.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Sullivan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Schneiders</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. R.</given-names>
            <surname>McCrory</surname>
          </string-name>
          .
          <article-title>Concussion information online: evaluation of information quality, content and readability of concussion-related websites</article-title>
          .
          <source>British journal of sports medicine</source>
          ,
          <volume>46</volume>
          (
          <issue>9</issue>
          ):
          <fpage>675</fpage>
          -
          <lpage>683</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Bruza</surname>
          </string-name>
          , G. Zuccon, and
          <string-name>
            <given-names>L.</given-names>
            <surname>Sitbon</surname>
          </string-name>
          .
          <article-title>Modelling the information seeking user by the decision they make</article-title>
          .
          <source>In MUBE 2013</source>
          , pages
          <fpage>5</fpage>
          -
          <lpage>6</lpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Carterette</surname>
          </string-name>
          .
          <article-title>System effectiveness, user models, and user utility: a conceptual framework for investigation</article-title>
          .
          <source>In SIGIR'11</source>
          , pages
          <fpage>903</fpage>
          -
          <lpage>912</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Craswell</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Soboroff.</surname>
          </string-name>
          <article-title>Overview of the TREC 2009 web track</article-title>
          .
          <source>In TREC'09</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fox</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Duggan</surname>
          </string-name>
          . Health online
          <year>2013</year>
          .
          <source>Tech. Rep., Pew Research Center's Internet &amp; American Life Project</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leveling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          , H. Mu¨ller, S. Salanter¨a, H. Suominen, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon.</surname>
          </string-name>
          <article-title>Share/clef ehealth evaluation lab 2013, task 3: Information retrieval to address patients' questions when reading clinical reports</article-title>
          .
          <source>In CLEF</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. R.</given-names>
            <surname>McCallum</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Peterson</surname>
          </string-name>
          .
          <article-title>Computer-based readability indexes</article-title>
          .
          <source>In ACM'82 Conf.</source>
          , pages
          <fpage>44</fpage>
          -
          <lpage>48</lpage>
          ,
          <year>1982</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moffat</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Zobel</surname>
          </string-name>
          .
          <article-title>Rank-biased precision for measurement of retrieval effectiveness</article-title>
          .
          <source>TOIS</source>
          ,
          <volume>27</volume>
          (
          <issue>1</issue>
          ):
          <fpage>2</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Smucker</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Clarke</surname>
          </string-name>
          .
          <article-title>Time-based calibration of effectiveness measures</article-title>
          .
          <source>In SIGIR'12</source>
          , pages
          <fpage>95</fpage>
          -
          <lpage>104</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Walsh</surname>
          </string-name>
          and
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Volsko</surname>
          </string-name>
          .
          <article-title>Readability assessment of internet-based consumer health information</article-title>
          .
          <source>Respiratory care</source>
          ,
          <volume>53</volume>
          (
          <issue>10</issue>
          ):
          <fpage>1310</fpage>
          -
          <lpage>1315</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Wiener</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Wiener-Pla</surname>
          </string-name>
          .
          <article-title>Literacy, pregnancy and potential oral health changes: The internet and readability levels</article-title>
          .
          <source>Maternal and child health journal</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y. C.</given-names>
            <surname>Xu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Relevance judgment: What do information users consider beyond topicality?</article-title>
          <source>JASIST</source>
          ,
          <volume>57</volume>
          (
          <issue>7</issue>
          ):
          <fpage>961</fpage>
          -
          <lpage>973</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Song</surname>
          </string-name>
          , and
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Concept-based document readability in domain specific information retrieval</article-title>
          .
          <source>In CIKM'06</source>
          , pages
          <fpage>540</fpage>
          -
          <lpage>549</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E.</given-names>
            <surname>Yilmaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Aslam</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Robertson</surname>
          </string-name>
          .
          <article-title>A new rank correlation coefficient for information retrieval</article-title>
          .
          <source>In SIGIR'08</source>
          , pages
          <fpage>587</fpage>
          -
          <lpage>594</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>