<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Optimizing Search Results for Educational Goals: Incorporating Keyword Density as a Retrieval Objective</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rohail Syed</string-name>
          <email>rmsyed@umich.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kevyn Collins-Thompson</string-name>
          <email>kevynct@umich.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Information, University of Michigan</institution>
          ,
          <addr-line>105 S. State St., Ann Arbor MI, 48109</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>While past research has shown that learning outcomes can be influenced by the amount of e↵ort students invest during the learning process, there has been little research into this question for scenarios where people use search engines to learn. In fact, learning-related tasks represent a significant fraction of the time users spend using Web search, so methods for evaluating and optimizing search engines to maximize learning are likely to have broad impact. Thus, we introduce and evaluate a retrieval algorithm designed to maximize educational utility for a vocabulary learning task, in which users learn a set of important keywords for a given topic by reading representative documents on diverse aspects of the topic. Using a crowdsourced pilot study, we compare the learning outcomes of users across four conditions corresponding to rankings that optimize for di↵erent levels of keyword density. We find that adding keyword density to the retrieval objective gave significant learning gains on some topics, with higher levels of keyword density generally corresponding to more time spent reading per word, and stronger learning gains per word read. We conclude that our approach to optimizing search ranking for educational utility leads to retrieved document sets that ultimately may result in more ecient learning of important concepts.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The Web has become a primary source of online
information for learning-related tasks [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. While current Web search
engines are tuned to give fast, high-quality results for single
queries, they are optimized for generic relevance, not
learning outcomes: many tasks involving educational goals
require significant time and multiple queries to complete with
current Web search engines [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and ideally, personalized
retrieval that can exploit representations of user history and
learning goals to be most e↵ective. Developing a search
algorithm that is optimized for the learning process is a natural
prerequisite to encouraging more Web-based learning.
      </p>
      <p>
        Exploring new topic areas and learning important domain
Search as Learning (SAL), July 21, 2016, Pisa, Italy
The copyright for this paper remains with its authors. Copying permitted
for private and academic purposes.
vocabulary is one popular instance of a learning task [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Ideally, a retrieval algorithm optimized for this task would not
only be e↵ective at teaching a user the important keywords
for a given topic by finding highly relevant representative
documents, but also enable them to do so eciently. While
user e↵ort itself could be defined in many ways when ranking
search results based on factors such as reading diculty [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
or other text properties [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], we consider the total amount
of text to be read in the search result documents as a
simple proxy for e↵ort. Given a desired count of exposure for
each keyword, by returning documents with higher keyword
density per document, we obtain more ecient keyword
coverage, thus reducing eo↵rt by reducing the total amount of
text that needs to be read. Thus, we explore the role of
keyword density as a component of educational retrieval.
      </p>
      <p>Toward that goal, the main contributions of this work are
a novel search algorithm that re-ranks for optimized
educational utility using keyword density as a proxy for eo↵rt,
and a study that evaluates the e↵ectiveness of this approach
on actual learning outcomes.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        While research on ranking algorithms to maximize the
relevance of generic or personalized search results is
wellestablished, few studies have focused on algorithms that can
optimize results with utility for an educational goal as the
retrieval objective. Researchers have recognized the
importance of going beyond traditional retrieval evaluation
measures to consider user progress over time [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] as well as degree
of eo↵rt [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], but little, if any, of that work has involved
learning assessment. Eickho↵ et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] investigated learning
behaviors of Web search users, but used only indirect evidence
via implicit indicators derived from Web search logs, rather
than direct assessment of users. They also did not develop
or assess new retrieval algorithms that could be adapted to
improve learning outcomes. Collins-Thompson et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
incorporated a form of e↵ort criterion into Web search ranking
by incorporating reading diculty as a personalized
ranking feature, but did not assess its e↵ectiveness for actual
learning outcomes. Similarly, Raman et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] showed how
‘intrinsically diverse’ (ID) sessions for exploring and
learning about a new, specific topic could be identified and
supported using a new diversity-based retrieval algorithm, but
without assessing learning outcomes. A subsequent study
by Collins-Thompson et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] examined the e↵ectiveness of
ID results presentation on actual high- and low-level
learning outcomes. We build on both of these previous studies
by exploring a modified variant of the ID algorithm in the
context of a vocabulary learning task.
      </p>
    </sec>
    <sec id="sec-3">
      <title>METHOD</title>
      <p>Our retrieval approach has three stages: (1) given a topic
expressed as a query, selecting appropriate aspects to be
learned for each topic, with each aspect represented by a
keyword, (2) for each aspect (keyword), determining the
total frequency with which the keyword should occur in the
retrieval results, and (3) developing a retrieval algorithm for
vocabulary learning that finds documents to ‘cover’ the
selected keywords eciently by including the keyword density
of the documents as an adjustable sub-objective.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Selecting Topic Aspects</title>
      <p>For each topic in our study, we manually collected a set of
exemplar Web documents D⇤ that were deemed to be
representative of useful knowledge about that topic. We then
represented the vocabulary learning goal for a given topic
as a weighted set K = {k1, . . . , kN } of keywords, which we
call the target keywords, derived from the topic’s exemplar
set. For this study, we chose the top N = 10 most
representative keywords for each topic, using a measure of term
frequency weighted by inverse term log-frequency in a global
corpus. As di↵erent aspects of a topic may have greater or
less relevance in understanding the topic, each keyword is
assigned an associated weight wi, where w are the parameters
of a multinomial distribution estimated from the frequency
counts of the keywords in the representative set D⇤ . Table 1
shows the top 5 out of 10 keywords generated for each topic,
along with their relative weight wi (in parentheses).
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Determining Total Words to Read</title>
      <p>We assume that a student’s knowledge of each topic
keyword ki monotonically increases with each instance of it that
they read. Now let T be the total keywords the learner reads.
The distribution of T among the N keywords will be
proportional to the importance of each keyword, given by wi.
Then, if si is the total instances of ki the learner reads, we
have: si = T · wi.</p>
      <p>Ideally, a student would learn the most with unlimited
instances of each keyword (T = 1 ). However, in reality a
student’s time and e↵ort will limit the amount of training
they experience, so the T value for each topic was
manually chosen to produce small document sets (less than 15
documents).
3.3</p>
    </sec>
    <sec id="sec-6">
      <title>Developing the Retrieval Algorithm</title>
      <p>
        As a baseline retrieval algorithm, we used the intrinsic
diversity algorithm developed by Raman et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], since it
was designed to provide optimal exploration of topics with
multiple sub-aspects. The intrinsic diversity objective
essentially rewards high quality documents from relevant and
representative subtopics, while penalizing redundant
documents and subtopics1. To account for user e↵ort, we added
a new sub-objective term (e↵✏ i ) to the existing intrinsic
diversity objective that influences the keyword density (and
thus, the eciency of keyword coverage) for results:
|D|
arg max X Rel(di|q) · Rel(di|qi) · eDiv (qi,Q) · e↵✏ i
D i=1
(1)
1We chose operational parameter settings
= 10, = 0.2
where the topic we want to teach is given by the base query q,
D is the resulting document set, Div(qi, Q) is a redundancy
penalty, qi is the ith sub-topic query and Rel(di|qi) is the
reciprocal rank of document di in the results page returned
for query qi.
      </p>
      <p>With this extension, setting ↵ = 0 recovers the original
intrinsic diversity results, while higher values of ↵ result in
document sets with increasingly dense keyword coverage.</p>
      <p>More specifically, ✏i is the normalized contribution that
document di o↵ers in terms of how much closer it brings
the student towards reading the total required number of
keyword instances (the sj counts). Let CD represent the
cumulative keyword counts the student has seen so far from
documents higher in the ranking, and Ci represents the
keyword frequency distribution of di. Then we have:
✏i =
1
|di| j=1</p>
      <p>XN ⇢ Cij
max(0, sj</p>
      <p>Cij + CDj  sj
CDj ) otherwise
The term ✏i e↵ectively is a measure of the keyword density
in di with respect to the target keywords for the topic. By
rewarding documents that have higher density, via the choice
of a higher ↵ setting, we enable the learner to reach the
target sj counts faster.</p>
      <p>Our implementation of the intrinsic diversity algorithm
determines the base query’s sub-topics by analyzing the
corresponding Wikipedia article on that query’s topic. It
generates sub-topic queries by extracting the main header topics
in the article and appending them to the base query. For
example, for the query “DNA”, some sub-topic queries were:
“DNA Properties” and “DNA Biological functions”. We then
fetch the top 70 Google search results for the base query and
the top 70 results for each of the sub-topics queries and run
optimization problem (1).</p>
      <p>We intend to refine our subtopic extraction methods to
generalize beyond those available in Wikipedia topics in
future work. In general, many di↵erent variables can
simultaneously influence learning. Some students may learn better
with multimedia aids, some will learn better with pure text
documents, some will benefit from more technically-worded
documents and so on. In this paper, we will specifically
evaluate only Web documents that contain only text and,
at most, supplementary pictures.
4.</p>
    </sec>
    <sec id="sec-7">
      <title>EVALUATION</title>
      <p>
        To assess the potential e↵ect on learning outcomes of
retrieved documents optimized using die↵rent levels of
keyword density (choices of ↵ ), we ran a crowdsourced user
study that involved a vocabulary learning task: learning the
target keywords. Participants first completed a
multiplechoice pre-test to assess their existing knowledge of the
keywords, then based on the condition, read through a
provided retrieval set of documents containing the keywords
to be learned, and then completed an immediate post-test
to assess their updated keyword knowledge. We ran five
separate crowdsourced jobs corresponding to five di↵erent
topics selected to cover a range of scientific topics: Igneous
rocks (geology), Tundra (environmental science), DNA
(genetics), Cytoplasm (biology) or GSM (telecommunications).
For each of these topic jobs, a participant was randomly2
assigned to one of four di↵erent keyword density conditions,
2Participants were sorted into conditions based on
Crowdflower’s random assignment to tasks.
corresponding to ↵ settings of [
        <xref ref-type="bibr" rid="ref1">0, 80, 120, 1</xref>
        ]. The ↵ = 1
condition simply means that we give full weight only to the
keyword density ✏i term and ignore all other terms in the
ID retrieval objective.
      </p>
      <p>The pre- and post- vocabulary tests consisted of a series of
multiple-choice questions, one for each of the K keywords.
Both the pre- and post-reading tests were constructed with
identical questions so that we could investigate the
participants’ learning gain for each vocabulary term by looking at
the di↵erence in scores 3.</p>
      <p>We used the Crowdflower platform for this study.
Participants were oe↵red US$0.04 per page (the equivalent of
US$3.20/hr) for completing the tasks. For quality control, in
addition to Crowdflower’s proprietary mechanisms and ‘gold
standard’ questions, we limited the participant pool to users
from the U.S. and Canada, given the vocabulary-centric
nature of the task and reliance on English reading skills. We
also o↵ered the tasks only to workers in the highest quality
(level 3) pool, and only kept responses from those workers
who spent at least four minutes on the task.</p>
      <p>The particular set of documents shown to each participant
was based on which ↵ condition they were assigned. We
gathered data for 35 participants per ↵ condition per topic,
resulting in a total of 140 participants per topic and 700
participants overall. After excluding those who didn’t pass
the test questions and those who didn’t complete the full
task, we ended up with 616 total participants.</p>
    </sec>
    <sec id="sec-8">
      <title>RESULTS</title>
      <p>Overall, our analysis showed that die↵rent choices of ↵
were in fact associated with di↵erences in learning, as
measured by both absolute and normalized gains from pre-test
to post-test.</p>
      <p>We first analyzed learning gains (sum of learning gains for
all K keywords) across the four ↵ conditions. Retrieval
results incorporating higher keyword density gave statistically
3In measuring ‘learning gain’, we assume no memory loss so
the learning gain is always non-negative.</p>
      <p>Keyword Difficulty Low High
0.20
significant mean learning gains for two out of the five
topics4 (Table 2). Both of these topics showed a peak learning
gain at the ↵ = 80 condition, suggesting that a
combination of lowering eo↵rt via the keyword density parameter
and rewarding intrinsic diversity in documents o↵ers better
learning gains than either factor alone. However, we also
found that the setting of ↵ = 120 yielded the worst learning
gains in those same topics. This suggests that the learning
gains are quite sensitive to the particular choice of ↵ and
that choosing an ↵ that combines both the ID objective and
the keyword density objective is not always going to improve
learning utility. It’s not entirely clear why the specific value
of ↵ = 80 oe↵red better performance but we intend to
investigate this further and how to algorithmically choose ↵
in future work, using an extended set of topics.</p>
      <p>Since the target keywords ranged from more familiar to
more technical, and learning gains could be expected to
interact with keyword diculty, we faceted the learning gain
results by low- and high-diculty keyword categories 5.
Figure 1 shows the result of averaging the learning gains for
each keyword in the two diculty categories and then
averaging the results across the five topics. We see that there
were learning gains in all conditions for both low- and
highdiculty keywords, but as expected, learning gains were
higher for the higher-diculty (and thus initially less
familiar) keywords.
4For all ANOVA analysis reported, the same significance
ranges were found using bootstrapped ANOVA and the
Kruskal-Wallis test.
5Keywords were split into two groups of five keywords
according to their age of acquisition (AoA) score in a standard
psychometric database. If a keyword didn’t have an AoA
score, it was assumed to be maximum diculty.</p>
      <p>Next, as a measure of learning eciency, we evaluated
absolute learning gain normalized by the total words read.
This measure incorporates e↵ort such that, for two students
scoring the same absolute gain, the one who achieved this
gain with less eo↵rt (reading less text) is rewarded more.
ANOVA analysis of the di↵erent ↵ levels shows that most
topics had strongly significant di↵erences in means. There
was a general trend of increasing gains with increasing ↵ and
several topics achieved maximum gains at ↵ = 1 (Table 3).</p>
      <p>We note that one topic, Cytoplasm, showed an opposite
trend where higher alpha values mostly lead to worse
normalized learning gains. We hypothesize that this may be
because the total number of words used in each condition
for Cytoplasm were significantly lower (almost half as many
for ↵ = 0 and ↵ = 80) compared to the four other topics. It
is thus possible that the positive impact of choosing higher
↵ values is only e↵ective after passing a certain threshold of
minimum reading material.
5.2</p>
    </sec>
    <sec id="sec-9">
      <title>Learning Gains per Unit Time</title>
      <p>When considering learning gains per unit time (Table 4)
instead of per word, the results were much less conclusive:
for example, two topics showed significant di↵erences in mean
learning per time, but with opposite extremes of ↵ values (0
and 1 ). To better understand the factors a↵ecting learning
gain per unit time (denoted T imLe ), consider the following
decomposition:
α = 80 α = 120</p>
      <p>Alpha Penalties
α = infty</p>
      <p>This relationship is visualized in Figure 2, with WT iomrdes on</p>
      <p>L
the x-axis and W ords on the y-axis. As the plot makes
evident, there is a positive correlation (r=.42, p=.06) between
these two components. However, while the slope of this
approximately linear relationship (which is exactly T imLe ,
learning per unit time), is relatively stable across conditions,
there are very die↵rent tradeo↵ regimes depending on the
value of ↵ : the ↵ = 0 condition is characterized by some
of the shortest reading times per word and lowest learning
gains per word, while the ↵ = 1 condition is characterized
by the highest times and learning gains. Thus, while the
overall learning gain per unit time (ratio of the two
components) may not change dramatically across conditions, the
underlying two components, representing the tradeo↵ users
choose between reading time and learning eciency, vary
greatly as keyword density changes greatly.</p>
    </sec>
    <sec id="sec-10">
      <title>5.3 Image Coverage vs. Keyword Density</title>
      <p>
        To gain more insight into why pages with increased
keyword density might contribute to more ecient learning, we
investigated additional properties of the page content that
might be correlated with keyword density. We found that
while few result documents made use of multimedia such as
animations, audio or video, a number did use images to
supplement the text. Thus, the picture superiority e↵ect [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], in
which people tend to remember things better when they see
pictures rather than words, could be relevant, since we were
testing fact-based learning, which relies at least partially on
recall. We thus examined whether there was a relationship
between image coverage – defined as total images divided by
total words – as a function of ↵ . We determined the number
of relevant images manually for each page, excluding
irrelevant images such as navigation icons and advertisements.
We found that pages with higher keyword density did indeed
tend to have increased image coverage, as shown in Fig. 3.
For three of the five topics, the highest image coverage is in
the ↵ = 1 condition.
      </p>
      <p>We consider the possibility that a heavier coverage of
images in teaching documents can improve learning outcomes
regardless of condition. There is partial evidence of this in
that ANOVA analysis of the topics “Igneous rock”,
“Tundra” and “DNA” showed no statistical significance in means
(Table 2) and these three topics had the top three average
image coverage (.0024, .0026 and .0034 respectively). On
the other hand, the two topics that showed significant
differences (“Cytoplasm” and “GSM”) had the lowest coverage
(.0015 and .0006 respectively). As such, it is possible that
a higher image coverage can collectively improve or worsen
learning gains regardless of conditions. Determining if the
presence or absence of images actually has such an ee↵ct
warrants further investigation.</p>
      <p>We observe informally that pages using a higher density
of keywords tend to be those that give an overview of topic
for instructional purposes, and thus are more likely to be
supplemented with images by the author. We intend to
investigate this phenomenon and other content properties that
may interact with learning in future work.</p>
    </sec>
    <sec id="sec-11">
      <title>6. CONCLUSION</title>
      <p>We introduced a novel algorithm for optimizing Web search
results for a learning-oriented objective – a vocabulary
learning task – by extending intrinsically diverse ranking to
incorporate a keyword density sub-objective. This keyword
density was controlled by a parameter ↵ that rewarded
documents containing a high density of topic-relevant keywords.
The result was an algorithm that not only gave relevant,
diverse results to explore new topics, but also emphasized
ecient keyword coverage in the results content, thus allowing
learners to potentially expend less e↵ort toward their
learning goal. We hypothesized that changing the keyword
density ↵ would be associated with positive changes in users’
vocabulary learning outcomes. We tested this hypothesis
with a crowdsourced pilot study based on five topics, across
four conditions that varied keyword density by using
die↵rent values of ↵ . We found that for some topics participants
did in fact show stronger learning gains per word with
nonzero ↵ settings. Of the four topics that showed significant
di↵erences of means, three were maximized at ↵ = 1 . This
is an interesting finding as the ↵ = 1 condition only
considers the keyword density as its objective which means that
our findings suggest that a search algorithm that is blind to
the rank or implicit quality of a document is o↵ering better
results than an algorithm that explicitly considers such
measures. We also examined learning gains per word and per
unit time, finding that users showed very di↵erent
tradeo↵s between reading time per word and learning gains per
word in low- vs high keyword density conditions. In future
work we intend to explore criteria for selecting optimal
operational settings of ↵ , and to incorporate more personalized
components in the retrieval model.</p>
      <p>Acknowledgements This research was supported in part
by the Institute of Education Sciences, U.S. Department of
Education, through Grant R305A140647 to the University
of Michigan. The opinions expressed are those of the
authors and do not represent views of the Institute or the U.S.
Department of Education.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bailey</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosenick</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reinholdtsen</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salada</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            <given-names>H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Wong</surname>
            <given-names>S.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>User task understanding: a web search engine perspective</article-title>
          .
          <source>In NII Shonan</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Collins-Thompson</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bennett</surname>
            <given-names>P. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            <given-names>R. W.</given-names>
          </string-name>
          , Chica S., de la, and
          <string-name>
            <surname>Sontag</surname>
            <given-names>D.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Personalizing Web Search Results by Reading Level</article-title>
          .
          <source>In Proceedings of the 20th ACM International Conference on Information and Knowledge Management (CIKM '11)</source>
          . ACM, New York, NY, USA,
          <fpage>403</fpage>
          -
          <lpage>412</lpage>
          . DOI:http://dx.doi.org/10.1145/2063576.2063639
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Collins-Thompson</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rieh</surname>
            <given-names>S. Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haynes</surname>
            <given-names>C. C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Syed</surname>
            <given-names>R.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Assessing Learning Outcomes in Web Search: A Comparison of Tasks and Query Strategies</article-title>
          .
          <source>In Proceedings of the 2016 ACM on Conference on Human Information Interaction and Retrieval (CHIIR '16)</source>
          . ACM, New York, NY, USA,
          <fpage>163</fpage>
          -
          <lpage>172</lpage>
          . DOI:http://dx.doi.org/10.1145/2854946. 2854972
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>De</given-names>
            <surname>Angeli</surname>
          </string-name>
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Coventry</surname>
          </string-name>
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Johnson</surname>
          </string-name>
          <string-name>
            <given-names>G.</given-names>
            , and
            <surname>Renaud</surname>
          </string-name>
          <string-name>
            <surname>K.</surname>
          </string-name>
          <year>2005</year>
          .
          <article-title>Is a picture really worth a thousand words? Exploring the feasibility of graphical authentication systems</article-title>
          .
          <source>International Journal of Human-Computer Studies 63</source>
          ,
          <issue>1</issue>
          (
          <year>2005</year>
          ),
          <fpage>128</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Eickho</surname>
            ↵
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teevan</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Dumais</surname>
            <given-names>S.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Lessons from the Journey: A Query Log Analysis of Withinsession Learning</article-title>
          .
          <source>In Proceedings of the 7th ACM International Conference on Web Search and Data Mining (WSDM '14)</source>
          . ACM, New York, NY, USA,
          <fpage>223</fpage>
          -
          <lpage>232</lpage>
          . DOI:http: //dx.doi.org/10.1145/2556195.2556217
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Raman</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bennett</surname>
            <given-names>P. N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Collins-Thompson</surname>
            <given-names>K.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Toward Whole-session Relevance: Exploring Intrinsic Diversity in Web Search</article-title>
          .
          <source>In Proceedings of the 36th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '13)</source>
          . ACM, New York, NY, USA,
          <fpage>463</fpage>
          -
          <lpage>472</lpage>
          . DOI:http://dx.doi.org/10.1145/2484028. 2484089
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Smucker</surname>
            <given-names>M. D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Clarke</surname>
            <given-names>C. L.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Time-based Calibration of Ee↵ctiveness Measures</article-title>
          .
          <source>In Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '12)</source>
          . ACM, New York, NY, USA,
          <fpage>95</fpage>
          -
          <lpage>104</lpage>
          . DOI:http://dx.doi.org/10.1145/ 2348283.2348300
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Yilmaz</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verma</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Craswell</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radlinski</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>and Bailey P.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Relevance and E↵ort: An Analysis of Document Utility</article-title>
          .
          <source>In Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management (CIKM '14)</source>
          . ACM, New York, NY, USA,
          <fpage>91</fpage>
          -
          <lpage>100</lpage>
          . DOI:http://dx.doi.org/10.1145/2661829.2661953
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>