<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>3 !European)Workshop)on)! Human"Computer)Interaction)and)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mixed)Reality)Lab)</string-name>
          <email>max.wilson@nottingham.ac.uk) )</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>The)Royal)School)of)Library)and))</string-name>
          <email>blar@iva.dk)</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dept.)of)Computer)&amp;)Systems)Sciences)</string-name>
          <email>preben@dsv.su.se) )</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information)Science,)Denmark)</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Stockholm)University</institution>
          ,
          <addr-line>)</addr-line>
          <country country="SE">Sweden)</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University)of)Nottingham</institution>
          ,
          <addr-line>)</addr-line>
          <country country="UK">UK)</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>5</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>EuroHCIR)2013)was)organised)with)the)specific)goal)of)better)engaging)the)IR) community,)who)have)been)underrepresented)at)previous)EuroHCIR) conferences.)Thus)we)proposed)to)have)the)workshop)at)the)ACM)SIGIR) conference)in)Dublin.)Research,)Industry,)and)Position)papers)were)invited,)and) although)very)few)industry)submissions)were)received,)we)received)a)number)of) research)and)position)papers)focusing)on)the)intersection)of)IR)and)HCI) evaluations,)several)focusing)on)adapting)the)TREC)paradigm.)Many)interesting) system)and)demonstrator)papers)were)also)accepted.)</p>
      </abstract>
      <kwd-group>
        <kwd>Information*Retrieval!</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Proceedings+of+the+!
rd</p>
    </sec>
    <sec id="sec-2">
      <title>Preface!</title>
    </sec>
    <sec id="sec-3">
      <title>Organised!by!</title>
      <sec id="sec-3-1">
        <title>Max"L."Wilson"</title>
      </sec>
      <sec id="sec-3-2">
        <title>Birger"Larsen"</title>
      </sec>
      <sec id="sec-3-3">
        <title>Preben"Hansen"</title>
      </sec>
      <sec id="sec-3-4">
        <title>Tony"RussellFRose"</title>
      </sec>
      <sec id="sec-3-5">
        <title>Kristian"Norling"</title>
        <p>Page"3"F"" Fading"Away:"Dilution"and"User"Behaviour"(Orally)Presented)"</p>
        <p>Paul%Thomas,%Falk%Scholer,%Alistair%Moffat%
Page"7"F"" Exploratory"Search"Missions"for"TREC"Topics"(Orally)Presented)"</p>
        <p>Martin%Potthast,%Matthias%Hagen,%Michael%Völske,%Benno%Stein%
Page"11"F"" Interactive"Exploration"of"Geographic"Regions"with"WebFbased"Keyword"
Distributions"</p>
        <p>Chandan%Kumar,%Dirk%Ahlers,%Wilko%Heuten,%Susanne%Boll%
Page"15"F"" Inferring"Music"Selections"for"Casual"Music"Interaction"(Orally)Presented)"</p>
        <p>Daniel%Boland,%Ross%McLachlan,%Roderick%MurrayESmith"
Page"19"F"" Search"or"browse?"Casual"information"access"to"a"cultural"heritage"collection"</p>
        <p>Robert%Villa,%Paul%Clough,%Mark%Hall,%Sophie%Rutter%
Page"23"F"" Studying"Extended"Session"Histories"</p>
        <p>Chaoyu%Ye,%Martin%Porcheron,%Max%L.%Wilson%
Page"27"F" Comparative"Study"of"Search"Engine"Result"Visualisation:"Ranked"Lists"Versus"
Graphs"</p>
        <p>Casper%Petersen,%Christina%Lioma,%Jakob%Grue%Simonsen"</p>
        <sec id="sec-3-5-1">
          <title>Position!Papers!</title>
          <p>Page"31"F"" Evolving"Search"User"Interfaces"(Orally)Presented)"</p>
          <p>Tatiana%Gossen,%Marcus%Nitsche,%Andreas%Nürnberger"
Page"35"F"" A"Pluggable"WorkFbench"for"Creating"Interactive"IR"Interfaces"(Orally)Presented)"</p>
          <p>Mark%M.%Hall,%Spyros%Katsaris,%Elaine%Toms"
Page"39"F"" A"Proposal"for"UserFFocused"Evaluation"and"Prediction"of"Information"Seeking"
Process"(Orally)Presented)"</p>
          <p>Chirag%Shah%
Page"43"F"" Directly"Evaluating"the"Cognitive"Impact"of"Search"User"Interfaces:"a"TwoF
Pronged"Approach"with"fNIRS"</p>
          <p>Horia%A.%Maior,%Matthew%Pike,%Max%L.%Wilson,%Sarah%Sharples%
Page"47"F"" Dynamics"in"Search"User"Interfaces"</p>
          <p>Marcus%Nitsche,%Florian%Uhde,%Stefan%Haun%and%Andreas%Nürnberger%</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>Demo!Descriptions!</title>
          <p>Page"51"F"" SearchPanel:"A"browser"extension"for"managing"search"activity"</p>
          <p>Simon%Tretter,%Gene%Golovchinsky,%Pernilla%Qvarfordt""
Page"55"F"" A"System"for"PerspectiveFAware"Search"</p>
          <p>M.%Atif%Qureshi,%Arjumand%Younus,%Colm%O’Riordan,%Gabriella%Pasi,%Nasir%Touheed"
Fading Away: Dilution and User Behaviour</p>
          <p>Paul Thomas</p>
          <p>CSIRO ICT Centre
paul.thomas@csiro.au</p>
          <p>Falk Scholer
School of Computer Science
and Information Technology</p>
          <p>RMIT University</p>
          <p>Alistair Moffat
Department of Computing and</p>
          <p>Information Systems</p>
          <p>The University of Melbourne
falk.scholer@rmit.edu.au
ammoffat@unimelb.edu.au
ABSTRACT
When faced with a poor set of document summaries on the first page
of returned search results, a user may respond in various ways: by
proceeding on to the next page of results; by entering another query;
by switching to another service; or by abandoning their search. We
analyse this aspect of searcher behaviour using a commercial search
system, comparing a deliberately degraded system to the original
one. Our results demonstrate that searchers naturally avoid selecting
poor results as answers given the degraded system; however, the
depth of the ranking that they view, their query reformulation rate,
and the amount of time required to complete search tasks, are all
remarkably unchanged.</p>
          <p>Categories and Subject Descriptors
H.3.4 [Information Storage and Retrieval]: Systems and
software—performance evaluation.</p>
          <p>General Terms
Experimentation, measurement.
Retrieval experiment, evaluation, system measurement.</p>
          <p>INTRODUCTION</p>
          <p>While carrying out a search, users have a number of tactics
available to them. Intuitively, it seems likely that these tactics or
behaviours will vary based on the quality of the results that are
returned by the retrieval system. For example, other things being
equal, a user who cannot find any relevant items on the first page
of search results might be more inclined to reformulate their query
(by entering another query into the search interface) than a user who
has found a large number of relevant items. Possible tactics when
using an apparently ineffective system include:
1. Looking further in the results list, visiting pages beyond the</p>
          <p>first, hoping that the results improve;
Presented at EuroHCIR2013. Copyright © 2013 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
2. Submitting another query, hoping for better results;
3. Switching to a different search engine and entering the same</p>
          <p>query, hoping that it provides better results;
4. Trying to find the information through other techniques, for</p>
          <p>example by browsing.</p>
          <p>We investigate the first two possibilities, reporting on differences
in user behaviour when a standard retrieval system is compared to an
adjusted system in which results are diluted by inserting non-relevant
answers. Our results indicate that searchers remained attentive to the
task in the degraded system, and adapted their behaviour to avoid
clicking on non-relevant snippets. However, all other aspects of
their behaviour were remarkably consistent, including the amount of
time spent on tasks; the number of query reformulations undertaken;
and their perceptions of search difficulty.
2.</p>
          <p>METHODS</p>
          <p>
            We designed a user experiment to explore ways in which
behaviour changes with retrieval quality. A total of n = 34 participants,
comprising staff and students from the Australian National
University, carried out six search tasks of differing complexity, covering
the remember, analyse and understand tasks of Wu et al. [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] but
modified for our context. On commencing a task, users were shown
a result page for an initial “starter” query that was constant across
users. They were then free to explore the results list, including being
able to open documents, to view further results pages, and to enter
follow-up queries. Once any document was opened for viewing,
participants were asked to indicate whether or not it was relevant to
their search task, before returning to the search results listing. The
search interface prevented tabbed browsing, and while a document
was being viewed it replaced the results page. Participants were not
given an explicit time limit for any task, but were told they could
move on when they felt ready.
          </p>
          <p>
            The search results displayed to participants were sourced from
the Yahoo! API, and presented in the usual way as an ordered list
consisting of query-biased summaries, with ten results per page. No
branding from the underlying search service was shown. Without
telling our participants, we simulated search systems of two
different effectiveness levels by showing results in one of two modes:
full, where the ranking obtained from the search service was
displayed in its original form; and diluted, where the original results
were interleaved with answers from a related but incorrect query [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ].
          </p>
          <p>Dilution was operationalised by leveraging the capacity-enhancing
(and obfuscatory) power of “management-speak”: the original
stakeholder information need was actioned going forward by enhancing
it through the win-win inclusion of a jargon competency chosen
randomly from a list of outside-the-box strategies, thereby
disempowering the results paradigm. For example, if the task was to “find
0
3
.
0
5
2
.
0
0
2
n .
itropo .150
roP .00
1
0
5
0
.
0
0
0
.
0
the Eurovision Song Contest home page”, a user’s initial full query
might be “eurovision”; whereas in the diluted system half of the
results displayed might instead be derived from the query
“eurovision best practice”. There were a small number of queries issued
for which it was not possible to generate five such results; these 22
out of 5930 page interactions are excluded from the analysis below.</p>
          <p>Most interactions with the search system were logged while
participants carried out the six search tasks, including: submitted search
queries; clicks on snippets in order to open documents for
viewing; assessments of document usefulness; and the point of gaze
on the screen, captured using an eye tracker. Task order was
balanced across the participants and topics so as to minimise the risk
of bias; similarly, whether the full or diluted approach got applied
for each participant-task combination was pre-determined as part of
the experimental design.</p>
          <p>RESULTS</p>
          <p>User behaviour, and the differences caused by the full and diluted
query treatments, can be measured in a range of ways.</p>
          <p>
            User click behaviour: The normalised click frequency at each rank
position in the answer pages is shown in Figure 1. In the diluted
retrieval system the “incorrect but plausible” documents were
inserted in positions 1, 3, 5, 7 and 9. The pattern of click behaviour
demonstrates that our experimental manipulation was successful:
for the full search results, the click distribution follows the expected
pattern of users clicking more frequently on items that are higher
in the ranked list [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ], whereas users of the diluted system were less
likely to click answer items in the odd positions. Note that position
bias – the propensity for searchers to select items that occur higher
in a ranking, possibly because they “trust” the underlying search
system [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ] – exists in both systems. In particular, all of the
oddnumbered rank positions in the diluted system are equally “bad”,
but participants still favoured items higher in the ranking.
          </p>
          <p>A second check to confirm that our system dilution had an impact
on search effectiveness is to consider the rates at which users saved
documents that they viewed (that is, the likelihood that a document
was found to be relevant after it was clicked). The mean rate is
0.733 for the full system, compared to 0.597 for the diluted system,
a statistically significant difference (t-test, p &lt; 0.05).</p>
          <p>While Figure 1 establishes that our user study participants
responded differently in terms of rank-specific click behaviour, the
high-level aggregated click behaviour across all participants and
search tasks was not distinctive: in total (all tasks, and all users)
there were 323 clicks for the full system, and 322 for the diluted
system. Unsurprisingly this difference is not statistically significant
1st results page</p>
          <p>2nd results page
full
diluted
(c 2 test, p = 0.97). The number of items that were determined as
being useful was also similar in the two conditions: 201 for full,
and 214 for diluted (c 2 test, p = 0.52). Our participants needed to
read a remarkably similar number of documents, and a remarkably
similar number of useful documents, to satisfy the (assigned) needs
regardless of the search system.</p>
          <p>Given this difference in click rates, it is reasonable to expect other
changes in behaviour and we consider this below.</p>
          <p>Depth of result page viewing: When presented with a search results
page, the user chooses which snippets require further evaluation.</p>
          <p>In line with commercial search engines, our experimental
participants were presented with ten answers per page, with the option of
accessing subsequent results pages.</p>
          <p>Faced with a relatively poor quality results list, a plausible strategy
for a user who is looking for an answer document is to look further
down the results page. Table 1 shows the frequency with which
results pages were viewed (that is, the user visited a results page
and looked at one or more items on the screen as recorded using
eye-tracking), summed across users and queries. When using the
full system, participants moved on to the second page of results for
15 out of 207 issued queries (with a corresponding mean page depth
of 1.07), while in the diluted system the second results page was
visited for 22 out of the total of 212 queries that were issued (a mean
page depth of 1.10). The difference in depth was not significant
(c 2 test, p = 0.34). No participants viewed results beyond the the
second page with either system.</p>
          <p>Figures 2 and 3 provide a more detailed view of gaze behaviour,
showing the deepest rank position that searchers examined while
carrying out a query, and the last rank position that was viewed
before finishing the query. The distributions of the lowest rank
positions viewed are similar between the full and diluted systems:
both show peaks at rank positions 7 (the last item above the fold)
and 10 (the last item in each page of search results). The distribution
of the last position viewed before finishing a query (which arises
when either enough relevant items have been found, or the user types
a fresh query) are also broadly similar. However, for the diluted
system, rank position 1 has a larger proportion of the probability
mass. A possible reason is that searchers mentally compare answers
as they view items in the results list, and most users scan at least the
top few items. The diluted system is likely to have a non-relevant
document in position one, and so reviewing that snippet may serve
as a final confirmation, before the user commits to a click on a
deeper-ranked snippet from the underlying full results.</p>
          <p>Query reformulation: A second way in which a user might respond
to search systems of differing quality is to change the rate at which
they stop looking through the current set of search results, and
instead enter a new query.</p>
          <p>The number of queries used by participants when carrying out
their search tasks is shown in Figure 4. Overall the number was
low for both systems, with a median of 1 and 2 queries (0 and 1
reformulations) for the full and diluted results, respectively. This
difference was not statistically significant (Wilcoxon signed-rank
test, p = 0.46).</p>
          <p>Ability to identify relevant answers: When a retrieval system serves
unhelpful answers, it might be that the ability of the searcher to
identify useful answers is similarly affected. However, based on our
experiments, the mean rate at which clicked items were saved as
being relevant was 0.787 for the full system and 0.747 for the diluted
system, showing no significant difference (t-test, p = 0.25). Thus
the ability of users to identify relevant answers, once documents
have been selected for viewing via their snippets, did not differ
between the experimental treatments.
4
1
2
1
irse 10
e
u
q
fro 8
e
b
m
uN 6
4
2
●
●
●
●
●
●
●
Ful</p>
          <p>Diluted
●
●
●
●</p>
          <p>Time spent on tasks: While depth of viewing and query
re-formulation do not show significant differences in searcher behaviour, it
could still be the case that using an inferior system makes querying
slower. Differences in system quality might alter the time spent
by users when viewing and processing result pages. However, the
average gaze duration when viewing snippets, measured as the sum
of fixation durations that occurred in the screen area defined by each
search result summary, was 0.586 second for full queries and 0.589
seconds for diluted queries. This difference was not statistically
significant (t-test, p = 0.89).</p>
          <p>Differences could also occur at a higher level of system
interaction. The mean time that participants spent working on each search
task, including viewing search result pages, viewing selected
documents, and making relevance decisions, was 2.70 minutes for the
full treatment, and 2.54 minutes for the diluted one. This difference
was not statistically significant (t-test, p = 0.62).</p>
          <p>Finally, we consider the interaction between time and query
reformulations. When using the full system, participants entered an
average of 1.50 queries per minute while completing each task.</p>
          <p>For the diluted system, the rate was 1.52 queries per minute. The
difference was not significant (t-test, p = 0.95).</p>
          <p>Overall, these results indicate that the quality of the search system
did not affect the rate at which participants were able to process
information on search results pages, or how much time they spent
working on tasks before feeling that they had achieved their goals.</p>
          <p>The only significant difference between the two treatments was the
click distribution, and the rate at which clicked documents were
judged to be useful.</p>
          <p>Searcher assessment of task difficulty: After carrying out each
search task, experimental participants were asked to answer two
questions: “How difficult was it to find useful information on this
topic?”, and “How satisfied were you with the overall quality of your
search experience?”. The 5-point response scale for these questions
was anchored with the labels “Not at all” (assigned a value of 1) and
“Extremely” (assigned a value of 5).</p>
          <p>Searchers found the tasks relatively easy to complete: the median
response rate for the search difficulty question was 2 for both the
diluted and full systems; this difference was not significant (Wilcoxon
test, p = 0.73). Satisfaction levels were also highly consistent
between the two systems, with a median response level of 4 for both
systems (Wilcoxon test, p = 0.91). Overall, there were no
systematic differences in participants’ perceptions of search difficulty or
the overall experience resulting from the two different treatments.</p>
          <p>DISCUSSION AND CONCLUSIONS</p>
          <p>It seems “obvious” that user behaviour will be influenced by the
quality of results that returned by a search service. Seeing many
poor results near the start of an answer list may influence the user’s
decision about whether to continue viewing subsequent answer
pages, to enter a new query, or to abandon the search altogether.</p>
          <p>
            Previous work has supported this view. For example, in a study of
36 users completing 12 search tasks with different search systems,
Smith and Kantor [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ] found that users adapted their behaviour:
when given a consistently degraded search system, they entered
more queries per minute than users of a standard system; similarly,
a higher detection rate (the ability to identify relevant answers) was
observed for users of degraded systems.
          </p>
          <p>However our study, in which 34 subjects carried out search tasks
using an evenly balanced combination of full and diluted search
systems, contrasts strongly with that intuition and previous findings.</p>
          <p>Overall, searchers took around the same amount of time to complete
their tasks in both experimental treatments; were able to save a
similar number of documents as being relevant; exhibited consistent
viewing behaviour when looking at the search results lists returned
by the treatments; and did not perceive significant differences in the
difficulty of carrying out tasks with both systems. The key difference
in participant behaviour was their click rate at particular ranks: in
essence, they successfully avoided poor answers, as demonstrated
by the shift in the click probability mass, shown in Figure 1.</p>
          <p>
            A possible explanation for the divergence in observed user
behaviour between the two studies may be the context in which the
searches were carried out. Participants in the Smith and Kantor
study were instructed to “find good information sources” for an
unspecified “boss”, with an incentive to find the most good and
fewest bad sources possible [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ]; participants were not constrained
in the amount of time that they could spend on a task. In contrast,
our subjects were instructed that they would complete a sequence of
. . . web search tasks and were advised to spend what feels to be an
appropriate amount of time on each task, until you have collected a
set of answer pages that in your opinion allow the information need
to be appropriately met. The overall expectations were therefore
different: in the Smith and Kantor study, participants were given the
goal of maximising relevance by finding as many good answers as
possible; in our study, participants were “satisficing”, having been
requested to decide for themselves when an appropriate number of
answers had been found.
          </p>
          <p>Alternatively, it may be that our diluted system, while certainly
poorer in overall quality (in the sense that non-relevant answers were
introduced into the ranking), was not poor enough to induce different
behaviour. Smith and Kantor used results typically from the 300th
position in Google’s results: even today, these are unreliable for
the simplest of our topics, and in 2008 will almost certainly have
produced a poor result set. Importantly, our diluted system always
included a few high-ranked results.</p>
          <p>Either way, our results raise an important question about how the
effectiveness of search systems should be analysed. While some
fine-grained aspects of user clicking behaviour differed between
the full and diluted treatments, the majority of behaviours did not.</p>
          <p>
            This outcome is in line with previous results that found little
relationship between user behaviour and system quality as measured
by common IR evaluation metrics such as MAP [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ]. The
question then becomes one of whether even a significant improvement
in effectiveness, as measured by some metric, actually results in
improved task performance. In future work, we therefore plan to
systematically investigate different levels of answer-page dilution,
to establish guidelines for the extent of practical differences that
need to be present in search systems for measurable disparities in
user behaviour to manifest. We also plan to explore the issue of the
impact that specific variations in task instructions have on searcher
behaviour through a controlled user study in a work task-based
framework [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ].
          </p>
          <p>Exploratory Search Missions for TREC Topics
Martin Potthast</p>
          <p>Matthias Hagen</p>
          <p>Michael Völske</p>
          <p>Benno Stein
Bauhaus-Universität Weimar</p>
          <p>99421 Weimar, Germany
&lt;first name&gt;.&lt;last name&gt;@uni-weimar.de
ABSTRACT
We report on the construction of a new query log corpus that consists
of 150 exploratory search missions, each of which corresponds to
one of the topics used at the TREC Web Tracks 2009–2011.
Involved in the construction was a group of 12 professional writers,
hired at the crowdsourcing platform oDesk, who were given the task
to write essays of 5000 words length about these topics, thereby
inducing genuine information needs. The writers used a ClueWeb09
search engine for their research to ensure reproducibility. Thousands
of queries, clicks, and relevance judgments were recorded. This
paper overviews the research that preceded our endeavors, details
the corpus construction, gives quantitative and qualitative analyses
of the data obtained, and provides original insights into the
querying behavior of writers. With our work we contribute a missing
building block in a relevant evaluation setting in order to allow for
better answers to questions such as: “What is the performance of
today’s search engines on exploratory search?” and “How can it be
improved?” The corpus will be made publicly available.</p>
          <p>Categories and Subject Descriptors: H.3.3 [Information Search
and Retrieval]: Query formulation</p>
          <p>INTRODUCTION</p>
          <p>Humans frequently conduct task-based information search, i.e.,
they interact with search appliances in order to conduct the research
deemed necessary to solve knowledge-intensive tasks. Examples
include long-lasting interactions which may involve many search
sessions spread out across several days. Modern web search
engines, however, are optimized for the diametrically opposed task,
namely to answer short-term, atomic information needs.
Nevertheless, research has picked up this challenge: in recent years, a
number of new solutions for exploratory search have been proposed
and evaluated. However, most of them involve an overhauling of
the entire search experience. We argue that exploratory search tasks
are already being tackled, after all, and that this fact has not been
sufficiently investigated. Reasons for this shortcoming can be found
in the lack of publicly available data to be studied. Ideally, for any
given task that fits the aforementioned description, one would have
a large set of search interaction logs from a diversity of humans
solving it. Obtaining such data, even for a single task, has not been
done at scale until now. Even search companies, which have access
to substantial amounts of raw query log data, face difficulties in
discerning individual exploratory tasks from their logs.</p>
          <p>In this paper, we contribute by introducing the first large corpus of
long, exploratory search missions. The corpus was constructed via
Presented at EuroHCIR2013. Copyright c 2013 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors..
crowdsourcing by employing writers whose task was to write long
essays on given TREC topics, using a ClueWeb09 search engine for
research. Hence, our corpus forms a strong connection to existing
evaluation resources that are used frequently in information retrieval.</p>
          <p>Further, it captures the way how average users perform exploratory
search today, using state-of-the-art search interfaces. The new
corpus is intended to serve as a point of reference for modeling users
and tasks as well as for comparison with new retrieval models and
interfaces. Key figures of the corpus are shown in Table 2.</p>
          <p>After a brief review of related work, Section 2 details the corpus
construction and Section 3 gives first quantitative and qualitative
analyses, concluding with insights into writers’ search behavior.
1.1</p>
          <p>Related Work</p>
          <p>
            To date, the most comprehensive overview of research on
exploratory search systems is that of White and Roth [19]. More
recent contributions not covered in this body of work include the
approaches proposed by Morris et al. [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ], Bozzon et al. [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ],
Cartright et al. [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ], and Bron et al. [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ]. Exploratory search is studied also
within contextual IR and interactive IR, as well as across disciplines,
including human computer interaction, information visualization,
and knowledge management.
          </p>
          <p>Regarding the evaluation of exploratory search systems, White
and Roth [19] conclude that “traditional measures of IR
performance based on retrieval accuracy may be inappropriate for the
evaluation of these systems” and that “exploratory search
evaluation [...] must include a mixture of naturalistic longitudinal studies”
while “[...] simulations developed based on interaction logs may
serve as a compromise between existing IR evaluation paradigms
and [...] exploratory search evaluation.” The necessity of user
studies makes evaluations cumbersome and, above all, expensive. By
providing part of the solution (a decent corpus) for free, we want
to overcome the outlined difficulties. Our corpus compiles a solid
database of exploratory search behavior, which researchers may use
for comparison purposes as well as for bootstrapping simulations.</p>
          <p>
            Regarding standardized resources to evaluate exploratory search,
hardly any have been published up to now. White et al. [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ]
dedicated a workshop to evaluating exploratory search systems in which
requirements, methodologies, as well as some tools have been
proposed. Yet, later on, White and Roth [19] found out that still no
“methodological rigor” has been reached—a situation which has not
changed much until today. The departure from traditional
evaluation methodologies (such as the Cranfield paradigm) and resources
(especially those employed at TREC) has lead researchers to devise
ad-hoc evaluations which are mostly incomparable across papers
and which cannot be reproduced easily.
          </p>
          <p>
            A potential source of data for the purpose of assessing current
exploratory search behavior is to detect exploratory search tasks
within raw search engine logs, such as the 2006 AOL query log [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ].
          </p>
          <p>
            However, most session detection algorithms deal with short term
tasks only and the few algorithms that aim to detect longer search
missions still have problems when detecting interesting semantic
connections of intertwined search tasks [
            <xref ref-type="bibr" rid="ref10 ref12 ref24 ref26 ref8">10, 12, 8</xref>
            ]. In this regard,
our corpus may be considered the first of its kind.
          </p>
          <p>
            To justify our choice of an exploratory task, namely that of writing
an essay about a given TREC topic, we refer to Kules and Capra [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ],
who manually identified exploratory tasks from raw query logs on
a small scale, most of which turned out to involve writing on a
given subject. Egusa et al. [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ] describe a user study in which they
asked participants to do research for a writing task, however, without
actually writing something. This study is perhaps closest to ours,
although the underlying data has not been published. The most
notable distinction is that we asked our writers to actually write,
thereby creating a much more realistic and demanding state of mind
since their essays had to be delivered on time.
          </p>
          <p>CORPUS CONSTRUCTION</p>
          <p>As discussed in the related work, essay writing is considered a
valid approach to study exploratory search. Two data sets form the
basis for constructing a respective corpus, namely (1) a set of topics
to write about and (2) a set of web pages to research about a given
topic. With regard to the former, we resort to topics used at TREC,
specifically to those from the Web Tracks 2009–2011. With regard
to the latter, we employ the ClueWeb09 (and not the “real web in the
wild”). The ClueWeb09 consists of more than one billion documents
from ten languages; it comprises a representative cross-section of the
real web, is a widely accepted resource among researchers, and it is
used to evaluate the retrieval performance of search engines within
several TREC tracks. The connection to TREC will strengthen the
compatibility with existing evaluation methodology and allow for
unforeseen synergies. Based on the above decisions, our corpus
construction steps can be summarized as follows:
1. Rephrasing of the 150 topics used at the TREC Web Tracks</p>
          <p>2009–2011 so that they invite people to write an essay.
2. Indexing of the English portion of the ClueWeb09 (about
0.5 billion documents) using the BM25F retrieval model plus
additional features.
3. Development of a search interface that allows for answering
queries within milliseconds and that is designed along the
lines of commercial search interfaces.
4. Development of a browsing interface for the ClueWeb09,
which serves ClueWeb09 pages on demand and which
rewrites links on delivered pages so that they point to their
corresponding ClueWeb09 pages on our servers.
5. Recruiting 12 professional writers at the crowdsourcing
plat</p>
          <p>form oDesk from a wide range of hourly rates for diversity.
6. Instructing the writers to write essays of at least 5000 words
length (corresponds to an average student’s homework
assignment) about an open topic among the initial 150, using our
search engine and browsing only ClueWeb09 pages.
7. Logging all writers’ interactions with the search engine and</p>
          <p>the ClueWeb09 on a per-topic basis at our site.
8. Double-checking all of the 150 essays for quality.</p>
          <p>After the deployment of the search engine and successfully
completed usability tests (see Steps 2-4 and 7 above), the actual corpus
construction took nine months, from April 2012 through
December 2012. The post-processing of the data took another four months,
so that this corpus is among the first, late-breaking results from
our efforts. However, the outlined experimental setup can
obviously serve different lines of research. The remainder of the section
presents elements of our setup in greater detail.</p>
          <p>Used TREC Topics.</p>
          <p>Since the topics from the TREC Web Tracks 2009–2011 were
not amenable for our purpose as is, we rephrased them so that they
ask for writing an essay instead of searching for facts. Consider for
example topic 001 from the TREC Web Track 2009:</p>
          <p>Query. obama family tree
Description. Find information on President Barack
Obama’s family history, including genealogy, national
origins, places and dates of birth, etc.</p>
          <p>Sub-topic 1. Find the TIME magazine photo essay
“Barack Obama’s Family Tree.”
Sub-topic 2. Where did Barack Obama’s parents and
grandparents come from?
Sub-topic 3. Find biographical information on Barack</p>
          <p>Obama’s mother.</p>
          <p>This topic is rephrased as follows:</p>
          <p>Obama’s family. Write about President Barack
Obama’s family history, including genealogy, national
origins, places and dates of birth, etc. Where did Barack
Obama’s parents and grandparents come from? Also
include a brief biography of Obama’s mother.</p>
          <p>In the example, Sub-topic 1 is considered too specific for our
purposes while the other sub-topics are retained. TREC Web track
topics divide into faceted and ambiguous topics. While topics of the
first kind can be directly rephrased into essay topics, from topics of
the second kind one of the available interpretations is chosen.</p>
          <p>A Search Engine for Controlled Experiments.</p>
          <p>
            To give the oDesk writers a familiar search experience while
maintaining reproducibility at the same time, we developed a tailored
search engine called ChatNoir [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. Besides ours, the only other
public search engine for the ClueWeb09 is hosted at Carnegie
Mellon and based on Indri. Unfortunately, it is far from our efficiency
requirements. Our search engine returns results after a couple of
hundreds of milliseconds, its interface follows industry standards,
and it features an API that allows for user tracking.
          </p>
          <p>
            ChatNoir is based on the BM25F retrieval model [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ], uses the
anchor text list provided by Hiemstra and Hauff [
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ], the PageRanks
provided by the Carnegie Mellon University,1 and the spam rank list
provided by Cormack et al. [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ]. ChatNoir comes with a proximity
feature with variable-width buckets as described by Elsayed et al. [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ].
          </p>
          <p>Our choice of retrieval model and ranking features is intended to
provide a reasonable baseline performance. However, it is neither
near as mature as those of commercial search engines nor does it
compete with the best-performing models proposed at TREC. Yet,
it is among the most widely accepted models in the information
retrieval community, which underlines our goal of reproducibility.</p>
          <p>In addition to its retrieval model, ChatNoir implements two search
facets: text readability scoring and long text search. The former
facet, similar to that provided by Google, scores the readability of a
text found on a web page via the well-known Flesh-Kincaid grade
level formula: it estimates the number of years of education required
in order to understand a given text. This number is mapped onto
the three categories “simple”, “intermediate”, and “expert.” The
long text search facet omits search results which do not contain at
least one continuous paragraph of text that exceeds 300 words. The
two facets can be combined with each other. They are meant to
support writers that want to reuse text from retrieved search results.</p>
          <p>Especially interesting for this type of writers are result documents
containing longer text passages and documents of a specific reading
1http://boston.lti.cs.cmu.edu/clueWeb09/wiki/tiki-index.php?page=PageRank
level such that reusing text from the results still yields an essay with
homogeneous readability.</p>
          <p>When clicking on a search result, ChatNoir does not link into
the real web but redirects into the ClueWeb09. Though ClueWeb09
provides the original URLs from which the web pages have been
obtained, many of these page may have gone or been updated since. We
hence set up an interface that serves web pages from the ClueWeb09
on demand: when accessing a web page, it is pre-processed before
being shipped, removing all kinds of automatic referrers and
replacing all links to the real web with links to their counterpart inside
ClueWeb09. This way, the ClueWeb09 can be browsed as if surfing
the real web and it becomes possible to track a user’s movements.</p>
          <p>The ClueWeb09 is stored in the HDFS of our 40 node Hadoop
cluster, and web pages are fetched with latencies of about 200ms.</p>
          <p>ChatNoir’s inverted index has been optimized to guarantee fast
response times, and it is deployed on the same cluster.</p>
          <p>Hired Writers.</p>
          <p>
            Our ideal writer has experience in writing, is capable of writing
about a diversity of topics, can complete a text in a timely
manner, possesses decent English writing skills, and is well-versed in
using the aforementioned technologies. This wish list lead us to
favor (semi-)professional writers over, for instance, volunteer
students recruited at our university. To hire writers, we made use of
the crowdsourcing platform oDesk.2 Crowdsourcing has quickly
become one of the cornerstones for constructing evaluation
corpora, which is especially true for paid crowdsourcing. Compared
to Amazon’s Mechanical Turk [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ], which is used more frequently
than oDesk, there are virtually no workers at oDesk submitting fake
results due to advanced rating features for workers and employers.
          </p>
          <p>Table 1 gives an overview of the demographics of the writers we
hired, based on a questionnaire and their resumes at oDesk. Most
of them come from an English-speaking country, and almost all of
them speak more than one language, which suggests a reasonably
good education. Two thirds of the writers are female, and all of them
have years of writing experience. Hourly wages were negotiated
individually and range from 3 to 34 US-dollars (dependent on skill
and country of residence), with an average of about 12 US-dollars.</p>
          <p>In total, we spent 20 468 US-dollars to pay the writers.</p>
          <p>CORPUS ANALYSIS</p>
          <p>This section presents the results of a preliminary corpus analysis
that gives an overview of the data and sheds some light onto the
search behavior of writers doing research.
2http://www.odesk.com
Corpus
Characteristic
Writers
Topics
Topics / Writer
Queries
Queries / Topic
Clicks
Clicks / Topic
Clicks / Query
Sessions
Sessions / Topic
Days
Days / Topic
Hours
Hours / Writer
Hours / Topic
Irrelevant
Irrelevant / Topic
Irrelevant / Query
Relevant
Relevant / Topic
Relevant / Query
Key
Key / Topic
Key / Query
min
Corpus Statistics.</p>
          <p>Table 2 shows key figures of the query logs collected, including
the absolute numbers of queries, relevance judgments, working days,
and working hours, as well as relations among them. On average,
each writer wrote 12.5 essays, while two wrote only one, and one
very prolific writer managed more than 30 essays.</p>
          <p>From the 13 651 submitted queries, each topic got an average
of 91. Note that queries often were submitted twice requesting
more than ten results or using different facets. Typically, about
1.7 results are clicked for consecutive instances of the same query.</p>
          <p>For comparison, the average number of clicks per query in the
aforementioned AOL query log is 2.0. In this regard, the behavior of
our writers on individual queries does not seem to differ much from
that of the average AOL user in 2006. Most of the clicks we recorded
are search result clicks, whereas 2457 of them are browsing clicks
on web page links. Among the browsing clicks, 11.3% are clicks
on links that point to the same web page (i.e., anchor links using a
URL’s hash part). The longest click trail observed lasted 51 unique
web pages but most click trails are very short. This is surprising,
since we expected a larger proportion of browsing clicks, but it
also shows our writers relied heavily on the search engine. If this
behavior generalizes, the need for a more advanced support of
exploratory search tasks from search engines becomes obvious.</p>
          <p>
            The queries of each writer can be divided into a total of 931
sessions with an average 12.3 sessions per topic. Here, a session is
defined as a sequence of queries recorded on a given topic which
is not divided by a break longer than 30 minutes. Despite other
claims in the literature (e.g., in [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ]), we argue that, in our case,
sessions can be reliably identified by means of a timeout because of
our a priori knowledge about which query belongs to which topic
(i.e., task). Typically, finishing an essay took 4.9 days, which fits
well the definition of exploratory search tasks being long-lasting.
          </p>
          <p>In their essays, writers referred to web pages they found during
their search, citing specific passages and topic-related information
used in their texts. This forms an interesting relevance signal which
allows us to separate irrelevant from relevant web pages. Slightly
different to the terminology of TREC, we consider web pages referred
to in an essay as key documents for its respective topic, whereas
web pages that are on a click trail leading to a key document are
relevant. The fact, that there are only few click trails of this kind
explains the unusually high number of key documents compared
to that of relevant ones. The remainder of web pages which were
accessed but discarded by our writers may be considered irrelevant.</p>
          <p>
            The writer’s search interactions are made freely available as the
Webis-Query-Log-12.3 Note that the writing interactions are the
focus of our accompanying ACL paper [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] and contained in the
Webis text reuse corpus 2012 (Webis-TRC-12).
          </p>
          <p>Exploring Exploratory Search Missions.</p>
          <p>To get an inkling of the wealth of data in our corpus, and how it
may influence the design of exploratory search systems, we analyze
the writers’ search behavior during essay writing. Figure 1 shows
for each of the 150 topics a curve of the percentage of queries at any
given time between a writer’s first query and an essay’s completion.</p>
          <p>We have normalized the time axis and excluded working breaks of
more than five minutes. The curves are organized so as to highlight
the spectrum of different search behaviors we have observed: in
row A, 70–90% of the queries are submitted toward the end of the
writing task, whereas in row F almost all queries are submitted at the
beginning. In between, however, sets of queries are often submitted
in short “bursts,” followed by extended periods of writing, which
can be inferred from the plateaus in the curves (e.g., cell C12). Only
in some cases (e.g., cell C10) a linear increase of queries over time
can be observed for a non-trivial amount of queries, which indicates
continuous switching between searching and writing.</p>
          <p>From these observations, it can be inferred that query frequency
alone is not a good indicator of task completion or the current stage
of a task, but different algorithms are required for different mission
types. Moreover, exploratory search systems have to deal with a
broad subset of the spectrum and be able to make the most of few
queries, or be prepared that writers interact only a few times with
them. Our ongoing research on this aspect focuses on predicting the
type of search mission, since we found it does not simply depend
on the writer or a topic’s difficulty as perceived by the writer.
4. SUMMARY</p>
          <p>We introduce the first corpus of search missions for the
exploratory task of writing. The corpus is of representative scale,
comprising 150 different writing tasks and thousands of queries,
clicks, and relevance judgments. A preliminary corpus analysis
shows the wide variety of different search behavior to expect from a
writer conducting research online. We expect further insights from
a forthcoming in-depth analysis, whereas the results mentioned
demonstrate the utility of our publicly available corpus.
3http://www.webis.de/research/corpora
5.</p>
          <p>Interactive Exploration of Geographic Regions with
Web-based Keyword Distributions</p>
          <p>Chandan Kumar
University of Oldenburg,</p>
          <p>Oldenburg, Germany
chandan.kumar@unioldenburg.de</p>
          <p>Wilko Heuten</p>
          <p>OFFIS – Institute for
Information Technology,</p>
          <p>Oldenburg, Germany
wilko.heuten@offis.de</p>
          <p>Dirk Ahlers
NTNU – Norwegian University
of Science and Technology,</p>
          <p>Trondheim, Norway
dirk.ahlers@idi.ntnu.no</p>
          <p>Susanne Boll
University of Oldenburg,</p>
          <p>Oldenburg, Germany
susanne.boll@unioldenburg.de
ABSTRACT
The most common and visible use of geographic information
retrieval (GIR) today is the search for specific points of
interest that serve an information need for places to visit.
However, in some planning and decision making processes, the
interest lies not in specific places, but rather in the makeup
of a certain region. This may be for tourist purposes, to find
a new place to live during relocation planning, or to learn
more about a city in general. Geospatial Web pages contain
rich spatial information content about the geo-located
facilities that could characterize the atmosphere, composition,
and spatial distribution of geographic regions. But the
current means of Web-based GIR interfaces only support the
sequential search of geo-located facilities and services
individually, and limit the end users on abstracted view,
analysis and comparison of urban areas. In this work we propose
a system that abstracts from the places and instead
generates the makeup of a region based on extracted keywords we
find on the Web pages of the region. We can then use this
textual fingerprint to identify and compare other suitable
regions which exhibit a similar fingerprint. The developed
interface allows the user to get a grid overview, but also to
drill in and compare selected regions as well as adapt the
list of ranked keywords.</p>
          <p>
            Categories and Subject Descriptors
H.3.3 [Information Storage and Retrieval]: Information
Search and Retrieval; H.5.2 [Information Interfaces and
Presentation]: User Interfaces
1. INTRODUCTION
Geospatial search has become a widely accepted search mode
o↵ ered by many commercial search engines. Their
interfaces can easily be used to answer relatively simple requests
such as “restaurant in Berlin” on a point-based map
interface, which additionally gives extended information about
entities [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. A corresponding strong research interested has
developed in the field of geographic information retrieval,
e.g., [
            <xref ref-type="bibr" rid="ref15 ref17 ref2">2, 17, 15</xref>
            ]. However, there are many tasks in which the
retrieval of individual pinpointed entities such as facilities,
services, businesses, or infrastructure cannot satisfy user’s
more complex spatial information needs.
          </p>
          <p>To support more complex tasks we propose a new retrieval
method based on entities. For example, sometimes the
distribution of results on a map can already inform certain
views about areas, e.g., a search for “bar” may show a
clustering of results that can be used for “eyeballing” a region
of nightlife even without sophisticated geospatial analysis.</p>
          <p>
            However, as users become more used to local search, more
complex search types and supporting analysis are desired
that enable a combined view onto the underlying data [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ].
          </p>
          <p>
            Exploration of geographic regions and their characterization
was found as one of the key desire of local search users in
our requirement study [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ]. A person who is moving to a
new area or city would like to find similar neighborhoods
or regions with a similar makeup to their current home. It
might not even be the concrete entities, but rather the
atmosphere, composition, and spatial distribution that make up
the “feeling” of a neighborhood that best capture the
intention of a user. To assess this similarity of regions we propose
a spatial fingerprint (query-by-spatial-example) that acts as
an abstracted view onto the same point-based data.
          </p>
          <p>We also aim to provide new visual tools for the exploration of
geographic regions. While the necessary multi-dimensional
geospatial data is already available, there is no suitable
interface to query them, let alone to deal with the multi-criteria
complexity. In this paper we describe a visual-interactive
GIR system to support the retrieval of relevant geospatial
regions and enable users to explore and interact with
geospatial data. We propose a new query-by-spatial-example
interaction method in which a user-selected region’s
characteristic is fingerprinted to present similar regions. Users can
interactively refine their query to use those characteristics
of a region that are most important to them. For a more
detailed overview, we use the full text of georeferenced Web
pages for queries and analysis. This work goes beyond
conventional GIR interfaces as it allows users to interact with
aggregated spatial information via spatial queries instead
of only textual querying, which is especially important to
define regions of interest. We discuss the necessary input,
visualization, comparison, refinement, and ranking methods
in the remainder of this paper.
2. USING THE GEOSPATIAL WEB TO
CHAR</p>
          <p>
            ACTERIZE GEOGRAPHIC REGIONS
The distribution of geo-entities is used to illustrate the
characteristics and dynamics of a geographic region. A
geoentity is a real life entity at a physical location, e.g., a
restaurant, theatre, pub, museum, business, school, etc. To
open these entities up for aggregate and multi-criteria region
characterization, they need a certain depth of information
associated with them. It is obvious that only position
information or the name of a place is insu cient, so categorial
or textual description is needed. For initial studies [
            <xref ref-type="bibr" rid="ref11 ref25 ref27 ref9">11, 9</xref>
            ]
we used OpenStreetMap (OSM)1 which uses a tagging
system for categories. To better characterize the geo entities
we now use their associated Web pages. The reason for this
is the massive increase of the amount of usable data. The
Web pages of entities contain a lot more than just the
basic information and can therefore be used to uncover much
more detailed information. This method can also include
additional sources such as events happening in the region
or user-generated content on third-party pages [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]. We later
describe how we identify the most meaningful keywords from
the pages for this task.
          </p>
          <p>
            To actually make the connection from a location to Web
pages, we assume that the presence of location references
on a page is a strong indication that the page is associated
with the entity at that location. We use our geoparser to
extract location references and thereby assess the geographical
scopes of a page. The geoparser is trained to the presence of
location references in the form of addresses within the page
content. This is a suitable approach for the urban areas
we are addressing in this work, because we need a
geospatial granularity at the sub-neighborhood level.
Knowledgebased identification and verification of the addresses is done
against a gazetteer extended with street names, which we
1http://www.openstreetmap.org/
fetched from OSM for the major cities of Germany. To
retrieve actual pages, we crawled the Web with a geospatially
focused crawler [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ] based on the geoparser and built a rich
geo-index for various cities of Germany, where each city
contains several thousand geotagged Web pages with their full
textual content.
3. INTERFACE FOR EXPLORATION OF
          </p>
          <p>GEOGRAPHIC REGIONS OF INTEREST
We have implemented two main interaction modes in the
Web interface as shown in Figure 1. A user intends to
compare multiple geographic regions of Frankfurt (target region,
right in the dual-map view) with respect to a certain relevant
region in Berlin (query region, left). The current reference
region of interest is specified via a visual query. The user
can then either select regions by placing markers onto the
map, or alternatively use a grid overview (right side of
Figure 1). In both cases, the system computes the relevance of
the target regions with respect to the characteristics of the
query region.
3.1 Query-by-spatial-example
Most GIR interfaces use a conventional textual query as
input method to describe user’s information need or use the
currently selected map viewport. We wanted to give users
the ability to arbitrarily define their own spatial region of
interest. The free definition of the query region is important,
as users may not always want a neighborhood that is
easily describable by a textual query. We therefore enabled to
query by spatial example, where users can define the query
region by drawing on map. Figure 1 shows an example of
a user selected region of interest via a polygon query (by
mouse clicks and drag) in the city of Berlin.
3.2 Visualization of suitable geographic regions
Users can select several location preferences in their target
region that they would like to explore by positioning markers
on the map interface. The system defines the targets with
a circle around the user-selected locations with the same
diameter as the reference region polygon. The target regions
obtain the ranking with respect to their similarity with the
reference region. Their relevance is shown by the
percentage similarity and the heatmap based relevance
visualization. We used a color scheme of di↵ erent green tones which
di↵ ered in their transparency. Light colors represented low
relevance, dark colors were used to indicate high relevance.</p>
          <p>The color scheme selection was aided by ColorBrewer 2.</p>
          <p>As an example, Figure 1 shows 4 user-selected locations on
the city map of Frankfurt, the circle regions around these
4 markers have the same diameter as the query region in
Berlin. The target region in the centre of the city is most
relevant with the similarity of 88%, and consequently has
the darkest green tone. If a user has not yet formed any
preference, we o↵ er an aggregate overview of geo-entities.</p>
          <p>
            We partition the map area using a grid raster [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ], as we
do not intend to restrict user exploration to only selected
areas. There could be situations when users look beyond the
specific target regions, and would like to have an overview
of the whole city with respect to a query region. The right
side of Figure 1 shows the aggregated ranked view of the
grid-based visualization. Each grid cell represents the overall
relevance with respect to the query region. The visualization
gives a good overview and assessment on relevant regions
which the user can then explore further. Users can select
the grid size, which is otherwise similar to the size of the
query region. The grid layout is fixed to the city boundaries
as we intend to give the overview of whole city. In the future
we would like to make it more dynamic where users should
be able to shift the grid layout, since a slight variation in
grid cell boundaries could alter the relevance results.
3.3 Exploration and interaction with geographic
          </p>
          <p>regions via keyword distributions
Interaction models should provide end users the
opportunity to explore the characteristics of selected regions, and
adapt it further to their requirements. We initially show
the most relevant keywords of the respective region using a
word cloud. The word cloud provides more detailed
information on keyword distribution when the mouse hovers over
it. The font size and order of the keywords signify their
relevance. Figure 2 shows the comparison of the query region
with the most relevant target region via both their keyword
distributions. In this case, the distributions of both regions
are very similar, leading to the high relevance score for the
target region.</p>
          <p>Since the keyword characteristics of a query region is
derived from the georeferenced Web pages, there are situations
where a user might not be satisfied with the spatial
descrip2http://colorbrewer2.org
tion and wants to influence the keywords. In the example
of Figure 3, a user decides that pubs are more important
than restaurant, fast food is not an aspect of his lifestyle
and should be replaced by education facilities near his new
home. In such scenarios users need to interact and adapt
the generated keyword distributions of query regions. We
make the word cloud interactive and editable. Users can
drag keywords to alter their position and thus their
significance. They can also edit, delete or replace keywords in the
word cloud to change the criteria. After modifying the
keyword distribution, users can revisualize the target regions to
update their ranking. Figure 3 shows this user interaction
with the word cloud, including the revisualization of the
updated ranking of target regions, which are visibly di↵ erent
from the previous ranking of Figure 2.
4. TEXT-BASED CHARACTERIZATION AND</p>
          <p>
            RANKING OF GEOGRAPHIC REGIONS
We adapt common IR methods for ranking and similarity
measures. In relevance-based language models, the
similarity of a document to a query is the probability that a given
document would generate the query [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. To be able to do
the same with geographic regions, we add a transitional step.
          </p>
          <p>
            Regions are considered as compound documents built from
the Web pages of the entities inside them. We can then
define the similarity of document clusters of regions based
on the probability that the target region can generate the
query region. The Kullback-Leibler divergence is used for
comparison [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ].
          </p>
          <p>For a geospatial document d, we estimate P (w|d) , which is a
unigram language model , with the maximum likelihood
estimator, simply given by relative counts: P (w|d) = tf(w,d) ,</p>
          <p>|d|
here tf (w, d) is the frequency of word w in the document d
and |d| is the length of the document d. A geographic region
contains several geospatial documents insides its footprint
area. We define a geographic region based on a document
cluster D which contains document {d1, d2....dk}, and the
distribution of a particular word w in the geographic
region would be estimated with its combine probability in the
collection P (w|D) = k1 Pk</p>
          <p>i=1 P (w|di). The word cloud
represents the most prominent keywords of the region with
respect to their ranked probability distribution P (w|D). The
comparison of regions is done with respect to their
probability distribution using KL-divergence. A target region x will
be compared to the query region as following</p>
          <p>Relevance(Regionx) =</p>
          <p>X P (w|Dq)log
w</p>
          <p>P (w|Dq)</p>
          <p>
            P (w|Dx)
The computation of this formula involves a sum over all
the words that have a non-zero probability according to
P (w|Dq). Each region Regionx gets a relevance score
according to its distribution comparison to the query region
Regionq. All target regions (user selected regions or grid
based divisions) are ranked with respect to their relevance
score for visualization.
5. RELATED WORK
The field of geographic information retrieval examines
documents’ geospatial features at a regional scale and also at
smaller granularities and usually supports keyword@location
queries [
            <xref ref-type="bibr" rid="ref15 ref17 ref2">2, 17, 15</xref>
            ]. Similarly, location-based services (e.g.,
FourSquare, Yelp, Google Maps) allow users to retrieve and
visualize geo-entities matching a category or search term.
          </p>
          <p>
            However, search for multiple categories or other complex
tasks is usually not supported. Some non-conventional
spatial querying methods have been proposed, e.g.,
query-bysketch on a map [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ]. Other work uses the density of
arbitrary user-supplied keywords to build a query region [
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ]. Tag
clouds have been adapted to maps, exploiting georeferenced
tags [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]. Locally characteristic keywords can be extracted
for map visualization and to show their spatial extent [19].
          </p>
          <p>
            None of these approaches make a larger word cloud available,
but only the main terms. Other geovisualization approaches
[
            <xref ref-type="bibr" rid="ref21 ref23 ref5 ref7">5, 7</xref>
            ] approach multi-criteria analysis, but are usually
targeted to specific domains and experts. The Inspect system
was tailored at geospatial analysts to visually filter and
explore multidimensional data [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. A multi-criteria
evaluation for home buyers was proposed in [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ]. The scenario of
spatial decision making is similar to ours, but it focused on
experts and spatial computation issues rather than interface
and visualization aspects.
          </p>
          <p>Our system interface di↵ ers in the granularity of
information need and representation, i.e., we focus on the ranking
of regions, but base it on high-granularity geo-entities that
have a very exact location, which ensures that the spatial
query does not produce overlap to neighboring regions and
makes the multi-criteria analysis more exact to be executed
at arbitrary region sizes.
6. CONCLUSIONS AND FUTURE WORK
Most current local search interfaces do not o↵ er adequate
support for the exploration and comparison of geographic
areas and regions. End users need visual and interactive
assistance from GIR systems for an abstracted overview and
analysis of geospatial data. We proposed interactive
interfaces for the characterization and assessment of relevant
geographic regions that enable end-users to query, analyze and
interact with the rich geospatial data available on the Web
in user-selected geographic regions. The relevance of regions
is based on the similarity of keyword distributions.</p>
          <p>The observation of results shows satisfactory performance by
uncovering realistic and meaningful keywords defining the
regions. We observed that the characterization and
comparison of geographic regions show good results with respect
to geo-located facilities and infrastructure of German cities,
e.g., clearly distinct characteristics for university, industrial,
or party districts. In the future we plan a more formal
qualitative and quantitative evaluation of these interfaces, to
examine the acceptance of these visualizations with regard
to user-centered aspects such as exploration ability,
information overload, and cognitive demand. We would also like
to explore more advanced interaction methods to enhance
the usability of the proposed visualizations.</p>
          <p>Additionally, we envision more powerful region similarity
measures such as landscape and topological similarity,
similarity via social media, and an integration of additional data
sources.</p>
          <p>Acknowledgments
The authors are grateful to the DFG SPP 1335 ‘Scalable
Visual Analytics’ priority program which funds the project
UrbanExplorer. The 2nd author acknowledges funding from
the ERCIM “Alain Bensoussan” Fellowship Programme.</p>
          <p>Inferring Music Selections for Casual Music Interaction</p>
          <p>Daniel Boland
University of Glasgow</p>
          <p>United Kingdom
daniel@dcs.gla.ac.uk</p>
          <p>Ross McLachlan
University of Glasgow</p>
          <p>United Kingdom
r.mclachlan.1@
research.gla.ac.uk</p>
          <p>Roderick Murray-Smith</p>
          <p>University of Glasgow</p>
          <p>United Kingdom
rod@dcs.gla.ac.uk
ABSTRACT
We present two novel music interaction systems developed
for casual exploratory search. In casual search scenarios,
users have an ill-defined information need and it is not clear
how to determine relevance. We apply Bayesian inference
using evidence of listening intent in these cases, allowing
for a belief over a music collection to be inferred. The first
system using this approach allows users to retrieve music
by subjectively tapping a song’s rhythm. The second
system enables users to browse their music collection using a
radio-like interaction that spans from casual mood-setting
through to explicit music selection. These systems embrace
the uncertainty of the information need to infer the user’s
intended music selection in casual music interactions.</p>
          <p>
            Categories and Subject Descriptors
H.5.2 [Information interfaces]: User Interfaces
General Terms
Design, Human Factors, Theory
1. INTRODUCTION
When interacting with a music system, listeners are faced
with selecting songs from increasingly large music
collections. With services like Spotify, these libraries can include
many songs the user has never heard of. This retrieval is
often a hedonic activity and may not serve a particular
information need. Users do not always have a song in mind
and are often just interested in setting a mood or finding
something ‘good enough’ [
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ]. This type of casual search has
recently been identified as not being well supported within
IR literature [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. In particular, the concept of relevance
becomes nebulous where the information need is not well
defined. By inferring a belief over a music collection using
the likelihood of a user’s input, we implement interactions
which incorporate this uncertainty. These interactions can
account for subjectivity and span from casual, serendipitous
listening through to highly engaged music selection.
          </p>
          <p>Presented at EuroHCIR2013. Copyright c 2013 for the individual papers
by the papers authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.</p>
          <p>
            Music listeners are not always fully engaged with the
selection of music - as evidenced by the success of the shu✏ e
playback feature. Large libraries of music such as Spotify are
available but users often just want background music, not a
specific song out of millions. In these casual search
scenarios, users often satisfice i.e. search for something which is
‘good enough’ [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ]. As this information need is poorly
defined, so too is relevance, placing these interactions outside
of typical Information Retrieval approaches.
          </p>
          <p>UNCERTAIN MUSIC SELECTION
By asking ‘What would this user do?’, we can develop a
likelihood model of user input within an interaction. With
Bayes theorem, this allows for an uncertain belief over a
music space to be inferred. Users can provide evidence of
their listening intent as part of a casual music interaction,
not needing to be fully engaged in the music retrieval. This is
an explicitly user-centered approach, focusing on how a user
will interact with the system. Both the systems discussed
here have been iteratively developed by comparing real user
behaviour against that predicted by the user input models.</p>
          <p>
            We present two novel music retrieval systems which explore
two challenges with this approach: i) how to correctly
interpret evidence which may be subjective and ii) how to allow
users to set their current level of engagement:
i) ‘Query by Tapping’ is a music retrieval technique where
users tap the rhythm of a song in order to retrieve it [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ].
          </p>
          <p>As part of a user-centred development process, we identified
that rhythmic queries are often subjective and so developed
a model of rhythmic input which captures some of this
subjective behaviour. This allows for the system to be trained
to the user’s tapping style, giving significant improvements
over previous e↵ orts at rhythmic music retrieval.
ii) FineTuner is a prototype of a radio-like music interface
that enables users to retrieve music at a level of
engagement suited to their current information need. Users
navigate their music collection using a dial, with the system
using prior knowledge of the user to inform the music
selection. A pressure sensor enables users to assert varying
levels of control over the system – with no pressure, users
can casually tune in to sections of their music collection to
hear recommended music with common characteristics. As
pressure is applied, the user is able to make increasingly
specific selections from the collection. The inferred music
selection is conditioned upon the asserted control, allowing
for the seamless transition from casual mood-setting to
engaged music interaction.
User 1</p>
          <p>User 2</p>
          <p>MODELLING SUBJECTIVITY
In this section we describe our e↵ orts to model the
subjectivity of rhythmic queries, yielding a query by tapping system
for casual music retrieval which can be trained to users to
account for their subjective querying style. After training
the system, a user can tap a rhythm to re-order their music
collection by rhythmic similarity to their query. The top 20
highly ranked results are listed on-screen as a music playlist,
from which the user could also then select a specific song.</p>
          <p>
            Query by tapping provides an example of a casual music
interaction which su↵ ers from subjective queries. In
mobile music-listening contexts, it can often be inconvenient
for users to remove their mobile device from their pocket
or bag and engage with it to select music. This tapping of
music as a querying technique for music is depicted in figure
2. Tapping a rhythm is already a common act and rhythm
is a universal aspect of music [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. In an exploratory design
session where users were asked to provide rhythmic queries,
it became apparent that users di↵ ered in querying style. We
describe this subjective behaviour and our approach to
modelling it in previous work [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. One of the key aspects of the
model is that users have preferences for which instruments
they tap to, as depicted in figure 1.
          </p>
          <p>
            In order to assign a belief to the songs in the music
collection given a rhythmic query, we compare the query to
those predicted by the user input model. This comparison is
done using the edit distance from string comparison
methods, scaling the mismatch penalty to the time di↵ erences
between the rhythmic sequences [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ].
‘Query by Tapping’ has received some consideration in the
Music Information Retrieval community. The term was
introduced in [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] which demonstrated that rhythm alone can
be used to retrieve musical works, with their system
yielding a top 10 ranking for the desired result 51% of the time.
          </p>
          <p>Their work is limited however in considering only
monophonic rhythms i.e. the rhythm from only one instrument,
as opposed to being polyphonic and comprising of multiple
instruments. Their music corpus consists of MIDI
representations of tunes such as ”You are my sunshine” which is
hardly analogous to real world retrieval of popular music.</p>
          <p>
            Rhythmic interaction has been recognised in HCI [
            <xref ref-type="bibr" rid="ref15 ref24 ref8">8, 15</xref>
            ]
with [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ] introducing rhythmic queries as a replacement for
hot-keys. In [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] tempo is used as a rhythmic input for
exploring a music collection – indicating that users enjoyed such a
method of interaction. The consideration of human factors
is also an emerging trend in Music Information Retrieval
[
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. Our work draws upon both these themes, being the
first QBT system to adapt to users. A number of key
techniques for QBT are introduced in [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ] which describes rhythm
as a sequence of time intervals between notes – termed
interonset intervals (IOIs). They identify the need for such
intervals to be defined relative to each other to avoid the user
having to exactly recreate the music’s tempo.
          </p>
          <p>
            In previous implementations of QBT, each IOI is defined
relative to the preceding one [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ]. This sequential
dependency compounds user errors in reproducing a rhythm, as
an erroneous IOI value will also distort the following one.
          </p>
          <p>100
)75
%
(
e
t
a
r
iitno50
n
g
o
c
re25
0
Dial position</p>
          <p>Dial position</p>
          <p>
            Dial position
(a)
(b)
(c)
The approach to rhythmic interaction in [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ] however used
kmeans clustering to classify taps and IOIs into three classes
based on duration. The clustering based approach avoids
the sequential error however loses a great deal of detail in
the rhythmic query and so we explore a hybrid approace.
3.2
          </p>
          <p>
            Evaluation
The most important metric for the system to be usable was
whether a rhythmic input produced an on-screen (top 20)
result. We asked eight participants to provide queries for
songs selected from a corpus of 300 songs which we had
complete note onset data for. Participants listened to the
songs first to ensure familiarity and were asked to provide
training queries for each song. These training queries were
used to train the generative model using leave-one-out
crossvalidation. We use a state-of-the-art onset detection
algorithm (based on measuring spectral flux [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ]) as a baseline
which does not account for subjectivity. Performance
typically improves with query length as seen in figure 3. Higher
rankings are achieved for all query lengths when using the
generative model. Interestingly, queries over 10 seconds lead
to a rapid fall-o↵ in performance - possibly due to errors
accumulating beyond the initial query the user had in mind or
due to users becoming bored.
          </p>
          <p>
            MODELLING ENGAGEMENT
We consider casual search interactions as spanning a range
of levels of engagement. How much a user is willing to
engage with a system and provide evidence of their listening
intent will undoubtedly vary with listening context. An
interaction which is fixedly casual would be as problematic
as one which requires a user’s full attention, with users
unable to take control when they wish to. An example of this
would be old analogue radios – whilst they o↵ er a simple
music interaction, users have limited control over what they
hear. Previous work by Hopmann et al. sought to bring the
benefits of interaction with vintage analog radio to modern
digital music collections [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ], however their work also required
explicit selection (a fixed level of engagement).
          </p>
          <p>We explore how the inference of listening intent can be
conditioned upon the user’s level of engagement, with the music
interaction spanning from casual mood-setting through to
specific song selection. While it would be desirable to bring
the simplicity of radio-like interaction to modern music
collections, mapping a modern music collection to a dial such as
in figure 5 would require prolonged scrolling. An alternative
would be to instead support scrolling through an overview of
the music space however this removes granularity of control
from the user, leaving them unable to select specific items.</p>
          <p>We developed a radio-like system called FineTuner that
allows users to navigate their music, which is arranged along
a mood axis. Users can ‘tune in’ to a mood to hear
recommended songs based on their listening history. FineTuner
allows the user to assert control over the music
recommendation by applying pressure to a sensor. This enables users to
seamlessly transition from a casual style of interaction akin
to a radio to controlling styles such as specifying a particular
sub-area of interest in a music space, or even selecting
individual songs. FineTuner provides a single interaction which
supports casual search through to fully engaged retrieval.
Our system enables both casual and engaged forms of
interaction, giving users varying degrees of control over the
selection of music. In casual interactions where users apply less
pressure, the system can become more autonomous – making
inferences from prior evidence about what the user intended.</p>
          <p>
            This handover of control was termed the ‘H-metaphor’ by
Flemisch et al. where it was likened to riding a horse – as
the rider asserts less control the horse behaves more
autonomously [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ]. By allowing users to make selections from
the general to the specific, the system supports both specific
selections and satisficing. Users can make broad and
uncertain general selections to casually describe what they want
to listen to. However, they can also assert more control over
the system and force it to play a specific song. Control is
asserted by applying force to a pressure sensor.
          </p>
          <p>As the user begins an interaction, they have not applied
pressure and therefore are not asserting control over the system.</p>
          <p>The inferred selection is thus broad, covering an entire
region of their collection and is biased towards popular tracks
(fig. 4a). The music in the inferred selection is visualised by
randomly sampling tracks from it and drawing beams from
the dial position to the album art. The user may press in the
knob to accept the selection and the sampled track is played.</p>
          <p>At low levels of assertion it is likely that most tracks played
would be highly popular tracks. This behaviour is a design
assumption, users may want the system to use other prior
evidence. When the user applies pressure, the system
interprets this as an assertion of control. The inferred selection is
smaller and the spread of beams becomes narrower, the
album art visualisation zooms in to show the smaller selection
(fig. 4b). This selection is a combination of evidence from
the dial position with prior evidence i.e. their last.fm music
history. When users fully assert control (max. pressure),
they navigate the collection album by album (fig. 4c) and
can make exact selections. By varying the pressure, users
seamlessly move through this continuous range of control.</p>
          <p>The smooth change in engagement is achieved using a
simple model of user input. We assume that in an engaged
interaction, users will point precisely at the song of
interest (as in fig. 4c). For more casual selection, we assume
that users will point in the general area (mood) of the music
they want, modelled using a normal distribution as in (fig.
4b). As less pressure is applied the distribution is widened,
leading to less precise selection and a greater role for a prior
belief over the music collection such as listening history.
5. SUMMARY
The scenarios explored here involve casual music retrieval,
where users have an ill-defined information need and browse
for hedonic purposes or to satisfice a music selection. In
these cases, considering what input a user would provide for
target songs and inferring selections is an intuitive approach
which avoids the issue of defining relevance. We show two
music interactions which support the uncertain selection of
music, inferred from casual user input such as tapping a
rhythm or turning a radio dial.</p>
          <p>We have shown that modelling user input for inferring
music selection can address issues of subjectivity by taking a
user-centered approach to model development. The model
can be iterated by comparing its predictions against actual
user behaviour. Accounting for this subjectivity can yield
significant improvements in retrieval performance as well as
creating a more personalised search experience. A key
feature of the second system, FineTuner, is its ability to span
seamlessly from casual search scenarios, such as satisficing,
through to more explicit selections of music. By
conditioning the inference upon the user’s level of engagement, we are
able to interpret the same input space (in this case the dial)
according to the current context.</p>
          <p>Our approach to casual music interaction empowers the user
to enjoy their music while expending as much or as little
e↵ ort in the retrieval as they wish, providing queries in their
own subjective style. Instead of focusing solely on optimising
the retrieval process, we consider it equally important to
design retrieval systems which suit how the user currently
wants to interact. By considering how users might provide
casual evidence for their listening intent, we achieve music
interactions as simple as tapping a beat or tuning a radio.</p>
          <p>ACKNOWLEDGMENTS
We are grateful for support from Bang &amp; Olufsen and the
Danish Council for Strategic Research.</p>
          <p>Search or browse? Casual information access to a cultural
heritage collection
Robert Villa, Paul Clough, Mark Hall, Sophie Rutter</p>
          <p>Information School
University of Sheffield</p>
          <p>Sheffield, UK</p>
          <p>S1 4DP
{r.villa, p.d.clough, m.mhall, sarutter1} @sheffield.ac.uk
ABSTRACT
Public access to cultural heritage collections is a challenging and
ongoing research issue, not least due to the range of different
reasons a user may want to access materials. For example, for a
virtual museum website users may vary from professionals or
experts, to interested members of the public visiting on a whim. In
this paper, we are interested in the latter user: a user who visits a
cultural heritage website without a clear goal or information need
in mind. In the user study reported here, carried out within the
context of the interactive task at CLEF (interactive CHiC), 20
participants explored a subset of Europeana with no explicit task
provided using a custom-built interface that offered both search
and browse functionalities. Results suggest that browsing is used
considerably more by the majority of users when compared to text
search (all participants used the category browser before carrying
out a text search). This highlights the need for cultural heritage
search interfaces to provide browsing functionality in addition to
conventional text search if they wish to support casual search
tasks.</p>
          <p>General Terms
Design, Experimentation, Human Factors.</p>
          <p>
            Keywords
Cultural heritage, virtual museums, information access.
1. INTRODUCTION
Providing public access to cultural heritage is an ongoing and
challenging area of research. Previous work suggests that visitors
to online cultural heritage collections (e.g. virtual museum
visitors) are not necessarily motivated by an explicit task, and that
interacting with cultural heritage collections is exploratory in
nature [
            <xref ref-type="bibr" rid="ref24 ref25 ref8 ref9">8, 9</xref>
            ]. Recent  work in the area of ‘casual search’ [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ] has 
also investigated situations where users are driven by the pleasure
of the search process itself, rather than an explicit information
need.
          </p>
          <p>The focus for this paper is how individuals explore a cultural
heritage collection when given no task. The results may be used
both to contrast with studies which have used explicit tasks, and
to motivate changes to cultural heritage systems to better support
a diverse range of user tasks.</p>
          <p>Presented at EuroHCIR2013. Copyright © 2013 for the individual papers
by the papers’ authors. Copying permitted only for private and academic 
purposes. This volume is published and copyrighted by its editors.</p>
          <p>
            The work reported here is based on initial results from the
Interactive CHiC (Cultural Heritage in CLEF) track of CLEF1 as
run at Sheffield University. The interactive CHiC track is based
on the CHiC Europeana data set as used in 2011 and 2012 [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. An
early prototype of an evaluation framework was used [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] which
allowed the interactive experiment to be semi-automated. In this
work, our focus is on how users explored the collection and in
particular how search and browse were used in this exploration.
          </p>
          <p>We consider three research questions:</p>
          <p>RQ1. How do participants initiate their exploration?
RQ2. Do participants use browse or search in their exploration</p>
          <p>of the collection?
RQ3. How do participants decide to search or browse, when</p>
          <p>
            given no explicit task?
With RQ1 we are particularly interested whether users start their
exploration by browsing categories, or by search. RQ2 then
considers how users access the collection over their whole
session. For RQ3 we will present some initial qualitative data
from our lab-based interactive study, where the aim is to identify
reasons for the use of either the search or browse functions.
2. PREVIOUS WORK
A general review of museum informatics is provided in [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ],
although the more specific area of museum visitor studies,
investigating why and how individuals visit museums, has a long
history [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ]. More recent work has focused on visitors to digital
museums [
            <xref ref-type="bibr" rid="ref21 ref22 ref23 ref5 ref6 ref7">5-7</xref>
            ]. In [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ] the information seeking behavior of
cultural heritage experts was studied through interviews, finding
that complex information gathering was required for the majority
of search tasks. In contrast [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] studied virtual museum visitors,
inspired by the work of [
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ] and [
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ] which suggest that museum
visitors are exploratory in their information seeking. This work
[
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] found that search occurred far more often than browse
behavior for three of the four tasks used in the study, the
exception being an open and broad task where browsing occurred
to a greater degree.
          </p>
          <p>
            Museum visitors can, in some respects, be considered as examples
of “casual  leisure” searchers, as outlined in [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ], where examples 
were found of “need -less” browsing (based on a diary study, and
analysis of Tweets, both outside the domain of cultural heritage).
          </p>
          <p>
            Darby and Clough [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ] investigated the information seeking
1 http://www.promise-noe.eu/unlocking-culture
behavior of genealogists, with an emphasis on the behavior of
amateurs and hobbyists, rather than professionals. In [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ] a
review of three digital libraries projects is carried out, from the
point of view of Ingwersen and Järvelin's Information Seeking
and Retrieval framework [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. Similar to [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ], it points out that
information behavior by end users may be the “end in itself”. 
The study reported here uses a conventional lab-based protocol.
          </p>
          <p>
            However, unlike in previous work, such as [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ], the participants
were not given an explicit task: the underlying aim being to model
a situation closer to that investigated in [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ], where there is no
explicit information need.
3. INTERACTIVE CHiC
A screenshot of the CHiC interactive system is shown in Figure 1.
          </p>
          <p>The interface is split into five main areas, clockwise from left to
right: a category browser, search box, item display, bookbag, and
search results. The search box operates in the conventional
manner, allowing free text queries with search results being
displayed as a grid below. When a result is clicked, it is displayed
in the “item display” on the right. This information will typically 
include  a  small  thumbnail,  textual  description,  and  the  item’s 
associated metadata. Metadata is clickable, e.g. if an item is listed
as being owned by the British Library, clicking on the field will
search for British Library objects. At the bottom of the item
display  is  a  “more  like  this”,  which  displays  the  images  of  up  to
eight similar objects, which can be viewed three at a time.</p>
          <p>
            On  the  left  of  the  interface  is  the  “category  browser”,  which 
allows the user to browse the Europeana collection through a
hierarchy of categories. This hierarchy is automatically generated,
and is based on the work of [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. The technique combines the
Wikipedia category hierarchy with topics derived from Wikipedia
articles into which items are mapped. When a category is clicked,
the main results are updated to list the category contents. Small
right arrows beside each non-leaf category allows the viewing of
sub-categories. The user can therefore search and browse the
collection in three main ways: using a text query, selecting a
category, or selecting item metadata or “more like this”. 
On the bottom right of the interface is the bookbag, into which
items can be placed. Book-bagged items are kept listed on the
display, and can be removed and redisplayed as required. The
underlying search system is based on Apache Solr2,
which provides the text search, spelling checker, and
the  “more  like  this”  suggestions (determined using
Solr’s standard more -like-this functionality. The data
set used was the same as that used in interactive
CHiC, a dump of the Europeana data set3.
4. EXPERIMENTAL SETUP
The search and browse interface was embedded into
an IR evaluation system, which automatically
administered pre- and post-questionnaires, and
displayed the experimental system. All data reported
here is from an in-lab study. This allowed a follow
up interview to be carried out, during which each
participant reviewed his or her search session. To
enable this reviewing, Morae screen recording
software  was  used  to  record  the  user’s  activity,  and 
during the interview, an audio recording was made of
the user’s comments. 
An important aspect of the interactive CHiC experimental design
was that no explicit task was provided to users. Instead
instructions asked the user to explore freely as they wished, until
they were bored. Users were informed after they had been active
for 10 minutes, and could then continue for a further 5 minutes if
they wished, at which point they would be asked to stop (these
timings were carry out by hand, and were approximate). Once this
was finished, the user’s search session would be replayed to th em,
and an interview conducted to investigate the user’s  search 
process. Participants were paid 10 pounds for taking part.
          </p>
          <p>In total 20 participants were recruited for the study, 11 male and 9
female. Eight participants were in the 18-25 year age band, nine in
the 26-35 band; the other 3 between 36-45. The majority were
students (13), with 5 employed, one unemployed, and one
“other”. 13 had completed a higher education degree, while six
were currently studying an undergraduate degree.
5. RESULTS
5.1 Initiation of exploration
RQ1 asks how users initiate their exploration of the collection. To
investigate this, we first looked at how users started their session,
and in particular, their searching. For example, did they select a
category or enter a query?
Over the whole data set four different actions were used by
participants to initiate their session (Table 1, column 2). For the
majority of users, the first action was to select one of the
categories (15 out of the 20 users). It should be noted that the
interface, on startup, showed a set of default results to all users.</p>
          <p>For three users, the first action was to display one of these default
results, another user clicked the “next page” to view the next page 
of default results, while the final  user’s  first  action  was  to 
bookmark one of the default result items.</p>
          <p>We also investigated the logs to find out each user’s first search or 
browse action, which could be one of category select, text query,
or metadata/more like this select. As shown in Table 1 (column
3), for all users this was a category select. In addition to counting
the first actions, we also investigated how long each user spent
before either clicking the interface, or starting a new
2 http://lucene.apache.org/solr/
3 http://www.europeana.eu/
it generates a query message, which is usually connected to
a Standard Results List.</p>
          <p>Standard Results List</p>
          <p>
            The Standard Results List component ([
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ], p. 50,
“Examine Results Interface” [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ], p. 77) provides a default 10 item
listing of search results. The Standard Results List includes
support for displaying snippets ([
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ], p. 51) and what Wilson
calls “Usable Information” ([
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ], p. 51) for each result
document. Unlike the other standard components, which can
be used out-of-the-box, the Standard Results List has to be
extended by the researcher in order to be able to access the
search-engine used to power the UI.
3.3
          </p>
          <p>Pagination</p>
          <p>
            The Pagination component ([
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ] p. 70) displays a
configurable number of pages around the current search-results
page. In response to user interaction it sends a start
message with the rank of the first document to paginate to.
3.4
          </p>
          <p>Category Browsing</p>
          <p>
            The Category Browsing component ([
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ], p. 54) provides a
hierarchical category structure that the participant can use
to explore a collection. Clicking on a category sends a query
message with the category’s identifier.
3.5
          </p>
          <p>Saved Documents</p>
          <p>The Saved Documents component provides an area where
the participant can save things that they have found
interesting, to support them in their current task. Documents
are added through a save_document message. The Saved
Documents component supports an optional tagging feature
enabling the participant to tag the document with values
specified by the researcher. This can be used to let the
participant specify why they have chosen that document or how
much it helps them in their current task.
3.6</p>
          <p>Task</p>
          <p>The Task component provides a static display of the task
information to show to the user. Two versions of this
component are provided, one that displays a static text set in
the configuration, and one that can fetch a task description
from the database, based on a parameter passed to it.</p>
          <p>The evaluation work-bench has so far been used to build
two IIR experiments, very di↵ erent in their nature, clearly
demonstrating the work-bench’s flexibility.</p>
          <p>The first experiment (fig. 4) re-uses the standard Task,
Search Box, Pagination, and Saved Documents components,
and extends the Standard Results List to work with the
specific search backend. This set-up re-creates what is
essentially a relatively standard search UI configuration, that is
being used to investigate query session behaviour.</p>
          <p>The second experiment (fig. 5) demonstrates a much
richer interface, with more modifications to the components
and an experiment-specific component. It re-uses the Task
and Category Browsing components, extends the default
Search Box, Pagination, Standard Results List, and Saved
Documents components, and adds a new Item View
component. The message-passing nature of the system made
it possible to quickly integrate the new component, so that
when the participant clicks on a meta-data facet in the Item
View, a query message is sent to the Standard Results List
to find items with the same bit of meta-data. The interface
was used to investigate un-directed exploration behaviour in
a large digital cultural heritage collection.</p>
          <p>WHERE TO GO NEXT?</p>
          <p>The stated aim of this paper was to present a novel,
pluggable, extensible, and configurable IIR interface work-bench,
that supports our wider aim of improving IIR experiment
comparability. The work-bench is su ciently flexible to
support the wide range of web-based IIR experiments that are
undertaken, while being su ciently simple and light-weight
to encourage wide-spread use of the workbench.</p>
          <p>To enable this wide-spread use, the system has been
released under an open-source license6. We are also moving
to engage with the wider research community to determine
to what degree the work-bench satisfies their needs for an
evaluation system and what needs to be done to achieve the
wide-spread use needed to improve IIR experiment
comparability.</p>
          <p>ACKNOWLEDGEMENTS</p>
          <p>The research leading to these results was supported by
the Network of Excellence co-funded by the 7th Framework
Program of the European Commission, grant agreement no.
258191.
6https://bitbucket.org/mhall/pyire</p>
          <p>A Proposal for User-Focused Evaluation and Prediction of
Information Seeking Process</p>
          <p>Chirag Shah
School of Communication &amp; Information (SC&amp;I)</p>
          <p>Rutgers University
4 Huntington St, New Brunswick, NJ 08901, USA</p>
          <p>chirags@rutgers.edu
ABSTRACT
One of the ways IR systems help searchers is by predicting or
assuming what could be useful for their information needs based
on analyzing information objects (documents, queries) and
finding other related objects that may be relevant. Such
approaches often ignore the underlying search process of
information seeking, thus forgoing opportunities for making
process-based recommendations. To overcome this limitation,
we are proposing a new approach that analyzes a searcher’s
current processes to forecast his likelihood of achieving a certain
level of success in the future. Specifically, we propose a
machine-learning based method to dynamically evaluate and
predict search performance several time-steps ahead at each
given time point of the search process during an exploratory
search task. Our prediction method uses a collection of features
extracted solely from the search process such as dwell time,
query entropy and relevance judgment in order to evaluate
whether it will lead to low or high performance in the future.</p>
          <p>Experiments that simulate the effects of switching search paths
show a significant number of subpar search processes improving
after the recommended switch. In effect, the work reported here
provides a new framework for evaluating search processes and
predicting search performance. Importantly, this approach is
based on user processes, and independent of any IR system
allowing for wider applicability that ranges from searching to
recommendations.</p>
          <p>Categories and Subject Descriptors
H.3: INFORMATION STORAGE AND RETRIEVAL H.3.3:
Information Search and Retrieval: Search process; H.3:
INFORMATION STORAGE AND RETRIEVAL H.3.4:
Systems and Software: Performance evaluation (efficiency and
effectiveness)
General Terms
Measurement, Performance, Experimentation
Keywords
Exploratory search, Evaluation, Performance prediction
Presented at EuroHCIR2013.</p>
          <p>Copyright © 2013 for the individual papers by the papers’ authors.</p>
          <p>Copying permitted only for private and academic purposes. This volume
is published and copyrighted by its editors.
2
1</p>
          <p>
            INTRODUCTION
IR evaluations are often concerned with explaining factors
relating to user or system performance after the search and
retrieval are conducted [20]. Most recommender systems,
however, operate with an objective to suggest objects that could
be useful to a user based on his/her or others’ past actions
[
            <xref ref-type="bibr" rid="ref2">2</xref>
            ][19]. We commenced our investigation by broadly asking
how we could take valuable lessons from both IR evaluations
and recommender systems to not only evaluate an ongoing
search process, but also predict how well it will unfold and
suggest a better path to the searcher if it is likely to
underperform. The motivation behind this investigation was
based on the following assumptions and realizations grounded in
the literature.
          </p>
          <p>
            The underlying rational processes involved in information
search are reflected in the actions users make while
searching. These actions include entering search queries,
skimming the results, as well as selecting and collecting
useful information [
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ][
            <xref ref-type="bibr" rid="ref14">14</xref>
            ][
            <xref ref-type="bibr" rid="ref15">15</xref>
            ].
          </p>
          <p>
            A searcher’s performance is a function of these actions
performed during a search episode [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ][22].
          </p>
          <p>With these assumptions, we propose to quantify a search process
using various user actions, and use it for user performance
(henceforth, ‘search performance’ or ‘performance’) prediction
as well as search process recommendations.</p>
          <p>BACKGROUND
Past research on predictive models that relates to the approach
we describe in this paper can be grouped into two main
categories: (1) behavioral studies and (2) IR approaches. In both
cases; however, the focus has been on end products instead of in
the process required to produce them.</p>
          <p>
            As far as the behavioral studies go, research has been conducted
to explore users models that help anticipating specific aspects of
the search process. One goal in this context has been the
determination of whether a search process will be completed in a
single or multiple sessions. For example, Agichtein et al. [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ]
investigated different patterns that can be identified in tasks that
require multiple sessions. As a result, the authors devised an
algorithm capable of predicting whether users will continue or
abandon the task. Similar work is described in Diriye et al. [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ],
which focuses on predicting and understanding of why and
when users abandon Web searches. To address this problem, the
authors studied features such as queries and interactions with
result pages. Based on this approach, the authors were able to
determine reasons for search abandonment such as accidental
causes (e.g. Web browser crashing), satisfaction levels, and
query suggestions, among others.
          </p>
          <p>There have been also attempts to understand past users'
behaviors in order to predict future ones in similar conditions.</p>
          <p>
            For example, Adar et al. [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] visually explored behavioral aspects
using large-scale datasets containing queries and other
information objects produced by users. The authors were able to
identify different behavioral patterns that seem to appear
consistently in different datasets. While not directly related to
performance prediction, this work focused on attributes of the
search process instead of in final products derived from it.
          </p>
          <p>
            Research like the ones described above often relies on historic
data from large populations and the use of trend and seasonal
components, which are used to model long-term direction and
periodicity patterns of time-series [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ]. For example, some have
explored seasonal aspects in Web search (e.g. weekly, monthly,
or annual behaviors) that provides useful information to predict
and suggest queries [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ].
          </p>
          <p>
            From an IR perspective, Radinski et al. [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] explored models to
predict users’ behaviors in a population in order to improve
results from IR systems. The authors also developed a learning
algorithm capable of selecting an appropriate predictive model
depending on the situation and time. As described by the
authors, applications of this approach could go from click
predictions to query-URL predictions. In contrast to this
approach, our method presented in this paper considers both the
population trends and an individual user behavior.
          </p>
          <p>
            In a similar track, several works have been conducted on query
performance prediction, focusing on developing techniques that
help IR system to anticipate whether a query will be effective or
not to provide results that satisfy users’ needs [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ][
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ][
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ]. For
example, Gao et al. [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ] found that features derived from search
results and interactions features offer better prediction results
than a prediction baseline defined in terms of query features.
          </p>
          <p>Results from this study have direct implications to individual
users by aiding the auto evaluation process of IR systems.</p>
          <p>In information search, users may be unaware of their individual
performance when solving an information search task. For
instance, Shah &amp; Marchionini [23] showed how lack of
awareness about different objects involved in searching (queries,
visited pages, bookmarks) could result in mistaken perception
about search performance during an exploratory search task.</p>
          <p>Even if an IR system is highly effective, users may run into
multiple query formulation and evaluation of several pages
before finding what they need. This process, which can be
related to search strategies, implies effort and time that is
usually underestimated by the users themselves. In this sense,
instead of predicting end products (i.e., overall performance),
the approach we introduce in this paper is oriented toward
predictions at different times in order to increase the level of
awareness of users about their own search process. Similar to
weather forecast, this information could help users to be aware
of possible trends based on past and current behavior.</p>
          <p>
            For a more recent discussion on IR evaluations and their
shortcomings, see [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. To the best of our knowledge, search
process performance prediction at different times from a user
perspective has not been explored. Similar approaches can be
found in weather and stock market studies. For example, using
machine learning approaches such as Support Vector Machine
(SVM), some models have been implemented to predict the
trends of two different daily stock price indices using NASDAQ
and Korean Stock prices [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ][
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]. In a similar fashion, our
approach is oriented to forecast users’ search performance
Nsteps ahead with the aim to aid their search process awareness
and performance trends.
          </p>
          <p>Unlike previous works in IR, we are not proposing to use time
series analyses or seasonal components of historic data. Instead,
we investigate predictive models based on machine learning
(ML) techniques; namely: SVM, logistic regression, and Naïve
Bayes which are trained over a set of features such as time,
number of queries, and page dwell time. In contrast to most IR
evaluations, our method focuses on user-processes. Also, unlike
most recommender systems, our approach could output
alternative strategies instead of similar/relevant products to help
the searcher. In essence, the work reported here takes several
lessons from tradition IR evaluations, recommender system
designs, and weather/stock forecasting to come up with a new
approach for evaluating and predicting search performance.</p>
          <p>In the next section we provide a detailed description of our
method, feature selection, and the measures we used in order to
create ML-based predictive models.
3 METHOD
In order to analyze the search processes followed by different
users/teams, we assume that the underlying dynamics of the
search processes are expressed by a collection of activities that
take place from the beginning to the end of the search processes.</p>
          <p>The first part of our method is a feature extraction step in which
we extract a wide array of features relating to webpages, queries
and snippets saved from the search processes for each unit of
time t. This step is performed in order to evaluate how well we
could use those features to capture the underlying dynamics
which would lead to recognizing whether a search process is
going to lead to high or low performance in the future time steps
at t+n (n=1,2,….,N), where N is the furthest time step.</p>
          <p>
            The decision to include or exclude a feature was based on
literature (e.g., [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ]) as well as our past experience [22] with
representing and evaluating search objects and processes. Each
feature is extracted for each user or team, u, up to time t from
the search processes and they are explained in detail as follows.
•
•
•
•
•
          </p>
          <p>Total coverage (u,t): The total number of distinct
Webpages visited by a user (u) up to time t. This feature
captures the Webpage based activity performed by a user
and provides a measure to see how much distinct
information has been found by the user up to this time.</p>
          <p>
            Useful coverage (u,t): The total number of distinct
webpages in which a user spent at least 30 seconds, up to
time t. This measure evaluates out of the total pages he/she
has visited how many of them were useful in finding
relevant information leading to satisfaction with their
context in completing the exploratory task [
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ][22][25].
          </p>
          <p>Number of queries (u,t): Total number of unique queries
executed by a user up to time t. This feature implicitly
relates to how much effort and cognitive thinking a user has
put in to this task.</p>
          <p>Number of saved snippets (u,t): Total number of snippets
saved by user u up to time t. This measures the amount of
information that the user thought that might be relevant in
the future to complete the task and needed to be
remembered. In other words, this feature is an indication of
explicit relevance judgments made by the user.</p>
          <p>Length of Query (u,q,t): Length of each query(q) executed
by a user u based on the character count of the query up to
time t. This feature captures how the user imposed the
•
queries and how long they were at different times of the
search process.</p>
          <p>Number of tokens in each query (u,q,t): This is the count of
tokens/words in each query(q) executed by user u up to
time t. This query based measure takes into account how
specific a user was in defining the query. By inspecting the
datasets, we realized that queries with a less number of
tokens tend to get general results. On the other hand,
composed queries with multiple terms are related to more
specific searchers. We also observed that typically the users
started with general queries with few words at the
beginning of the search process but then went into more
detailed queries to find more specific information later. For
all these reasons, we found it to be useful to capture the
number of token used in a query.</p>
          <p>Query entropy (u,q,t): This measures the information
content in a given query (q), by finding the expected value
of information contained in a query. We used the widely
recognized notion of Shannon entropy [24] in Information
Theory to calculate the information content of a query. We
calculated the number of unique characters appearing in
each of the queries, which represent the observed counts of
the random variable. This was used as the input to Shannon
entropy calculation and we used to the
maximumlikelihood method to calculate the entropy. Query entropy
feature has been used in the past to predict goodness of a
query for making query expansion decision [21].</p>
          <p>
            The method used to assess the search performance of a user is
described below. We define a measure called Efficiency (u,t), for
each user u up to time t in order to predict whether a given
search process is going to yield in high/low performance in the
future We first define Effectiveness of user u up to time t as the
ratio of useful coverage and total coverage (both defined
earlier). A similar measure was used in [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] and [22].
          </p>
          <p>Effectiveness(u,t) =</p>
          <p>Useful coverage(u,t)
Total coverage(u,t)
(1)
(2)
We then calculated Efficiency as defined in Equation 2.</p>
          <p>Efficiency(u,t) =</p>
          <p>Effectiveness(u, t)</p>
          <p>NumberofQueries(u,t)
In other words, Efficiency is defined as the Effectiveness
obtained per query, or how effective a query is in terms of
achieving a certain level of useful coverage.</p>
          <p>The performance for each user u at each time t was classified in
to the two classes; high performance and low performance based
on the following criteria:
Class = {
high ;if
low
;else</p>
          <p>Efficiency(u, t) ≥ Efficiency(u, t)</p>
          <p>(3)
Using various user studies data available to us, we constructed
feature matrices which consist of all aforementioned features for
each minute of time t for all the users in each dataset, and
converted in to a long vector of features which we fed as the
input to the classification models used.1 The class labels were
generated as high/low performance at minute t+n based on the
1 In the interest of space and scope of work here, details of these
experiments have been omitted, but will be available for
discussion at the workshop.
above mentioned criteria and threshold and used as the output
class labels to be used in the n-step ahead prediction model. If a
class label at n-step ahead was correctly predicted based on the
features extracted up to time t from the classification model it
was considered as correctly classified and if not as misclassified.
4 EXPERIMENTS
In order to evaluate whether users who are predicted to perform
at low performance in the future based on the current search
process, could benefit from this analysis to improve their search
process, we conducted some simple simulation analysis.</p>
          <p>We considered the individual user search processes as a
collection of search paths, where each search path is defined as
the search process from the time a user issued a query up to the
time user issued another quite different query. This was found
out using generalized Levenshtein (edit) distance, which is a
commonly used distance metric for measuring the distance
between two character sequences. If the Levenshtein (edit)
distance between two subsequent queries were greater than 2
(assuming less than 2 was when there were changes in the
queries due to simple spelling mistakes or refining of the query),
we considered the search process from the former query to the
next query as a single search path.</p>
          <p>Following this method, we found the first search path of each
user and based on the features extracted up to the end of the first
search path, and based on the classification model learnt from
that corresponding n-step ahead prediction we predicted whether
the user is going to have low/high performance at the end of the
session. If the user was going to have low performance, then out
of the users who predicted to have high performance, we looked
at which high performing user has the lowest Levenshtein (edit)
distance between the queries issued by low performing user
within the first search path and considered it as a pair of users,
whom we are going to use in the simulation. Then, for each low
performing user and high performing user that was matched, we
switched the search process of low performing user at the end of
the first search path with the high performing user’s search path
up to t=T minutes, where T is the total number of minutes for a
session. Then we evaluated by switching the search process
early during the overall process whether it would benefit each
low performing user to improve their performance. We found
that we were able to move most of the underperforming search
processes to higher performance by early detection and
switching, while keeping the higher performing processes
unharmed.</p>
          <p>These simulations provide verification that by realizing early
during the search process whether a user is going to perform
well or not, one could recommend better search
processes/strategies for that user which would lead to uplifting
the search performance of a previously destined to low
performing user.
5 CONCLUSION
When it comes to prediction, information retrieval and filtering
systems are primarily focused on objects while assessing what
and if something could help the users. These approaches are
often system-dependent even though the process of information
seeking is usually user-specific. Personalization and
recommendations are frequently exercised as methods to address
user-specific IR and filtering, but still limited to comparing and
recommending objects, not focusing on underlying IR processes
that are carried out by the searchers. We presented a new
approach to address these shortcomings. We began by asking
whether we could model a user’s search process based on the
actions he/she is performing during an exploratory search task
and forecast how well that process will do in the future. This
was based on a realization that an information seeker’s search
goal/task can be mapped out as a series of actions, and that a
sequence of actions or choices the searcher makes, and
especially the search path he/she takes, affects how well he/she
will do. Thus, in contrast to approaches that measure the
goodness of search products (e.g., documents, queries) as a way
to evaluate the overall search effectiveness, we measured the
likelihood of an existing search process to produce good results.</p>
          <p>Here we presented simulations to demonstrate what could
happen if one can make process-based predictions, but one could
develop an actual recommender system using the proposed
method. Another potential application of such prediction-based
method would be to use such approach in IR systems to provide
the awareness to users how their future performance will be
based on the current/past search process. The system could
identify that a user will have low performance if, he continues
this manner at an early stage of the process, and what could be
done to provide suggestions to improve overall performance.</p>
          <p>Given that the proposed technique is independent of any specific
kind of system, and solely focused on user-based processes, it
will presumably be easy to apply it to a variety of IR systems
and situations irrespective of retrieval, ranking, or
recommendation algorithms. Finally, while we have used
datasets borrowed from previous user studies, one could easily
apply the proposed method to Web logs, TREC data, and other
forms of datasets with various user actions recorded over time.
6 ACKNOWLEDGEMENTS
The work reported here is supported by The Institute of Museum
and Library Services (IMLS) Cyber Synergy project as well as
IMLS grant # RE-04-12-0105-12. The author is also grateful to
his PhD students Chathra Hendahewa and Roberto
GonzalezIbanez for their valuable contributions to this work.
7
in collaborative IR systems. In Proceedings of the 75th Annual
Meeting of the Association for Information Science and</p>
          <p>Technology (ASIS&amp;T). Baltimore, MD, USA.</p>
          <p>Directly Evaluating the Cognitive Impact of Search User
Interfaces: a Two-Pronged Approach with fNIRS
Horia A. Maior1,2, Matthew Pike1, Max L. Wilson1, Sarah Sharples3
1Mixed Reality Lab, 2Horizon DTC, 3Human Factors - School of Engineering</p>
          <p>University of Nottingham, UK
{psxhama,psxmp8,max.wilson,sarah.sharples}@nottingham.ac.uk
ABSTRACT
Recent research has pointed towards further understanding
the cognitive processes involved in interactive information
retrieval, with most papers using secondary measures of
cognition to do so. Our own research is focused on using direct
measures of cognitive workload, using brain sensing
techniques with fNIRS. Amongst various brain sensing
technologies, fNIRS is most conducive to ecologically valid user
studies, as it is less a↵ ected by body movement and can be worn
while using a computer at a desk. This paper describes our
two pronged approach focusing on a) moving fNIRS research
beyond simple psychological tests towards actual interactive
IR tasks and b) evaluating real search user interfaces.</p>
          <p>Categories and Subject Descriptors
H5.2 [Information interfaces and presentation]:
Evaluation/methodology, Theory and methods
Keywords
Functional near-infrared spectroscopy(fNIRS), Brain-computer
interface(BCI), Human cognition, Information processing
system, Multiple resource model, Limited resource model
1. INTRODUCTION</p>
          <p>
            The cognitive aspects of Information Retrieval (IR) have
repeatedly received focus over time, from Ingwersen’s
Cognitive Model [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ], to recent analyses of cognitive workload
during search tasks [
            <xref ref-type="bibr" rid="ref10 ref2 ref26">2, 10</xref>
            ]. The recurring interest is in what
users think about at di↵ erent task stages, and how much
mental workload is involved. The benefits of knowing more
about the searcher’s cognitive state would come from
providing better support for their needs, with Wilson et al
suggesting that better designed Search User Interfaces (SUIs)
could reduce unnecessary workload on the user [23].
          </p>
          <p>
            Although some prior work (e.g. [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]) have used indirect
techniques to analyse workload during search tasks, the
decreasing cost of brain sensing hardware has meant that more
recent research is using more objective techniques. Pike et
al [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] and Gwizdka et al [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ] used EEG technology, while
Moshfeghi et al used fMRI to measure workload when
making relevance judgements [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. Each of these technologies
have known limitations for studying actual interactive IR
behaviour, with EEG being highly a↵ ected by even tiny body
movement, and fMRI requiring users to lay in tunnel void
of any metal objects. Recent Human-Computer Interaction
research has listed the benefits of fNIRS brain sensing
techniques, which are less a↵ ected by body movement, and can
be more easily used in ecologically valid study conditions.
          </p>
          <p>Functional Near Infrared Spectroscopy (fNIRS) is an
emerging neuroimaging technique that is non-invasive, portable,
inexpensive and suitable for periods of extended
monitoring. fNIRS measures the hemodynamic response - the
delivery of blood to active neuronal tissues. fNIRS is designed
to be placed directly upon a participants scalp, typically
targeting the prefrontal cortex. This paper describes our
two-pronged approach to using fNIRS to study the
cognitive workload created by SUIs, focused on a) task analysis
and b) SUI analysis.</p>
          <p>RELATED WORK</p>
          <p>
            Understanding the cognitive aspects of interactive
searching (as well as interaction in general) has been a long-standing
goal for researchers in the field of Interactive IR. In the 1970s
Bates suggested that searchers employ both search tactics
and idea tactics [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ]. In an attempt to explain an individual’s
path during IR, Bates’ “Berrypicking” model [
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ] argued that
search will vary as the user recognises information and has
new ideas and questions.
          </p>
          <p>
            In the main cognitive evolution of information seeking
research, Ingwersen proposed a cognitive model of IR [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ],
where the searcher’s understanding of the document
collection, system, and task that would determine which path a
search would take. The model again put the user’s cognition
as the central point of interest. More recently, Joho [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]
argued that the cognitive e↵ ects typically observed in
Psychology could provide a potential building block of theoretical
development for evaluating interactive IR. Back et al [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ], for
example, examined the cognitive demands on users during
the relevance judgement phase, suggesting that the amount
of workload involved was the reason behind searchers rarely
providing relevance judgements in previous work. Using a
secondary measure, the Stroop task, Gwizdka [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ] mapped
varying levels of workload at multiple stages of search.
          </p>
          <p>
            More recently, researchers have focused on objectively
measuring interactive IR phases, in line with Back et al’s work,
Moshfeghi et al measured workload during relevance
assessments by asking people to make judgements while lying in
an fMRI machine. As making relevance judgements can be
performed without directly interacting with a computer, this
made use of an fMRI machine more realistic. Using more
commercialised tools, Anderson [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] used an EEG sensor to
compare visualization techniques in terms of the burden they
place on a viewer’s cognitive resources. Similarly, Pike et al
[
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] developed a prototype tool named CUES that was
capable of collecting a variety of data including EEG whilst
interacting with a website. Pike et al used this to
monitor aspects such as frustration and concentration, but their
work demonstrated the variability of EEG data across the
several minutes involved in an interactive IR task.
          </p>
          <p>
            Using fNIRS, as introduced above, Peck [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] performed a
similar study of di↵ erent visualisation techniques, while a
system called Brainput [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] was able to identify and
correlate brain activity patterns among users during multitasking
studies, and intervene when it sensed workload exceeding a
certain level. Our work intends to build upon these HCI
studies, to study interactive IR tasks and SUIs in more
ecologically valid user study situations.
          </p>
          <p>RESEARCH PATHS</p>
          <p>
            Pike et al [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] highlighted the challenges of using brain
sensing technologies to evaluate IIR tasks: that tasks have
di↵ erent stages, that behaviour quickly diverges after the
first interaction (and thus is hard to compare), and that
brain measurements vary dramatically over time. In order
to address these challenges, we have initiated two clear
research paths, both utilising fNIRS technology: 1) evaluating
the cognitive aspects of Interactive IR tasks and 2)
methods to evaluate the design of SUIs. The aim of the first
path, is to move beyond using fNIRS to measure workload
in simplistic psychology memory tasks (like Peck et al [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]),
towards being able to break down real search tasks into
primary components. This implies three considerations:
• Collected data would be meaningless if is not related
to existing knowledge. Therefore, to interpret sensed
fNIRS data we use proposed theories and models.
• It is known that fNIRS can sense cognition information
[
            <xref ref-type="bibr" rid="ref16">19, 16</xref>
            ] related to so called working memory (if placed
on the forehead). Assuming this is correct, we are
using models of working memory.
• The proposed models will help us interpret the sensed
data with fNIRS and have a better understanding of
the cognitive impact of various complex tasks (such as
a IR).
          </p>
          <p>Such a technique would allow researchers to analyse data by
stage, and find e↵ ective points of comparison during several
minutes of continuous measurements. The second path is
focused on identifying which aspects of working memory are
a↵ ected by di↵ erent features of SUIs, such that researchers
can objectively evaluate the e↵ ect of di↵ erent SUI design
decisions. A combination of both paths works towards being
able to proactively evaluate how SUIs support searchers.</p>
          <p>PATH 1: WORKLOAD MODELS</p>
          <p>To understand the cognitive aspects of IIR, it is essential
to learn about user’s capabilities and limitations in terms
of their cognition: how people perceive, think, remember,
and process information. This path of research focuses on
existing models from Cognitive Psychology and Human
Factors, models that conceptualize and highlight aspects that
typically describe or influence elements of human cognition.</p>
          <p>One important part of cognition during interactive
searching involves human memory systems. There are two
different types of memory [21]: working memory (sometimes
called short-term memory) and long-term memory.
Wickens describes working memory as the temporary holding of
information that is “active”, while long-term memory
involving the unlimited, passive storage of information that is not
currently in working memory.</p>
          <p>
            Working memory. Working memory, proposed by
Baddeley and Hitch (1974) [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ], refers to a specific system in the
brain which “provides temporary storage and manipulation
of information...” [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ]. Working memory [
            <xref ref-type="bibr" rid="ref20 ref21 ref22 ref4 ref5 ref6">6, 4, 5</xref>
            ] processes
information in two forms: verbal and spatial, and has four
main components (Figure 1):
• A central executive managing attention, acting as
supervisory system and controlling the information from
and to its “slave systems”.
• A visuo-spatial sketch pad holding information in
an analogue spatial form (e.g. Colours, shapes, maps,
etc.), specialised on learning by means of visuospatial
imagery.
• A phonological loop holding verbal information in
an acoustical form (e.g. Numbers, words, etc.);
specialised on learning and remembering information
using repetition.
• A episodic bu↵ er dedicated to linking verbal and
spatial information in chronological order. It is also
assumed to have links to long-term memory.
          </p>
          <p>Information processing system. As humans, we are
exposed to large amounts of information via our sensory
systems. One of our strengths is in selecting information
from our environment, perceiving it, processing it, and
creating a response. Therefore we can use this understanding
of brain activity to identify which elements of an
interactive IR environment need to be considered when measuring
brain activity, and how we can reduce rather than increase
a user’s mental workload via interface and system design.</p>
          <p>Wicken’s Information Processing Model [21] aims to
illustrate how elements of the human information processing
system such as attention, perception, memory, decision
making and response selection interconnect. We are interested in
observing how and when these elements interconnect during
IR. He describes three di↵ erent ‘stages’ (see STAGES
dimension in Figure 2) at which information is transformed:
a perception stage, a processing or cognition stage, and a
response stage, the first two being processes involved in
cognition. The first stage involves perceiving information that
is gathered by our senses and provide meaning and
interpretation of what is being sensed. The second stage represents
the step where we manipulate and “think about” the
perceived information. This part of the information processing
system takes place in working memory and consists of a
wide variety of the mental activities. In relation to IR, it
is interesting to observe how elements of cognition, such as
rehearsal of information, planning the search strategy and
deciding on the search keywords interconnect.</p>
          <p>Multiple Resource Model. One model of mental
workload that has been widely accepted in Human Factors is
Wickens Multiple Resource Model [20] (Figure 2). The
elements of this model overlap with the needs and
considerations of evaluating complex tasks (such as IR). He describes
the aspects of human cognition and the multiple resource
theory in four dimensions:
• Avoid unnecessary zeros in codes to be remembered;
• Encourage regular use of information to increase
fre</p>
          <p>quency and redundancy;
• Encourage verbalization or reproduction of
informa</p>
          <p>tion that needs to be reproduced in the future;
• Carefully design information to be remembered;
Resource vs Demands. One other model that is of
interest is the limited resource model [22] describing the
relationship between the demands of a task, the resources allocated
to the task and the impact on performance.
• The STAGES dimension refers to the three main stages</p>
          <p>of information processing system (Wickens, 2004 [21]).
• The MODALITIES dimension indicating that
audi</p>
          <p>tory and visual perception have di↵ erent sources.
• The CODES dimension refers to the types of memory</p>
          <p>encodings which can be spatial or verbal.
• The VISUAL PROCESSING dimension refers to a nested
dimension within visual resources distinguishing
between focal vision (reading text) and ambient vision
(orientation and movement).</p>
          <p>Our aim is to understand how these elements link together
and compose more complex components/tasks. Additionally
we want to consider how complex tasks (such as a search
task) can be divided into primary components according to
the models described. This will help identify possible
problems in SUI design as well as indicating a possible solution
to the problem (suggested implications by Wickens [21]):
• Minimize working memory load of the SUI system and</p>
          <p>consider working memory limits in instructions;
• Provide more visual echoes (cues) of di↵ erent types</p>
          <p>
            during IR (verbal vs spatial);
• Exploit chunking (Miller, 1956 [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]) in various ways:
physical size, meaningful size, superiority of letters
over numbers, etc;
• Minimize confusability;
          </p>
          <p>The graph from Figure 3 is used to represent the
limited resource model. The X-axes represent the resources
demanded by the primary task and as we move to the right
of the axes, the resources demanded by the primary task
increase. The axes on the left indicate the resources being
used, but also the maximum available resources point (if we
think of working memory that is limited in capacity). The
right axes indicate the performance of the primary task (the
dotted line on the graph). The key element of this model is
the concept of a limited set of resources which, if exceeded,
has a negative impact on performance. However, it does not
distinguish between resource modality, therefore we propose
to use both the limited and multiple resources models to
inform our work.</p>
          <p>PATH 2: SUI EVALUATION</p>
          <p>Relating quantitative data from brain sensing devices into
feedback about SUI designs is one of our ultimate goals in
conducting this research. SUIs are inherently information
rich and thus a↵ ect both visual (results page layout) and
verbal (text based results) memory. Detecting a change in
either verbal or spatial working memory would help determine
if a workload di↵ erence was caused by SUI design (spatial)
or the amount of information the design provides (verbal).</p>
          <p>
            Our first in-progress study has stimulated each memory type
in di↵ erent tasks - Verbal memory was tested by performing
an n-back [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] number memory task, whereas spatial
memory was tested using an n-back visual block matrix task.
          </p>
          <p>
            Other studies have also looked at each type of memory and
confirmed fNIRS ability to detect changes in heamodynamic
responses accordingly [
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ].
          </p>
          <p>In addition to developing an understanding of the
extent to which we can monitor di↵ erent memory, our
initial study also sought to measure the e↵ ect of artefacts on
the fNIRS data. Controlling the environment and human
derived sources of noise is a potentially di cult factor to
control without e↵ ecting the ecological validity of a study.</p>
          <p>
            Solovey et al [19] showed that fNIRS is relatively resilient to
motion derived artefacts when compared to EEG [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] for
example, but still required some consideration by researchers
conducting studies. In our own experience, we found that
asking participants to remain still as much as possible was
fairly successful. We are additionally looking at possible
methods for correcting motion derived artefacts using an
external gyroscope connected to the participant.
          </p>
          <p>Designing tasks for experiments that measure cognitive
effect via a brain sensor require careful consideration in order
to ensure that results can be attributed to a cause.
Thankfully this problem space has been well explored in the field
of Psychology and we are able to adapt the approaches
described in the literature to suit our task type requirements.</p>
          <p>
            A primary example of this adaptation is demonstrated by
Peck et al [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ], where 2 data visualisations techniques were
compared using a methodology based loosely on the n-back
task - a widely used psychology task that is designed to
increase load on working memory.
          </p>
          <p>
            Additionally, we are interested in exploring standard search
studies (without following a psychological study layout) and
seeing whether interesting states can be detected. Solovey
et al [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] performed a similar function by utilising a
machine learning algorithm that had classified “states of
interest” prior to performing a task.
          </p>
          <p>Using a similar approach, we could evaluate a SUI to
determine whether a particular change in layout has a positive
or negative impact on visual memory. Alternatively, to test
the relevance of a results page (which would be dependant
on the textual results), we could analyse the e↵ ects on verbal
memory between 2 varied results pages, we could then
reflect these changes to the Wickens Multiple Resource Model
[20]. We are also working towards enabling the
interpretation of data within the context of complex multimodal tasks
to further extending our knowledge of the processes involved
during IR and how they interact and e↵ ect one another.
6. SUMMARY</p>
          <p>This paper has aimed to summarise our two-pronged
approach towards actually evaluating the design of search user
interfaces, in realistic ecologically valid study conditions,
using fNIRS technology. The approach first involves braking
down interactive IR tasks into how they e↵ ect the di↵
erent elements of working memory, and second understanding
how SUIs are processed by di↵ erent parts of working
memory. Our two paths of research will build towards a stage
where we can combine them and objectively evaluate
cognitive workload involved in interactive IR. We believe that this
research will provide a novel new direction that SUI’s and
indeed HCI in a broader sense can benefit from. The
association of physical recordings in ecological valid settings, to
an existing theoretical model, provides a new measure from
which future SUI development and evaluation could benefit.</p>
          <p>Dynamics in Search User Interfaces
Marcus Nitsche, Florian Uhde, Stefan Haun and Andreas Nürnberger</p>
          <p>Otto von Guericke University, Magdeburg, Germany
{marcus.nitsche, stefan.haun, andreas.nuernberger}@ovgu.de,</p>
          <p>florian.uhde@st.ovgu.de
ABSTRACT
Searching the WWW has become an important task in today’s
information society. Nevertheless, users will mostly find static search
user interfaces (SUIs) with results being only calculated and shown
after the user triggers a button. This procedure is against the idea
of flow and dynamic development of a natural search process. The
main difficulty of good SUI design is to solve the conflict between
good usability and presentation of relevant information. Serving a
UI for every task and every user group is especially hard because
of varying requirements. Dynamic search user interface elements
allow the user to manage desired information fluently. They offer
the possibility to add individual meta information, like tags, to the
search process and enrich it thereby.</p>
          <p>Keywords
Search User Interface, User Experience, Exploratory Search.</p>
          <p>Categories and Subject Descriptors
H.3.3 [Information Storage and Retrieval]: Information Search
and Retrieval.; H.5.2 [Information Interfaces and Presentation]:
User Interfaces.</p>
          <p>
            General Terms
Design, Human Factors, Management.
1. MOTIVATION
Since the launch of the WWW, users accumulated a vast amount of
information. With broadband technologies becoming a part of
everyday life1 the WWW offers a great opportunity in terms of
learning and education. University courses, for instance, are available
online and nearly every topic is handled somewhere in the great
amount of blogs, Q&amp;A pages, fora, web pages or databases. Yet
there is no map, no guide leading through this vast amount of
information. Users need to search for information, to locate the bits
fitting to their specific information need, indexing the amount of
1http://www.internetworldstats.com/images/
world2012pr.gif, 02.05.2013
knowledge available online. Therefore, a proficient tool to
analyse the structure of the web and to provide guidance to specific
sources of information is needed. This task is accomplished by
modern search engines like Google2, Bing3, Yahoo4 and other
local or topic centred search engines. By the increase of
computational power in smart phones and wider access to online resources
the demand for these search tools has risen and the quality of the
search terms has changed. Instead of single-query-searches, users
tend to request complex answers5, trying to learn about topics in
deep. While the need for information and the expectations of users
increased, matching the broader knowledge base contained in the
Internet in the last few years. About 300 Mio. websites were added
in 20116. Search engines mainly remain the same. This leads to the
fact that a “significant design challenge for web search engine
developers is to develop functionality that accommodates the wide
variety of skills and information needs of a diverse user population”
[
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. Therefore, this paper proposes the concept of using dynamic
elements in SUIs, that focus on fluent work flow characteristics, a
high grade of interactivity and an adequate answer-time-behaviour.
2. INFORMATION GATHERING
Looking at users’ habits in search, they no longer perform
simple lookup searches. There is an increasing need to answer
complex information needs. Therefore, we mainly consider
information gathering processes, searches where users are not familiar with
the domain. Users need to refine search queries, branch out into
other queries to gain additional understanding and collect results to
merge them into a single topic. This kind of search process is called
exploratory search and is contrary to a known-item search task as
stated in [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]. Exploratory search processes “depend on selection,
navigation, and trial-and-error tactics, which in turn facilitate
increasing expectations to use the Web as a source for learning and
exploratory discovery” [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ]. Search tasks are fragmented,
consisting of single queries and search requests. The search requests may
yield additional data or parts of the final information which in the
end form the information requested by the user. While
performing such a complex search task, a pattern called berry picking [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ]
can be observed. While reading through a source of data, looking
for qualified information the user discovers new traces leading to
other sources, which have to be handled one after the next. By
re2http://www.google.com, 02.05.2013
3http://www.bing.com, 02.05.2013
4http://www.yahoo.com, 02.05.2013
5see the 2009 HitWise study for more details:
//image.exct.net/lib/fefc1774726706/d/1/
SearchEngines_Jan09.pdf, 10.07.2013
6http://royal.pingdom.com/2012/01/17/
internet-2011-in-numbers/, 02.05.2013
fining the search and gaining deeper information the user satisfies
the initial need for it. These different traces span a map in the end,
representing the whole search and its processing. When someone
is learning about something this map is refined and expanded. The
learner may track back to a certain node and deepen the
understanding about it by adding new queries, and therefore new branches. Or
he may discard a whole part of the map because it turned out that
the contained information was not relevant to him. When the user
is satisfied with the gained information this map is encapsulated
and represents the whole development of this complex information.
          </p>
          <p>According to this concept the result is not a single object. It is a set
of sources, representing the learning process for a specific user.</p>
          <p>
            Looking at the current process of information gathering in the
Internet there are only two places. The Internet itself, containing the
pool of existing information, in an unstructured form and a mental
model about the information (space) that is constructed. This
system may work perfect when dealing with short, exact search queries
like postal code New York City, but when it comes to complex
information needs, where the user needs to access a lot of information
and generate more detailed search queries while looming through
pages this system reaches it boundaries. The user might retrieve
only partial facts. For example, if the user needs explanation of a
term used in its initial query. The user is now in need of another
place, where he can store information, reorder it and put it into the
context of other information pieces.
3. STATE OF THE ART
Looking at Google, the most used search engine today [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ], the user
interface of a modern search engine is mostly static. Google’s
features include some dynamic elements like real time search. For
example “[..] Google Suggest which interactively displays
suggestions in a drop-down list as the searcher types in each character of
his/her query. The suggestions are based on similar queries
submitted by other users.” [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] Dynamic previews of results will be offered
when clicking on the double arrow beside a result. But the core of
the interface has not changed a lot since its launch in 19977. While
adopting fast to new information sources like Facebook and Twitter,
Google discarded the adoption of new HCI methods in favour of a
clean, slim interface. With increasing touch support on the devices,
a richer user interface can be designed to provide the user with
immediate feedback and allows haptic interaction with the search
process. Some mobile clients take advance of the additional
information available, like the iOS search client, which switches to
voice queries when the phone is lifted to the head, but there is no
full extension of Google’s search services. While Google is an
adequate tool for short queries and queries calling for a direct answer,
features for deep research on complex topics are missing.
          </p>
          <p>One way to integrate dynamic elements into existing SUI
infrastructure is to build an overlay. Thereby, dynamic UI utilize existing,
well known search engines and provide a benefit by enriching them.</p>
          <p>
            This approach is shown in the Boolify8 search engine, which
provides a dynamic drag and drop interface on top of Google’s search
engine. This engine is relatively new and was build to promote the
understanding of boolean queries. Users build a query by
dragging jigsaw like parts onto a search surface. These parts contain
words (general or exact) and linkers like AND and OR. Additional
parts have been added to provide search on a specific page or for
7http://www.google.com/about/company/
history/, 02.05.2013
8http://www.boolify.org/, 02.05.2013
synonyms. By adding and linking those parts the user constructs
a boolean query which will be submitted to the Google search
engine. Boolify was built for children and elderly. Tests in a third
grade technology class showed that children without any
knowledge of boolean queries were able to construct complex queries
just by pulling them together piece by piece9. A similar approach
was implemented at SortFix10. This tool offers the user the
“ability to drag and drop search terms in between several buckets” [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ]
to in- and exclude them in the query. With a Standy Bucket users
are “able to keep track of all [their] inspirations and alternative
search words off to the side, ready to be dragged and dropped into
your search box if needed.” [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ] Another possible use of dynamic
interface elements is the weighting of search terms based on their
font size as used at SearchCloud.net11. The ranked keywords are
shown in a Tag Cloud like manner and additionally the site shows,
based on the ranking, “the calculated relevance score for each
[result]” [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ]. Not only the query building process can be enchanted
by dynamic elements, also the presentation of the result can benefit
from it. Dynamic side loading can provide the user a lens like view
to parts of the result where keywords occur. Microsoft’s WaveLens
“[...] fetches a longer sample for the page containing your
keywords, without you having to download it.” [
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ] Microsoft Research
shows that in a study using WaveLens, presenting the participants
with a normal interface and two versions of WaveLens’ UI (instant
zoom and dynamic zoom), “participants were not only slower with
the normal view than the other two, but they were more than twice
as likely to give up” [
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ]. Another way of result presentation was
shown at SearchMe12: “Fragmentation into multiple sites, domains
and identities becomes a huge distraction. User don’t know which
site to visit for which purpose, and the lack of consistent, intuitive
inter-site search and navigation makes it hard to find content [..]”
[
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ]. All these dynamic features can be used as a mask over
traditional SUIs to extend them. By hiding the dynamic part, dynamic
elements can be added to an existing search engine and let the user
make a choice which part should be shown and used. The proposed
concept is similar to Byström &amp; Hansen’s approach in [19].
          </p>
          <p>
            Issues. Comparing the state of the art with the process of
information gathering some issues appear, which may be resolved or at
least damped by using of dynamic elements. While collecting
information pieces for solving complex questions the user discovers
new sources, containing more information. These sources may not
form a linear search process every time. Sometimes there will be a
split and the user needs to decide which trace to follow first. This
issue is also noted in [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ]. Today’s search engines offer only little
support for this. The user needs to save web pages to favourites or
organize them himself for later reading. Searching different terms
one by one allows users to follow new pages like traces through
the Internet. By connecting these traces and setting them into
relation the user can retrieve the whole information needed to cover
his query. Most modern search engines discard this feature, it is
again something the user needs to do by himself. This leads to
another more general problem, the enclosing of search queries.
          </p>
          <p>
            Google for example handles every search term as a new
operation. Data is stored, but contains only general information about
the user, queries are not related to each other and therefore
miss9http://ed-tech-axis.blogspot.de/2009/03/
boolified.htm, 02.05.2013
10SortFix.com, offline since 11/2011, Firefox plugin:
https://addons.mozilla.org/en-us/firefox/
addon/sortfix-Extension, 02.05.2013
11http://searchcloud.net/, 02.05.2013
12http://www.searchme.com, offline since 2009
ing its broader context. But when learning about a complex topic
refining the search query is more important to the user. In the
iteration of search processes, to narrow down the mass of information
and to tap new sources, the searcher needs to rewrite and modify
the query, to link it to other related search tasks. Building a
connection between parts of information and evaluating it against each
other is a core principle of learning. This leaves the user targeting
a broader, intense search, in the need to build a custom solution to
extract knowledge and manage it. This is strictly against the
guideline for online interfaces which suggests to “[..] not require users
to remember information from place to place on a Web site” [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ] as
this is a distraction from the main process of searching and destroys
the interaction flow triggered by the search process.
4. COMPOSING A DYNAMIC SUI
The proposed approach shows a design based on today’s search
engines, enriched with dynamic UI elements to provide a plus for the
user. The design includes principles to form web based learning
applications [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ] to focus on the completion of complex search tasks.
          </p>
          <p>
            By adding dynamic elements internal states can be visualized for
the user to give a better overview about the current position in the
search process. Furthermore it will allow the serialization of search
processes and to step in at every point of the process later on. As
stated in Beyond Box Search “different interfaces (or at least
different forms of interaction) should be available to match different
search goals” and “[t]he interface should facilitate the selection of
appropriate context for the search” [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. Both of this quality
measurements should be regarded when conceptualizing a SUI. The
first point will be covered by a modular UI, the user may move,
hide and scale elements to fit his current need. The second point
is strongly bounded to the use of dynamic items in the UI design.
          </p>
          <p>By giving immediate feedback to the user it is easier to classify
the current results. The context of the whole search process will
be persistent over multiple search queries and provide a method of
accumulation parts of the search process into a single object.</p>
          <p>Four features are proposed and explained in this paper, showing a
use-case for dynamic search interfaces and giving a suggestion how
this can be accomplished. Together these features build up a mid
instance to accumulate into a bigger context for a search process.</p>
          <p>This clipboard (Fig. 1) reshapes the search process and provide the
place to store information between search queries. Instead of trying
to accumulate knowledge and information directly the user is able
to construct a solution of the search query in this buffer and save it
as a complete collection of the information retrieval process.</p>
          <p>Reordering. Giving users the opportunity to reorder and therefore
to rate a search result is an important step towards dynamics in
SUIs. Every result is handled as a single item and can be picked
by the user and dropped in another place. The other items reorder
fluently, giving user feedback while the user moves on. The SUI
holds an array of parameters, which is used to evaluate every item.</p>
          <p>
            Possible criteria are Accuracy, Clarity, Currency and Source
Novelty. These and more criteria are mentioned and explained in [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ].
          </p>
          <p>
            When a user reorders items to fit his preferences the search engine
may use the information provided by this ranking to weight the
existing parameters to yield better results in the future. The engine
will be able to present results ranked according to the user’s
preference. This can be done for all users and also search process wide, as
some search tasks require documents and papers while others may
focus on web pages or media. This addition to classical user
interfaces can make great use of the up-trend for touch based devices, in
2012 89% of mobile phones and smart-books support touch [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ].
          </p>
          <p>Designing the SUI responsive to touch and gesture is maybe one
of the most natural solutions for human computer interaction and
adds an amount of possible actions based on gestures.</p>
          <p>
            Workbench. The workbench targets the issue of loosing
information while switching between different searches. It adds a third
place to the proposed search process, located outside of the search
scope but still related to it. The user may drop queries here to keep
them throughout the whole search process. When entering a query,
indicators show how relevant items on the bench are. This allows
the user to classify new results in terms of integrity towards already
selected snippets. The workbench acts as a buffer between search
queries, adding a broader context to every entry. Like a frame, it
contains information exclusively attached to the current search
process, leading to the possibility of customization and user centred
search environments. When the user switches between queries he
can immediately determine how well the new results fit into already
selected items. This allows identifying false positive as well as
exploratory search [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] results. Users may just enter queries that lead
to a peripheral topic and check the indicators whether the result is
relevant to his initial information.
          </p>
          <p>
            Tag Cloud. The tag cloud is another feature to guide the user in the
search process. As shown in [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] a tag cloud supported retrieval
system can increase the find rate of adjacent data nodes by nearly
15%. When adding an item to the workbench its most relevant tags
are extracted and visualized in the tag cloud. It is able to show how
often a tag occurs and how different tags are related to each other.
          </p>
          <p>When entering a new search query the tag cloud displays the
relevant tags and reorders the cloud to revolve around the current tags.</p>
          <p>By combining distance and size of the entered tag with their direct
neighbours the user can directly spot how homogeneous its current
query is in terms of the whole process. The tag cloud can also use
the existing tags to show the user other closely related tags and
suggest query refinement based on tag proximity. Colours can indicate
the state a tag is currently in. A possible color scheme for western
culture can be based on the three colors used in traffic lights. The
concept of three-coloured traffic lights also work for color-blind
people, since they do have a given position. Therefore, we also use
second coding paradigm: form. A green triangle is proposed for
tags resulting from the current query, which are contained in the
overall tag cloud spanned by the workbench. An orange circle
indicates a warning for tags, either in the current query result or the
bench, which are not related to the rest of the cloud. A red square is
avoided for the reason that uncontained tags may not be bad, they
can lead to a new direction or add a reasonable value to the whole
search process. The tags are scaled depending on their frequency.</p>
          <p>When the user selects any item from the bench or the search
result the corresponding tags are centred. The other tags are located
based on their coherence with the selected tags; closer means the
tag is in a direct relation to the selected item. A user can quickly
check the integrity of his search process by looking at the tag cloud.</p>
          <p>A slim, packed cloud means the results are all related to each other,
an open, wide cloud indicates a broad result field, covering many
aspects. False positives may be filtered out, when enough items
exist, as they stick out the rest of the cloud.</p>
          <p>
            Search Map Support. The search map (Fig. 2) acts as a
representation of the whole search process, by storing every query and
following up querying and visualize it in a chronological order. The user
may select single nodes in the map to get into the state of search
process at this moment and refine it. The map provides a kind of top
view to the path of the search and shows where the user branched
out into new queries. It allows the user to cut off nodes and whole
branches if they are not needed any more to fulfil the need for
information. As it contains every action and some data in the current
search process, the search map might be serialized and stored to
retrieve the search process later on. With this map at hand a user can
save whole search tasks just like he saves favourite web pages. He
can step back into the process at any time and reconstruct the whole
learning process or correct parts of the search which has proven to
be not correct. This kind of Story Telling helps to visualize the
given data, “[...] lead to findings, which prompt actions [...] [and]
can indicate the need to forage for new data.” [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] The search map
[
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] features two ways of expanding. The user may follow a result to
expand it vertically. The result is added as a new node and resides
in the map until it is processed further. When the user selects an
existing node he steps back to the vertical position of this node and
can now branch out horizontally. This deals with an issue of
berrypicking [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ], where the new sources has to be processed one by one.
          </p>
          <p>While not abolishing this the search map provides a visual
representation to simulate parallelism. The map also allows scoping of
the analysis by creating a horizontal or vertical bound. Only tags
and items inside this bound will be considered, the rest is greyed
out. This allows the user to dig deep into a certain topic (small
vertical bounds) or create a better understanding of a certain term
and add more results to a certain query (horizontal boundary). This
can help the user to concentrate on smaller pieces of a big search
process and to narrow down problems one by one.
5. CONCLUSION
This paper has shown certain design flaws of today’s search engines
and some proposed dynamic design principles to counter them. The
application of the envisioned elements can extend a search engine
towards a software capable of complex research tasks. With the
current up-trend of online learning this unlock a new way of using
them. The surplus resides not only in the dynamic and vivid
interface, it prepares a whole new tier of online search solutions. The
process of learning can be preserved and shared with others. One
can come back at any time, jump right into the saved search process
and reconstruct the development of certain knowledge. With this
tool chain at hand learning becomes a social and an integrative part
of the WWW. The next step in deploying dynamic elements into
search user interfaces would be prototyping them. Design snippets
need to be tested for usability and acceptance in the real world.</p>
          <p>Starting as overlays and additional feature of existing search
engines may develop and emerge into independent solutions.</p>
          <p>Acknowledgement
Part of the work is funded by the German Ministry of Education and
Science (BMBF) within the ViERforES II project (01IM10002B).</p>
          <p>SearchPanel: A browser extension</p>
          <p>for managing search activity</p>
          <p>Simon Tretter</p>
          <p>University of Amsterdam
Amsterdam, The Netherlands</p>
          <p>Gene Golovchinsky
FX Palo Alto Laboratory, Inc.</p>
          <p>3174 Porter Drive</p>
          <p>Palo Alto, CA
gene@fxpal.com</p>
          <p>Pernilla Qvarfordt
FX Palo Alto Laboratory, Inc.</p>
          <p>3174 Porter Drive</p>
          <p>Palo Alto, CA
pernilla@fxpal.com
ABSTRACT
People often use more than one query when searching for
information; they also revisit search results to re-find
information. These tasks are not well-supported by search
interfaces and web browsers. We designed and built a Chrome
browser extension that helps people manage their ongoing
information seeking. The extension combines document and
process metadata into an interactive representation of the
retrieved documents that can be used for sense-making, for
navigation, and for re-finding documents.
1. INTRODUCTION</p>
          <p>
            Broder et al. [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ] proposed a taxonomy of web search that
included transactional and navigational searches in addition
to the more traditional (from an IR perspective)
informational searches. To this taxonomy we might add re-finding
[
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ], the task of locating a previously-found document.
          </p>
          <p>From a theoretical perspective, it is not clear whether
refinding is a di↵ erent kind of search activity or an orthogonal
dimensions. Regardless, while major web search engines o↵ er
simple and e cient interfaces for navigational and
transactional searches, relatively little support is available for more
complex informational search or re-finding.</p>
          <p>
            These seemingly neglected activities are not unimportant,
however: Teevan et al. [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] reported that 39% of queries are
re-finding queries; furthermore, 20-30% of searches represent
open-ended informational needs [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. Related, Qvarfordt et
al. [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ] found query overlap rates of 50-60% in exploratory
search, and suggested that awareness of this overlap may be
useful in supporting more e cient searching behavior. Thus
we decided to explore ways in which searchers’ interactions
with search engines could be enhanced to support these more
complex information-seeking tasks.
          </p>
          <p>We created a web browser extension that enriches
common web search engine interfaces and addresses important
deficits with respect to open-ended (exploratory) search and
re-finding. Our extension visualizes search results to help
users find the right document or documents by visualizing
metadata of the retrieved pages.</p>
          <p>
            Following Golovchinsky et al. [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] we distinguish
document metadata from process metadata. Document metadata
– dates of publication, titles, hosting web sites, etc. – are
basic characteristics of documents that are independent of
the means by which these documents were retrieved.
Process metadata, on the other hand, characterize aspects of
documents in relation to the searcher’s activity: how many
times was a document retrieved, whether it was viewed
before, etc. This kind of information can help searchers to
remember, understand and plan their search processes.
          </p>
          <p>The browser plugin enhances the searcher’s ability to use
process metadata to understand their search results and to
plan subsequent activity by displaying surrogates for the
current set of retrieved documents. We represent prior
retrieval state, whether a document was opened, and whether
it was bookmarked in an integrated overview that appears
at the side of the browser window. We also make it
possible for searchers to examine multiple documents without
returning to the search results or using multiple tabs.</p>
          <p>The remainder of this paper is organized as follows: we
review the relevant related work, describe the browser
extension, and conclude with a discussion of the design space.</p>
          <p>RELATED WORK</p>
          <p>There are two broad categories of related work: the
management of search history and the representation of search
results. Refinding has received increasing attention recently.</p>
          <p>
            While the browser implements some history mechanisms,
these are typically not well-suited to users’ needs [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ].
Elsweiler and Ruthven [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ] described di↵ erent patterns of
refinding; Teevan [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] proposed a mechanism for merging
previously-found and newly-retrieved documents. More explicit
management of search history has also been investigated in
the literature; see [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ] for a succinct summary.
          </p>
          <p>
            Information overload due to large numbers of results is
a common problem in information seeking [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]. This
problem can be addressed in a variety of ways. MetaSpider
[
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ] uses a 2D map to display and classify retrieved
documents. Grokker [
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ] uses nested circular and rectangular
shapes to present results and also shows them in a
hierarachical grouped way. Sparkler [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ] uses a star plot for the
result presentation, where every star represents a document.
          </p>
          <p>One potential issue with the systems above is that the
overall organization of the interface itself may induce
usability problems. Complex interfaces allow more individual
settings to be specified by a user, but simple interfaces allow
a broader spectrum of users to use them. This tradeo↵ is
not trivial to handle, and as we see nowadays, most Web
search interfaces tend to be quite simple.</p>
          <p>
            Supporting the searcher’s decision making process can be
crucial for e↵ ective search performance for complex
information needs. This support can take the form of enhanced
surrogates for documents. One type of information often
used for this purpose is document metadata (author, date,
images of the document, etc.). Even et al. [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ] has shown
that the decision making process can be highly improved by
adding process metadata (in our case information that is
related to the search process) to the user interface. Research
has shown that presenting simple tasks in a slightly di↵
erent way may help the user to understand how the search
is performing and what can be done to gain better results
[
            <xref ref-type="bibr" rid="ref18">18</xref>
            ]. One common example of incorporating process
metadata in web browsers is the practice of changing the color of
a traversed link anchor.
          </p>
          <p>
            Spoerri [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ] showed that users can benefit from di↵ erent
or additional visualizations of web search results. However,
none of the techniques above have been integrated by major
search engines into their main interfaces. In some cases,
extension developers have enhanced the user experience of web
search. Examples include: SearchPreview[
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ] that fetches
screen shots of the result pages and shows them directly
next to the each search result. Bettersearch[
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] is a Firefox
extension that performs a similar task, but also enriches the
result page with more features and links. For example, this
extention allows users to open a result in a new tab, or adds
links to a search result to quickly show the web page on the
”Wayback Machine”1. WebSearch Pro [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ] is also a Firefox
extension that adds the ability to look up a text by
highlighting it on a page. Another feature is drag&amp;drop zones
to search for things directly from any website.
          </p>
          <p>To compensate for the deficiencies of SERPs we created a
browser extension called SearchPanel. This extension
combines document and process metadata in a visual
representation of search results to help people manage their
information seeking. We chose the browser extension approach
rather than creating a proxy for several reasons. While both
o↵ er the potential of parsing and augmenting SERP and
document pages, a browser extension has some advantages.</p>
          <p>It scales better with respect to storing user history data. It
ensures a higher level of data privacy, since data that might
potentially reveal user interests (e.g., query keywords,
selected URLs, etc.) can be logged as hashed values. Finally,
it has access to bookmarks and local browsing history.
3.1</p>
          <p>Design space</p>
          <p>When performing search tasks, searchers may need di↵
erent kinds of information to support their information
seeking. We represent the design space as consisting of three
categories of activities: search activity, navigation activity,
and organization activity.</p>
          <p>Historically, web UI support for the search process, or
search activity, has been focused on query formulation and
understanding the current query. Web browsers o↵ er
limited support for comparing current results set with earlier
activity by marking the visited status of documents.</p>
          <p>When engaged with a search task, users need to shift their
attention between the SERP and the retrieved pages. In
some cases, the searcher does not find the desired
information in a retrieved document, but rather in links to other
documents containing relevant information. This
navigation activity can be an important part of the information
seeking process.</p>
          <p>When searchers find useful web pages, they may wish to
save those documents for future access. More specialized
search engines sometimes support this capability directly,
but it is most often supported only by the browser’s
bookmarking capability.</p>
          <p>We can consider these search and sense-making activities
in light of the kinds of information required to satisfy them.</p>
          <p>In particular, Table 1 shows when document and process
metadata might be pertinent for the di↵ erent categories of
search activities. A representation of the number of visits
to a retrieved result (process metadata) could be used by a
searcher to decide how to interact with that result. In a
refinding sub task, for example, searchers might want to ignore
newly-found documents or pages that were not opened.</p>
          <p>The purpose of the search panel is to complement the
SERP and to be available when exploring search results; we
wanted the design to be simple and unobtrusive but still
convey useful information. Some features (e.g.,
organizating bookmarks) listed in Table 1 are too complex to be
integrated into the extension. Others, such as favicons, while
seemingly trivial, may still provide useful information for
navigating search results.</p>
          <p>SearchPanel displays automatically on the right side of the
browser window when it is enabled (Figure 1). The right side
of the content page has been chosen because this location is
frequently free of document content. In cases of overlap, its
vertical position can be adjusted manually to accommodate
page content that may be occluded.</p>
          <p>SearchPanel displays immediately after a search has been
performed on a supported web search engine (currently, they
are Google, Google Scholar, Yahoo, Bing and Microsoft
Academic Search). SearchPanel remains visible even if the
searcher follows links from retrieved documents. In addition,
searchers can return directly to the original query, or re-run
it on a di↵ erent search engine.</p>
          <p>A short tutorial page is displayed at installation, and can
also be reached through the option menu. This page also
allows logging (see 3.2.4) to be disabled, and can be used to
delete the recorded history.
1The Wayback Machine is a service that provides access to
archived and historical versions of web sites.</p>
          <p>SearchPanel displays several kinds of document metadata.</p>
          <p>Documents are represented by bars arranged in order
correa newly-found page; 3 favicon representing the site from
which the page was retrieved; 4 bar representing page that
has been visited; 5 highlighted bar based on cursor
position; 6 bookmark indicator; 7 currently-selected page.
sponding to the retrieved list; clicking on a bar is equivalent
to clicking on a link on the SERP. Almost all websites have
icons (favicons) to help re-identify the web page quickly;
these icons are shown to the right of the bar (see Figure 1,
item 3 ). A tooltip with the title of the document is added
to each bar as well. We considered identifying other
metadata such as document MIME type, but that would incur
the overhead of a separate HTTP request for each document.</p>
          <p>At least initially, we chose not to pursue this strategy.
3.2.2</p>
          <p>Process metadata</p>
          <p>Process metadata is also incorporated into SearchPanel.</p>
          <p>First, the icon of the search engine that ran the search is
highlighted in the top bar (item 1 ). Other icons
represent available comparable search engines. Clicking on one
of these icons re-runs the query with the selected search
engine. Search engines are grouped into two categories (web
search and academic research) and only the relevant ones are
shown. The current selection (highlighted with a black
border) links back to the search result page if the user navigates
to one of the retrieved documents.</p>
          <p>Each bar can have one of three di↵ erent colors, depending
on the link history. If a link has never been retrieved before,
the state of the link is ”new” and the color will be teal.
Results that have been retrieved by prior queries but have not
been clicked on are colored blue. Visited links are colored
violet. The local browser history is examined to retrieve the
link status. This allows us to incorporate page views that
occurred before SearchPanel was installed.</p>
          <p>Each bar’s length reflects the frequency of retrieval of the
corresponding page. The more frequently a page has been
retrieved, the shorter the bar gets (item 3 ). The retrieval
history is stored locally in the browser for privacy reasons
and can be deleted through SearchPanel’s option page.</p>
          <p>In SearchPanel, the bookmarking function serves two
purposes (item 6 in Figure 1). First, searchers can click on
the star to bookmark the corresponding page. Second,
previously bookmarked documents in the SERP will show a
yellow star next to them. This allows to re-find a web page
quicker, as the user does not need to navigate to a document
to know if they have previously bookmarked it.</p>
          <p>The selection indicator (see item 7 in Figure 1) indicates
the currently-selected result page. If a link on a result page
is clicked, the page indicator will stay on the last retrieved
document page to indicate that navigation started with it.</p>
          <p>Hovering over the result highlights the associated bar (item
5 ), and also highlights the corresponding snippet in the
SERP (Figure 2); the SERP is scrolled as necessary to bring
highlighted snippet into view. Conversely, when the mouse
is over a snippet on the SERP, the related bar jiggles
leftright to reinforce the connection between the two.</p>
          <p>When the user navigates o↵ the SERP to a search
result, SearchPanel remains active. Clicking on bars navigates
among the retrieved documents, bypassing the intermediate
step of reloading the search results. When the mouse is over
a bar in SearchPanel, the SERP snippet of that result will
be shown. This can be seen in Figure 3, where a preview
of the Wolfram Alpha snippet is shown. If the snippet is
not available, a tooltip with the document title is shown
instead. Both of these features should make it easier and
more e cient to navigate the search results without
necessarily creating a large number of tabs in the process.</p>
          <p>The extension was created to study people’s information
seeking behaviors. The goal of the project is to understand
how people use the web when looking for information to
improve their search experience. Therefore logging of user
activity was necessary. To encapsulate it from the basic
functionality it was designed as plugin that could be
connected or disconnected from SearchPanel. It collects
information related to the use of SearchPanel for the purposes of
statistical analysis of patterns of behavior.</p>
          <p>To maximize searchers’ privacy, no personally-identifying
information is saved. Queries and found URLs are recorded
as MD5-hashed values only. This allows us to identify
recurring queries and documents, without being able to read
the content of the query or to observe which pages people
view. Specifically, the following information is recorded:
• The IP address and the time the event was logged
• When a search result was clicked and where this
hap</p>
          <p>pened (SearchPanel or SERP)
• Hash strings that represent the queries and found web</p>
          <p>pages.
• Time spent with the mouse on di↵ erent interface parts</p>
          <p>(SearchPanel vs SERP)
• Various actions related to the extension (adding
book</p>
          <p>marks by clicking the start, moving it, etc.).</p>
          <p>After an in-house pilot deployment, SearchPanel has been
made available through the Google Chrome store. The goal
of the deployment is to understand whether the extension
helps people with their search tasks, and to assess the
relative utility of document vs. process metadata. We also
expect to collect a dataset that characterizes people’s browsing
and searching behaviors in terms of patterns of retrieval and
re-retrieval, search result navigation, etc.</p>
          <p>Web search engines are used for many di↵ erent kinds of
search tasks. While navigational and transactional uses of
search engines are well-supported by current interfaces and
algorithms, searchers are left to their own devices for more
open-ended information seeking and re-finding. We created
a Google Chrome browser extension to help people manage
their search activity. We explored the design space of
document and process metadata related to the wide range of
activities searchers may engage in during information
seeking. The extension keeps track of retrieval, page visits, and
bookmarking, and integrates traces of these activities with
document metadata to give people a more complete
impression of their search activity. An upcoming deployment will
explore the e↵ ect that this extension has on how people
interact with search results.
M. Atif Qureshi*†!, Arjumand Younus*†!, Colm O’Riordan*, Gabriella Pasi!, Nasir</p>
          <p>Touheed†
*Computational Intelligence Research Group, Information Technology, National University of Ireland,</p>
          <p>Galway, Ireland
!Information Retrieval Lab, Informatics, Systems and Communication, University of Milan Bicocca,</p>
          <p>Milan, Italy
†Web Science Research Group, Faculty of Computer Science, Institute of Business Administration,</p>
          <p>Karachi, Pakistan
muhammad.qureshi,arjumand.younus@nuigalway.ie, colm.oriordan@nuigalway.ie,
pasi@disco.unimib.it, ntouheed@iba.edu.pk
Traditional search engines fail to capture the notion of
“perspective” in their search results and at times present the
results skewed towards a particular topic. Under most of these
cases even query reformulation fails to retrieve desired search
results and the underlying reason for such failure is often
the bias within the document collection itself (e.g., news
articles). A perspective-aware search interface enabling users
to look into search results for some “perspective” terms may
be of great use for certain information needs. In this paper
we describe such a system.</p>
          <p>Categories and Subject Descriptors
H.1.2 [User/Machine Systems]: Human factors; H.3.3
[Information Search and Retrieval]: Search process
General Terms
Human Factors, Performance
Keywords
Perspective, Wikipedia, Bias
1. INTRODUCTION AND RELATED WORK</p>
          <p>
            It is often the case that when using a search engine for
information seeking users have an underlying intent [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ].
Traditional search interfaces fail to capture the user intent for
certain topics and at times return results that may be skewed
towards a certain perspective. Here, perspective as defined
by the Oxford Dictionary refers to a “point of view”1 within
the search results that may or may not be something what
user is looking for. We explain further through the following
motivating examples:
• Consider the case of a user who wishes to find more
about a certain event (say, a bomb attack in a certain
region). The search results returned contain a
majority of news reports blaming Islam relating it with
1This may also be seen as topic drifts within a document.
          </p>
          <p>Presented at EuroHCIR2013. Copyright !c 2013 for the individual
papers by the papers’ authors. Copying permitted only for
private and academic purposes. This volume is published and
copyrighted by its editors..
terrorism in most of the cases. This prompts the user
to explicitly evaluate how much Islam is related to
terrorism in the returned search results.
• Consider the case of a user who wishes to find out
about roles and rights of women in Islam but the search
engine returns articles that contain a high amount of
terms highlighting oppression against women instead
of women rights and roles. In this case the user is
prompted to check the correlation between women and
oppression within the search results that have been
returned.</p>
          <p>Note that the perspective given by most search results
(Islam in our motivating example (1) and oppression in our
motivating example (2)) may or may not be aligned with
the user’s query intent. In case of search results not being
aligned with his/her query intent he/she may be interested
in observing the amount of perspective tendencies in various
news reports.</p>
          <p>
            This paper proposes the concept of a “perspective-aware”
search interface that enables the user to explicitly analyse
search results for information from a particular
perspective with respect to an issued query. To the best of our
knowledge, previous research within Human-Computer
Interaction and Information Retrieval has failed to capture
the notion of “perspective” within the information retrieval
process. Early research related to Interactive Information
Retrieval by Belkin [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] and Ingwersen [
            <xref ref-type="bibr" rid="ref22 ref6">6</xref>
            ] suggests the
integration of cognitive aspects within the information retrieval
process: in line with this suggestion we argue for
incorporating the essential cognitive element of “perspectives”2 within
the search engine interface.
          </p>
          <p>
            Recently the information retrieval community has turned
attention to diversification of search results which aims to
tackle the issue of query ambiguity on the user side [
            <xref ref-type="bibr" rid="ref24 ref8">8</xref>
            ].
However, even when formulating a non-ambiguous query users
may have an intent that influences the perspective from
which the query terms can be interpreted in a text; in case of
2According to Wikipedia the definition of perspective states
the following: “Perspective in theory of cognition is the
choice of a context or a reference (or the result of this choice)
from which to sense, categorize, measure or codify
experience, cohesively forming a coherent belief, typically for
comparing with another.”
perspective mismatch between the user intent and the
documents returned in first positions by a search engine, users
may find the retrieved results annoying or subjective to a
non-agreed perspective [
            <xref ref-type="bibr" rid="ref23 ref7">7</xref>
            ]. One may argue that a query
reformulation technique could be employed to tackle this
problem [
            <xref ref-type="bibr" rid="ref21 ref5">5</xref>
            ]; e.g. considering the motivating example (2), the user
could issue a reformulated query such as “roles and rights of
women in islam”. However, for some topics query
reformulation may fail to retrieve the desired search results, and the
underlying reason for such failure is often the bias within the
document collection itself (e.g., news articles) [
            <xref ref-type="bibr" rid="ref10 ref26">10</xref>
            ]. Under
such a scenario it would be interesting to provide a search
interface that would enable the users to look into the search
results for some “perspective” terms and we describe such a
system in this paper.
          </p>
          <p>
            This section presents the essential details of the proposed
perspective-aware search interface along with the underlying
implementation details. We keep the interface as simple as
possible on account of research suggesting users’ reluctance
in switching from a simple search form [
            <xref ref-type="bibr" rid="ref19 ref3">3</xref>
            ]. Figure 1 shows
the entry point of the interface which resembles the standard
type-keywords-in-entry-form interface with the
augmentation of an additional input text box for entry of perspective
terms.
          </p>
          <p>The underlying perspective detection algorithm makes use
of the encyclopedic structure in Wikipedia; more
specifically the knowledge encoded in Wikipedia’s graph structure
is utilized for the discovery of various perspectives in
documents returned by the search engine. Wikipedia is organized
into categories in a taxonomy-like3 structure (see Figure 2).</p>
          <p>
            Each Wikipedia category can have an arbitrary number of
subcategories as well as being mentioned inside an arbitrary
number of supercategories (e.g., category C4 in Figure 1 is
a subcategory of C2 and C3, and a supercategory of C5, C6
and C7.) Furthermore, in Wikipedia each article can belong
to an arbitrary number of categories, where each category is
a kind of semantic tag for that article [
            <xref ref-type="bibr" rid="ref11 ref27">11</xref>
            ]. As an example,
in Figure 2, article A1 belongs to categories C1 and C10,
article A2 belongs to categories C3 and C4, while article A3
belongs to categories C4 and C7. It can be seen that the
articles and the Wikipedia Category Graph are interlinked
and our system makes use of these interlinks for the
detection of a certain perspective within a document retrieved by
the search engine.
          </p>
          <p>The underlying perspective detection algorithm within our
system requires the perspective term/phrase to match the
title of a Wikipedia article. This may seem to impose a
cognitive load on the user at search time. However, this is not
the case: as shown in Figure 3 the entered text
automatically turns green when a certain user-specified perspective
term matches the title of a Wikipedia article, and
symmetrically the entered text automatically turns red in case of a
mismatch.</p>
          <p>Once the perspective term is entered correctly the system
fetches the Wikipedia article corresponding to the
perspective term referred to as Seed Perspective Article (PAseed)
along with the categories to which it belongs and we use
3We say taxonomy-like because it is not strictly
hierarchical due to the presence of cycles in the Wikipedia category
graph.</p>
          <p>DISCUSSION
“terrorism” is shown in Figure 4. As evident from the top
search result, there is a high perspective of terrorism within
the returned document and perspective terms that our
algorithm fetches are as follows: a) the war on terrorism, b)
ayman al zawahiri, and c) osama bin laden.</p>
          <p>PC04 to refer to these categories. After fetching of Wikipedia
categories in PC0, the system retrieves sub-categories of PC0
until depth 2 i.e., PC1 and PC25 and collectively these
categories related to PAseed are referred to as PC (where PC
is union of PC0, PC1 and PC2.). Next, the set of all
articles within the Wikipedia category set PC is retrieved
and we refer to this set as Expanded Perspective Article Set
(PAexpanded). The system then retrieves all categories
associated with the set PAexpanded which we refer to as WC ;
note that PC is a subset of WC. Finally, the intersection
between PC and WC is retrieved which is a set of categories
representative of the domain of the perspective term
originally input by the user, we refer to this set of representative
categories as RC.</p>
          <p>After building the Wikipedia category sets as defined above6
i.e., PC, RC and WC we match variable-length n-grams
within a document with articles in the set PAexpanded, and
we check for cardinality of RC and WC. The cardinality
scores along with n-gram frequencies are used to compute a
perspective score for each document.</p>
          <p>There have been many efforts in the information retrieval
research to present to users information regarding the
relationship between the query and the answer set and the query
and document collection. Capturing this information during
the retrieval process provides the user with much valuable
information (e.g. whether a term is overly specific, or whether
a term is ambiguous etc.). Various attempts have been made
to tackle this problem, ranging from the definition of
snippets to the definition of approaches to cluster search results
(Clusty.com), to the presentation of diversified search results
in the first position of the ranked list offered to the users.</p>
          <p>Recently there has been a resurgence of interest in defining
visualization techniques of search results that offer an
effec2.2 Search Results Presentation tive and more informative alternative to usual and scarcely
informative ranked lists. Pioneer visualization systems are</p>
          <p>
            The perspective scores computed in section 2.1 are dis- represented by Tilebar [
            <xref ref-type="bibr" rid="ref20 ref4">4</xref>
            ], and Infocyrstal [
            <xref ref-type="bibr" rid="ref25 ref9">9</xref>
            ], and these
played within the search results, and based on the perspec- attempts have been aimed to provide the user with more
tive score a document receives , we define four levels of information than that provided by the traditional ranked
perspective adherence as follows: a) High, b) Medium, c) list.
          </p>
          <p>Low, and d) Neutral. Moreover, in case of documents with This additional information can help the user in their
high, medium and low scores we also report the top-scoring search task (e.g. allowing them to navigate the collection
perspective terms that were extracted using the Wikipedia more easily or providing evidence to allow the user to
reforgraph structure as explained previously. A sample search mulate their query more efficiently).
corresponding to search query “india pakistan relations” and Our proposed system, although related in that we also
attempt to give the user an insight into the answer set and its
4These are basically perspective categories at depth zero. relation to the query, differs in a fundamental manner. Our
t5wToh.ese are basically perspective categories at depth one and system, we posit, allows the user to gain insight into the
an6The set building phase is performed through a cus- swer set and its relation to the query, but moreover, allows
tom Wikipedia API that has pre-indexed Wikipedia to the user to gain an insight into a perspective inherent in
data and hence, it is computationally fast. For details the answer set. Our system uses an external and collectively
http://www3.it.nuigalway.ie/cirg/prj/WikiMadeEasy.html created knowledge resource (which is less likely to be biased
in a given direction) to obtain extra terms to represent the
perspective of interest to the user. This knowledge
(perspective term and related terms) does not modify the query
(as would an additional query term), but is instead used to
highlight the presence of a perspective in the answer set.</p>
          <p>In this paper we have proposed a novel approach for
capturing the relationship between a user’s query and the
returned answer set. We do not rely on evidence in the
document collection or the query stream, but rather instead
extract terms from an external source of evidence to help
users quickly see the presence of a particular perspective in
the document collection and answer set.</p>
          <p>Having built the system and undertaken preliminary user
evaluations7, we aim at undertaking a complete and
systematic review of the approach. This will comprise a number
of separate user evaluation tasks. The initial experiments
will involve comparing our search approach with and
without the perspective-aware component over a number of tasks
to see if the additional context and information provided by
our perspective aware system aids the users in a range of
information-seeking tasks. Our second planned experiments
will be focussed on persons seeking information from
newspaper articles, a domain wherein a degree of bias often exists.</p>
          <p>We wish to explore the users’ experience with regards to any
perceived bias in the considered corpora.</p>
          <p>REFERENCES</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>ABAKUS.</given-names>
            <surname>Bettersearch</surname>
          </string-name>
          <article-title>a firefox addon for enhancing search engines</article-title>
          . http://mybettersearch.com/,
          <year>2010</year>
          . [Online; accessed 06/06/2013].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ribeiro-Neto</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , et al.
          <source>Modern information retrieval</source>
          , vol.
          <volume>463</volume>
          . ACM press New York,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Broder</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>A taxonomy of web search</article-title>
          .
          <source>SIGIR Forum 36</source>
          ,
          <issue>2</issue>
          (Sept.
          <year>2002</year>
          ),
          <fpage>3</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chau</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zeng</surname>
          </string-name>
          , D. Metaspider:
          <article-title>Meta-searching and categorization on the web</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          <volume>52</volume>
          ,
          <issue>13</issue>
          (
          <year>2001</year>
          ),
          <fpage>1134</fpage>
          -
          <lpage>1147</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Elsweiler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ruthven</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Towards task-based personal information management evaluations</article-title>
          .
          <source>In Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          (New York, NY, USA,
          <year>2007</year>
          ),
          <source>SIGIR '07</source>
          , ACM, pp.
          <fpage>23</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Even</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shankaranarayanan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Watts</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Enhancing decision making with process metadata: Theoretical framework, research tool, and exploratory examination</article-title>
          .
          <source>In System Sciences</source>
          ,
          <year>2006</year>
          .
          <source>HICSS'06. Proceedings of the 39th Annual Hawaii International Conference on (2006)</source>
          , vol.
          <volume>8</volume>
          , IEEE, pp.
          <fpage>209a</fpage>
          -
          <lpage>209a</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Golovchinsky</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diriye</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Dunnigan</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>The future is in the past: designing for exploratory search</article-title>
          .
          <source>In Proceedings of the 4th Information Interaction in Context Symposium</source>
          (New York, NY, USA,
          <year>2012</year>
          ), IIIX '12, ACM, pp.
          <fpage>52</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Hong-li</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <article-title>A novel visual search engines: Grokker</article-title>
          .
          <source>Journal of Library and Information Sciences in Agriculture 8</source>
          (
          <year>2008</year>
          ),
          <fpage>047</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>KG</surname>
          </string-name>
          , P. U. . C.
          <article-title>Searchpreview, the browser extension previously known as googlepreview</article-title>
          . http://searchpreview.de/,
          <year>2013</year>
          . [Online; accessed 06/06/2013].
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Martijn</surname>
          </string-name>
          .
          <article-title>Web seach pro, search the web the way you like</article-title>
          ... http://websearchpro.captaincaveman.nl,
          <year>2012</year>
          . [Online; accessed 06/06/2013].
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Qvarfordt</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Golovchinsky</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunnigan</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Agapie</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <article-title>Looking ahead: Query preview in exploratory search</article-title>
          .
          <source>In Proceedings of the 36th international ACM SIGIR conference on Research and development in Information Retrieval</source>
          (New York, NY, USA,
          <year>2013</year>
          ),
          <source>SIGIR '13</source>
          , ACM.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boukhelifa</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Rodgers</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>Multiform glyph based web search result visualization</article-title>
          .
          <source>In Information Visualisation</source>
          ,
          <year>2002</year>
          . Proceedings. Sixth International Conference on (
          <year>2002</year>
          ), IEEE, pp.
          <fpage>549</fpage>
          -
          <lpage>554</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Rose</surname>
            ,
            <given-names>D. E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Levinson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Understanding user goals in web search</article-title>
          .
          <source>In Proceedings of the 13th international conference on World Wide Web</source>
          (
          <year>2004</year>
          ), ACM, pp.
          <fpage>13</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Spoerri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>How visual query tools can support users searching the internet</article-title>
          .
          <source>In Information Visualisation</source>
          ,
          <year>2004</year>
          .
          <article-title>IV 2004</article-title>
          .
          <article-title>Proceedings</article-title>
          . Eighth International Conference on (
          <year>2004</year>
          ), IEEE, pp.
          <fpage>329</fpage>
          -
          <lpage>334</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Tauscher</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Greenberg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>How people revisit web pages: empirical findings and implications for the design of history systems</article-title>
          .
          <source>Int. J. Hum.-Comput. Stud</source>
          .
          <volume>47</volume>
          ,
          <issue>1</issue>
          (
          <year>July 1997</year>
          ),
          <fpage>97</fpage>
          -
          <lpage>137</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Teevan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>The re:search engine: simultaneous support for finding and re-finding</article-title>
          .
          <source>In Proceedings of the 20th annual ACM symposium on User interface software and technology</source>
          (New York, NY, USA,
          <year>2007</year>
          ),
          <source>UIST '07</source>
          , ACM, pp.
          <fpage>23</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Teevan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adar</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Potts</surname>
            ,
            <given-names>M. A. S.</given-names>
          </string-name>
          <article-title>Information re-retrieval: repeat queries in yahoo's logs</article-title>
          .
          <source>In Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          (New York, NY, USA,
          <year>2007</year>
          ),
          <source>SIGIR '07</source>
          , ACM, pp.
          <fpage>151</fpage>
          -
          <lpage>158</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deshpande</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Shneiderman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>A temporal pattern search algorithm for personal history event visualization. Knowledge and Data Engineering</article-title>
          , IEEE Transactions on
          <volume>24</volume>
          ,
          <issue>5</issue>
          (
          <year>2012</year>
          ),
          <fpage>799</fpage>
          -
          <lpage>812</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          .
          <article-title>'natural' search user interfaces</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>54</volume>
          (
          <issue>11</issue>
          ):
          <fpage>60</fpage>
          -
          <lpage>67</lpage>
          , Nov.
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. O.</given-names>
            <surname>Pedersen</surname>
          </string-name>
          .
          <article-title>Visualizing information retrieval results: a demonstration of the tilebar interface</article-title>
          .
          <source>In Conference Companion on Human Factors in Computing Systems</source>
          , pages
          <fpage>394</fpage>
          -
          <lpage>395</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. N.</given-names>
            <surname>Efthimiadis</surname>
          </string-name>
          .
          <article-title>Analyzing and evaluating query reformulation strategies in web search logs</article-title>
          .
          <source>In Proceedings of the 18th ACM conference on Information and knowledge management</source>
          ,
          <source>CIKM '09</source>
          , pages
          <fpage>77</fpage>
          -
          <lpage>86</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ingwersen</surname>
          </string-name>
          .
          <article-title>Cognitive perspectives of information retrieval interaction: Elements of a cognitive IR theory</article-title>
          .
          <source>Journal of Documentation</source>
          ,
          <volume>52</volume>
          (
          <issue>1</issue>
          ):
          <fpage>3</fpage>
          -
          <lpage>50</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Jansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Booth</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Spink</surname>
          </string-name>
          .
          <article-title>Determining the informational, navigational, and transactional intent of web queries</article-title>
          .
          <source>Inf</source>
          . Process. Manage.,
          <volume>44</volume>
          (
          <issue>3</issue>
          ):
          <fpage>1251</fpage>
          -
          <lpage>1266</lpage>
          , May
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Ounis</surname>
          </string-name>
          .
          <article-title>Intent-aware search result diversification</article-title>
          .
          <source>In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval, SIGIR '11</source>
          , pages
          <fpage>595</fpage>
          -
          <lpage>604</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Spoerri</surname>
          </string-name>
          .
          <article-title>Infocrystal: A visual tool for information retrieval &amp; management</article-title>
          .
          <source>In Proceedings of the second international conference on Information and knowledge management</source>
          , pages
          <fpage>11</fpage>
          -
          <lpage>20</lpage>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Younus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Qureshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Kingrani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saeed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Touheed</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. O'Riordan</surname>
            ,
            <given-names>and P.</given-names>
          </string-name>
          <string-name>
            <surname>Gabriella</surname>
          </string-name>
          .
          <article-title>Investigating bias in traditional media through social media</article-title>
          .
          <source>In Proceedings of the 21st international conference companion on World Wide Web, WWW '12 Companion</source>
          , pages
          <fpage>643</fpage>
          -
          <lpage>644</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zesch</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Gurevych.</surname>
          </string-name>
          <article-title>Analysis of the Wikipedia Category Graph for NLP Applications</article-title>
          .
          <source>In Proceedings of the TextGraphs-2 Workshop (NAACL-HLT)</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>