<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>2 !European)Workshop)on)! Human"Computer)Interaction)and)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mixed,Reality,Lab,</string-name>
          <email>max.wilson@nottingham.ac.uk, , Tony"Russell&gt;Rose"</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>The,Royal,School,of,Library,and,,</string-name>
          <email>blar@iva.dk, , James"Kalbach"</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information</institution>
          ,
          <addr-line>Science,</addr-line>
          ,
          <country country="DK">Denmark</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University,of</institution>
          ,
          <addr-line>Nottingham,</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <fpage>45</fpage>
      <lpage>86</lpage>
      <abstract>
        <p>EuroHCIR2012,was,the,second,workshop,in,the,Europea n,series,focusing,on,the, combined,aspects,of,Human=Computer,Interaction,and,Information,Retrieval, (HCIR).,The,MUMIA,WG3,supported,event,saw,significant,growth,in,interest, from,the,first,year,,attracting,over,30,submissions.,9,key,research,and,position, papers,were,accepted,for,Oral,Presentation,,while,a,further,13,demos,and, posters,were,accepted.,All,are,included,in,these,proceedings,so,that,they,can,be, accessed,by,those,unable,to,attend,the,event,in,Nijmegen. ,</p>
      </abstract>
      <kwd-group>
        <kwd>Information*Retrieval!</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Proceedings+of+the+!
nd</p>
    </sec>
    <sec id="sec-2">
      <title>Executive!Summary!</title>
    </sec>
    <sec id="sec-3">
      <title>Organised!by!</title>
      <sec id="sec-3-1">
        <title>Max"L."Wilson"</title>
      </sec>
      <sec id="sec-3-2">
        <title>Birger"Larsen"</title>
        <p>Supported!as!a!MUMIA!WG3!event!</p>
        <sec id="sec-3-2-1">
          <title>Oral!Presentations!</title>
          <p>Page"3"&gt;"" Using"Card"Sorts"to"Understand"how"Users"Think"of"Personal"Information"</p>
          <p>Paul%Thomas%and%David%Elsweiler"
Page"7"&gt;"" Using"Semantic"Differentials"for"anE"valuative"View"of"theS"earch"Engine"as"an"
Interactive"System"</p>
          <p>Frances%Johnson"
Page"11"&gt;"" The"Fault,"Dear"Researchers,"is"not"in"Cranfiedl,"But"in"ourM"etrics,"that"they"
are"Unrealistic."</p>
          <p>Mark%D.%Smucker%and%Charles%L.%A.%Clarke%
Page"13"&gt;"" A"Model"of"Consumer"Search"Behaviou"r</p>
          <p>Tony%Russell@Rose%and%Stephann%Makri%
Page"17"&gt;"" Revisiting"User"Information"Needs"in"Aggregated"Sear"ch</p>
          <p>Shanu%Sushmita,%Martin%Halvey,%Robert%Villa%and%Mounia%Lalmas%
Page"21"&gt;"" Improving"Search"Experience"on"Distributed"Leisure"Events"</p>
          <p>Richard%Schaller,%Morgan%Harvey%and%David%Elsweiler%
Page"25"&gt;"" Opinion"Mapping:"Information"Visualization"Approaches"for"Comparative"
Sentiment"Analysis"</p>
          <p>William%Hsu%and%Praveen%Koduru%
Page"29"&gt;"" Search"System"Functions"for"Supporting"Search"Mod"es</p>
          <p>Thomas%Beckers%and%Norbert%Fuhr%
Page"33"&gt;"" Ingredients"for"a"User"Interface"to"Support"Media"Studies"Researchers"in"Data"
Collection"</p>
          <p>Marc%Bron,%Frank%Nack,%Maarten%De%Rijke%and%Jasmijn%Van%Gorp"</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Poster!Papers!</title>
          <p>Page"37"&gt;"" Exploring"Italian"Wine:""aCase"Study"of"Aesthetics"and"Interaction"in"a"</p>
          <p>Generative"Information"Visualization"Method"</p>
          <p>Luca%Buriano""
Page"41"&gt;"" From"Task&gt;based"Evaluation"to"Featur&gt;ebased"Evaluation"in"Personal"Searc"h</p>
          <p>Seyedeh%Sargol%Sadeghi,%Mark%Sanderson%and%Falk%Scholer""
Page"43"&gt;"" Visualization"of"Clandestine"Labs"from"Seizure"Reports:"Thematic"Mapping"and"
Data"Mining"Research"Directions"</p>
          <p>William%Hsu,%Mohammed%Abduljabbar,%Ryuichi%Osuga,%Max%Lu%and%Wesam%Elshamy""
Page"47"&gt;"" Towards"Detecting"Wikipedia"Task"Context"s</p>
          <p>Hanna%Knäusl,%David%Elsweiler%and%Bernd%Ludwig""
Page"51"&gt;" CUES:"Cognitive"Usability"Evaluation"Syste m"</p>
          <p>Matthew%Pike,%Max%L.%Wilson,%Anna%Divoli%and%Alyona%Medelyan""
Page"55"&gt;"" Collaborative"Environment"of"the"PROMISEIn"frastructure:"an""ELEGantt""</p>
          <p>Approach"
Marco%Angelini,%Claudio%Bartolini,%Gregorio%Convertino,%Guido%Granato,%Preben%Hansen%
and%Giuseppe%Santucci"
Page"59"&gt;"" Search"User"Interface"Design"for"Children:"Challenges"andl"uStoions"</p>
          <p>Tatiana%Gossen,%Marcus%Nitsche%and%Andreas%Nuernberger""
Page"63"&gt;" EyeGrab:"A"Gaze&gt;based"Game"with"a"Purpose"to"Enrich"Image"Context"</p>
          <p>Information"</p>
          <p>Tina%Walber,%Chantal%Neuhaus%and%Ansgar%Scherp""
Page"67"&gt;"" Using"Wordclouds"to"Navigate"and"Summarize"Titwter"Search"Results"</p>
          <p>Rianne%Kaptein""
Page"71"&gt;"" Do"Users"Benefit"from"Controlled"Vocabularies"in"Search"Interfaces?"</p>
          <p>Ying@Hsang%Liu,%Paul%Thomas,%Jan@Felix%Schmakeit%and%Tom%Gedeon""
Page"75"&gt;"" User&gt;Centred"Design"to"Support"Exploration"and"Path"Creoanti"in"Cultural"
Heritage"Collections"
Paula%Goodale,%Paul%Clough,%Nigel%Ford,%Mark%Hall,%Mark%Stevenson,%Samuel%Fernando,%</p>
          <p>Nikolaos%Aletras,%Kate%Fernie,%Phil%Archer%and%Andrea%De%Polo%
Page"79"&gt;"" Supporting"Serendipitous"and"Focused"Searc"h</p>
          <p>Junte%Zhang""
Page"83"&gt;"" Vague"Query"Formulation"by"Desig"n</p>
          <p>Marcus%Nitsche%and%Andreas%Nuernberger%</p>
          <p>Using card sorts to understand
how users think of personal information</p>
          <p>Paul Thomas</p>
          <p>CSIRO, Canberra
paul.thomas@csiro.au</p>
          <p>David Elsweiler
University of Regensburg
david@elsweiler.co.uk
ABSTRACT
Understanding how users think of personal information, and
how they mentally categorise or classify the objects they
work with, should inform the design of personal information
management (PIM) or personal retrieval systems. However,
most investigations of this topic predate widespread
multimedia, websites, and social media—objects that a contemporary
PIM or retrieval system should work with.</p>
          <p>We describe a pilot study that has used a variant of card
sorts to elicit categories for personal information such as
files, email, tweets, and websites. Our early results suggest
that there are common categorisations which are not yet
supported by PIM software, but which might reward further
work. Our results also suggest that—with some caveats—
card sorts are useful for understanding users’ categories.</p>
          <p>Categories and Subject Descriptors: H.5.2
[Information Interfaces and Presentation]: User Interfaces
General Terms: Human Factors
Keywords: Facets, classification, card sorts
1. INTRODUCTION</p>
          <p>
            Tools for personal information management (PIM) and
search support the archival, retrieval, and management of
“personal” data: the files, email, photos, videos, and other
digital objects a person creates or uses [
            <xref ref-type="bibr" rid="ref27">15</xref>
            ]. Several studies
show that PIM can be challenging [
            <xref ref-type="bibr" rid="ref11 ref17 ref23 ref27 ref5">5, 11, 15</xref>
            ] and it has been
suggested that tools could be easier to use and more useful, if
the way they represent objects matches the way users think
of them [
            <xref ref-type="bibr" rid="ref10 ref22">10, 20</xref>
            ].
          </p>
          <p>We are interested in how users think of the wide range of
digital objects they interact with—the objects conventionally
considered by PIM tools, objects less commonly considered
such as websites and applications, and newer objects such as
messages from social media. There are three linked questions:
1. What properties do users think “personal” digital
objects have? That is, in which ways do users think of
the objects they use?
Presented at EuroHCIR2012. Copyright © 2012 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
2. Can we expose these properties in a PIM or search tool?</p>
          <p>Can the properties of an object be determined
algorithmically? How should the properties be presented?
3. Assuming we can expose some or all of these properties,
would we expect that to make management or retrieval
easier?</p>
          <p>At present, we are considering the first question. In
particular, in this work, we have experimented with card sorts
to elicit users’ own descriptions of personal information.
2.</p>
          <p>CLASSIFICATION AND TOOLS</p>
          <p>Past work has investigated the properties users assign to
files, and elicited categorisation schemes. This work has not
however considered as wide a range of object types as we do
here; we may expect that with di↵ erent types, sources, and
quantities we would see di↵ erent categorisations. Existing
PIM and file management systems also support, or impose,
particular faceting schemes.
2.1</p>
          <p>Classifications</p>
          <p>There is a rich tradition, in information science and
information behaviour, of studies that try to understand how
people organise, classify, and think about their information—
that is, how people understand their information independent
of any particular software capabilities or restrictions.</p>
          <p>Three studies of note are by Cole, Kwasnik and Case.</p>
          <p>
            Cole [
            <xref ref-type="bibr" rid="ref19 ref7">7</xref>
            ] studied how 30 o ce workers classified their
document collections. Six aspects of documents were important in
filing decisions: “type”, “form”, “volume”, “complexity”,
“functions”, and “levels of information”. Similarly, Kwa´snik [19]
examined the categorisation behaviour of eight researchers
and identified seven dimensions: “situation”, “document”,
“disposition”, “order/scheme”, “time”, “value”, and “cognitive
state”. Case [
            <xref ref-type="bibr" rid="ref18 ref6">6</xref>
            ] investigated the behaviour of twenty
historians and identified three main factors by which objects were
classified in o ces; “ease of access”, “form” and “topic”. While
there is considerable overlap in the findings of these studies,
particularly the criteria “form” and “topic”, the studies all
predate the rich digital landscape we have today, and focus
on physical information objects.
          </p>
          <p>
            Other research relating to our work has tried to learn about
how people think about digital information by investigating
how they behave with information in practice. For example,
people have been shown to organise email messages and files
based on projects [
            <xref ref-type="bibr" rid="ref29">17, 23</xref>
            ] and prefer to refind objects by
location than using search facilities [
            <xref ref-type="bibr" rid="ref14 ref15 ref2 ref3">2, 3</xref>
            ]. These kinds of
studies provide strong hints at how people may think about
digital information, but are influenced by the tools they have
available to them.
          </p>
          <p>
            More recently, Gonc¸alves and Jorge [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ] asked participants
to tell stories about three of their personal documents by
describing each, from memory, in terms of its features, its
content and the context in which it was created or used.
          </p>
          <p>It was discovered that time, location, and purpose of the
document were the most common attributes used in stories.</p>
          <p>
            Similarly, Blanc-Brude and Scapin [
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ] used semi-structured
interviews to examine participants’ recollection of their
documents. They found that location, format, time, keywords
and associated events were remembered most frequently, but
many of these attributes, particularly keywords, time and
location were often only partially remembered or the
recollections o↵ ered by the participants were incorrect. Both of these
studies add a rich understanding of how people perceive their
documents by examining a small number of documents in
great detail, but do not explore how documents are related.
          </p>
          <p>Our aim here is to add to and complement this previous
work by using a technique that can deal with rich variety of
information objects we interact with today; be tool agnostic;
and allow insight into how di↵ erent documents can be
associated in di↵ erent ways. We would also like to understand
whether this has any impact on the design of PIM tools.
2.2</p>
          <p>Tools</p>
          <p>
            Tools for desktop search typically expose not just
filesystem attributes such as name, size, and timestamp, but also
extracted metadata. For example, Phlat [
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ] uses title, date,
author, recipient, media type and tags; Haystack [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ] has
extensible facets but the authors have discussed media type,
people named in email, text in a document, and URL.
          </p>
          <p>
            More elaborate PIM tools have exposed other attributes to
support di↵ erent interactions. Some, such as Lifestreams [
            <xref ref-type="bibr" rid="ref12 ref24">12</xref>
            ],
have supported time-based browsing and searching; an
interface to Stu↵ I’ve Seen [21] extended this by indexing
documents according to contemporaneous events. Other
tools have taken a more personal view of time, or document
lifecycle, and supported information management by context.
          </p>
          <p>
            Here, objects are organised according to tags for the
context in which they are used [
            <xref ref-type="bibr" rid="ref28 ref30">16, 18</xref>
            ] or grouped according to
patterns of use [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ].
          </p>
          <p>These systems o↵ er variety of projections, across a number
of media and storage types, but it is hard to know whether
these match the way people naturally think of their objects.</p>
          <p>Alternative presentations that were natural for users, and
easy to implement, would be worth further thought.</p>
          <p>METHOD</p>
          <p>In this work, we have experimented with repeated
singlecriterion card sorts to elicit users’ mental categorisations.</p>
          <p>Cards represented digital objects on each participant’s
computer.
3.1</p>
          <p>Card sorts</p>
          <p>Repeated single-criterion card sorts—or just “card sorts”—
are a common technique for eliciting users’ categorisations
(see e.g. Rugg and McGeorge [22] for an overview). Compared
with interview-based techniques, card sorts are less flexible
but are very lightweight: in our experience participants
grasped the idea very quickly, many found it enjoyable, and
the entire protocol took little time. Coding card sorts for
later analysis is also relatively straightforward.</p>
          <p>In a typical exercise, each participant is given a number of
cards, each representing an object or a concept. They are
asked to partition these cards according to any criteria they
like; the criteria used for the sort, the categories (piles or sets),
and the cards in each category are recorded. This is repeated
several times, with participants suggesting a di↵ erent criteria
each time. For example, given cards labelled as follows:</p>
          <p>1. pig; 2. chicken; 3. snake; 4. horse; 5. spider
a participant may sort cards according to the criterion “raised
on a farm”, with cards 1, 2, and 4 in category “yes” and
cards 3 and 5 in category “no”. A second sort, according
to the criterion “where eaten”, might have cards 1 and 2
in category “almost everywhere”, card 3 in category “Asia”,
card 4 in category “Asia and Europe” and card 5 in category
“don’t know”.</p>
          <p>Records of the sorts may then be analysed with qualitative
or quantitative methods.
3.2</p>
          <p>Our approach</p>
          <p>For this early experiment, we recruited a convenience
sample of ten participants from two institutions. All were heavy
computer users.</p>
          <p>As preparation for the experiment participants were asked
to select several information objects they had seen, used
or created in the recent past. An “information object” was
defined by giving as examples computer files, emails,
websites, tweets or Facebook updates, documents or articles
read, photographs or images, videos, music, and computer
applications. However, participants were not restricted to
these objects and could choose anything they wanted using
these as a guideline. We encouraged participants to label 10
to 15 cards, which we believe balances the need for broad
coverage with practical limits on participants’ time.</p>
          <p>While choosing objects, participants were asked to create
index cards with the name of each (or some other reminder
of its identity or contents).</p>
          <p>With cards made, each participant was introduced to sorts
using a set of cards with pictures of buildings; they were
taken through some example sorts which included criteria
clear from the pictures themselves (colour, material), criteria
which were not immediately clear (insulation), criteria which
were subjective (good place for a party). “Can’t tell” or
“don’t know” categories were included in these examples.</p>
          <p>The participants’ own cards were then used for repeated
sorts. Participants were asked to make piles according to
a criterion of their choosing, and we noted the criterion
(sometimes this was implicit), the categories used, and the
cards in each category.</p>
          <p>
            After collecting individual classifications from all
participants, the full dataset, i.e. the criteria used to associate
information objects, was analysed qualitatively using an
a nity diagramming technique. This is a group-based
process, which allows the discovery and validation of patterns in
the data [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ]. The researchers, as a team, looked for patterns
in the data and grouped related criteria; we then related the
formed groups in a way that creates a hierarchical coding
scheme.
          </p>
          <p>RESULTS OF THE SORTS</p>
          <p>The results from this pilot are promising. Card sorts
elicited a variety of criteria; some of these di↵ er from those
seen before, and many are not well supported by PIM or
retrieval tools.
document lifecycle
events
object’s form
object’s a↵ ective qualities
object’s cost and value
object static/dynamic distinction
people and community
properties of associated tasks
topics covered
work/leisure distinction
(three other object-related groups)</p>
          <p>Participants</p>
          <p>General observations</p>
          <p>Our ten participants provided 64 sorts, a median 5.5 sorts
each (first/third quartile 5.0/8.0 sorts each). At the leaves
of our hierarchy, there was in general little overlap: 13 of 25
criteria were used by only one participant. However, 12 were
used by two or more participants, 8 by three or more, and one
criteria (discussed below) was used by every participant for at
least one sort. Figure 1 summarises how many participants
were represented in each top-level group, that is each group
at the top of our hierarchical coding.</p>
          <p>The single most common criterion was a distinction
between objects used for work and objects used for leisure—all
our participants used this criteria, and typically early on.</p>
          <p>Following this there were four common groups: to do with
the object itself, especially the form (data type and other
surface features), which 7 participants used at least once;
the topics an object is connected with (5 participants); the
properties of tasks associated with an object (4); and criteria
describing people and community (4). A striking finding
is the diversity in the criteria derived by the participants.</p>
          <p>Although we were able to group the criteria into 13 distinct
high-level cateorgies, only two of these, work/leisure and
form, were named by more than half the population.</p>
          <p>After our initial grouping, which was based on labels
proffered by participants and not on any statistics of the sorts
themselves, there were no clear correlations between criteria—
that is, “work” did not look the same as “important” or “Word
files” as “makes me angry”.</p>
          <p>On our analysis, five of Kwa´snik’s seven groups were
represented in our data: situation, document, time, value, and
cognitive state. However, they were very unevenly distributed:
criteria we classified as “situation” were used by all ten
participants, for one or two sorts each; “document” was used by
nine participants, for a median 3.5 sorts each; while at the
other end of the scale, “time” was used by only three
participants and “cognitive state” was used by two participants,
once each in each case.
4.2</p>
          <p>Criteria, groups, and tool support</p>
          <p>In many previous studies two groups of criteria—form and
topic—were found to be central, and our data reinforces this.</p>
          <p>There are also, however, notable contrasts.</p>
          <p>Work/leisure. To the best of our knowledge, previous
studies of classification behaviour have not found a work/leisure
distinction. However, every participant in our sample used
this criteria. This may be because we most of the objects in
our study were digital, not paper documents—it is very easy
to mingle work- and leisure-related objects online—but it
is clearly important and is not explicitly supported by PIM
tools.</p>
          <p>One participant reported that he used two top-level folders
in his file system, and two email accounts, to keep work and
leisure information separate. No other participants reported
as clear a distinction, however. It should be possible in a
PIM/search tool to tag files, or e.g. learn a classifier, to help
maintain this distinction. Distinguishing work from leisure
contexts might also allow di↵ erent technologies to be used
in each case.</p>
          <p>People. Our participants did associate their objects in terms
of specific people, but not in the way we might have expected.</p>
          <p>Rather than linking objects to particular, specified people,
our dimensions relate to relationships with the community:
“popular with many people”, for example, “things I will/won’t
talk about”, or “involvement of other people”. Unfortunately
it is not clear how a PIM or search tool could support this
sort of classification.</p>
          <p>
            Task, time, and workflow. Users in our study did not group
objects by particular tasks—objects related to task A, to
task B, etc—but four users did group objects according to
whether an object had an associated task, and by properties
of that task (state, importance, and cost or di culty). This
could be used to extend the work of Jones and his colleagues
[
            <xref ref-type="bibr" rid="ref28">16</xref>
            ], who advocate project organisation, but do not allow
tasks within projects to be annotated with properties such
as cost or importance.
          </p>
          <p>Time was mentioned by four participants. However rather
than categorise objects according to time of use (or receipt),
as supported by a number of tools, three participants derived
criteria from the lifecycle of an object. Criteria such as “when
I need to act on this” or “when this is important” will change
over time for each object. This is related to Cole’s “level of
information” dimension. Only one participant used objects’
importance to an event, at a particular time, as a criteria.</p>
          <p>Tools which support an explicit notion of document or task
lifecycle, or approximate this e.g. by recording patterns of
use, would presumably suit these participants.
5.</p>
          <p>DISCUSSION OF THIS APPROACH</p>
          <p>The results above suggest that card sorts, in this variant,
are useful for eliciting criteria: it does seem possible to
gain some insight into how users think of personal objects,
and how we might support this. This pilot has, however,
highlighted some limitations.</p>
          <p>The objects represented by each participant’s cards were
familiar—that is, they tended to choose objects they had used
recently or frequently. They were also selected for sharing,
since although we did not record the card titles we did see
them. We cannot be sure that the chosen objects represent
the sorts of things users may search for in a PIM system,
and of course they are not representative of unfamiliar
objects. We could instead choose objects from a participant’s
computer, for example by choosing randomly from the file
system and labelling cards with file icons and names, and
similar. There is a tradeo↵ , however: if participants did not
recognise these objects, the only possible criteria would be
file icons and names, and we would learn little. By allowing
users to choose their own objects the cards hopefully acted
as prompts for other, richer, associations.</p>
          <p>On a related point, some media types, such as video or
audio, are di cult to represent on cards. It is not clear what
this means for eliciting criteria. We are possibly unlikely to
get criteria such as “out of focus” (for photos) or “scratchy
bit in the middle” (for audio), but participants’ familiarity
with the objects may mitigate this to some extent.</p>
          <p>
            There are of course properties that are not captured by
this method: links between documents, for example (except
implicit links of the type “sorted into the same pile”). It is
also possible that our presence, and the apparatus we used,
made it hard for participants to think naturally. They may
have been inspired to create other categories if prompted
(as in Gonc¸alves and Jorge [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ]); on the other hand, our
approach has the advantage that we can see which properties
were immediately obvious.
          </p>
          <p>We also note that the objects chosen varied greatly from
participant to participant, and this may have played a role
in the criteria that were chosen—although we did see some
overlap, possibly there would have been more if the objects
were more similar. We are considering constraining
participants more in future, for example by prompting them to
make a certain number of cards for each media or perhaps
having subpopulations sort a shared set of cards, e.g. emails,
web pages etc. they have all seen or received.</p>
          <p>CONCLUSIONS</p>
          <p>We hope to extend this work by scaling to a larger group of
participants, but we will consider some methological changes:
constraining the objects chosen, for example, or careful
prompts to elicit more classifications. Nonetheless card sorts,
in the variant here, have proved useful for starting to
understand how users think of “personal” digital objects from a
wide range of sources and media. Some classifications were
both common and expected, but we did observe interesting
di↵ erences both with the criteria found in earlier studies and
with the criteria exposed in PIM and search tools.</p>
          <p>ACKNOWLEDGEMENTS
We would like to thank our participants for their time.</p>
          <p>Using semantic differentials for an evaluative view of the
search engine as an interactive system</p>
          <p>Frances Johnson
Department of Languages, Information &amp; Communications</p>
          <p>Manchester Metropolitan University</p>
          <p>Geoffrey Manton
ABSTRACT
In this paper, we investigate the use of semantic differentials in
obtaining the evaluative view held by users of the search engine.</p>
          <p>The completed scales of bipolar adjectives were analysed to
suggest the dimensions of the user judgment formed when asked
to characterize a search engine. These were then used to obtain a
comparative evaluation of two engines potentially offering
different types of support (or assistance) during a search. We
consider the value of using the semantic differential as a technique
in the toolkit for assessing the user experience during information
interactions in exploratory search tasks.</p>
          <p>Categories and Subject Descriptors H3.3 [Information search
and retrieval]; Search process. H.5.2 [User interfaces]:
Evaluation/methodology
General Terms
Measurement, Performance, Design, Human Factors
1. INTRODUCTION
The design of interfaces to support exploratory search seeks to
provide users with the tools for and the experience of an
interactive and engaging search. This is a departure from the
classic model of information retrieval wherein the user submits a
keyword query to the system and scans the list of retrieved results
for relevance, either stopping with relevant results or refining the
query to get results that are closer to the information need.</p>
          <p>Exploratory search does not necessarily assume that the user has a
well defined information need (at least one that can be articulated
as  a  keyword  query)  or  indeed  that  the  query will  be ‘static’ and 
thus satisfied by a single list of retrieved results.</p>
          <p>Presented at EuroHCIR2012. Copyright © 2012 for the individual
papers by the papers' authors. Copying permitted only for private
and academic purposes. This volume is published and copyrighted
by its editors.</p>
          <p>
            Accordingly, search engine developments have focused on
providing query assistance drawing on contextual aspects to the
search, such as personal history and/or current context [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ]. At the
interface, developments focus on improving the search process via
richer information representations and interactions, such as
previews and facets through to tools that allow the user to view
and  explore  connections  in  the  results,  for  example  ‘ the relation
browser  data  analysis  tool’  [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ]. These shifts into HCIR are
intended to help in the various stages of search, from starting the
task and understanding the query topic, throughout the search in
deciding what to do next, and to stopping with a sense of
confidence. In short, developments aim to support true
exploration of the search and, whilst many efforts may fall short,
they will provide some form of user support in query assistance
and in improving the search process as an interactive experience.
          </p>
          <p>
            The context for evaluation is predicated on White and Roth’s [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ] 
model of the exploratory search process. This involves the
searcher in a dynamic interplay between the cognition of their
‘problem  space’  and  their  exploratory  activities  in  the  iterative 
search process including the query formulation, results
examination and information extraction. Data collected on the
searcher’s information intera ctions may confirm this model [
            <xref ref-type="bibr" rid="ref19 ref7">7</xref>
            ] as
well as attempt to systematically evaluate the effectiveness of
exploratory search systems. In evaluation, a framework is used to
attempt to assess performance during the search stages and to
relate aspects of the system to its role in supporting information
exploration, including sense making or query visualisation [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ].
          </p>
          <p>The challenge for the evaluation of exploratory search is the
assumption that the user is willing or able to make an evaluative
judgment throughout the search or that valid measures can be
found through their actions, for example of usage of query terms.</p>
          <p>In general, evaluation draws from established HCI measures of
effectiveness (can people complete their tasks?) efficiency (how
long  do  people  take?),  an  assessment  of  the  user’s  overall 
satisfaction or other affective responses. Where possible, and
increasingly so, the user actions are observed and recorded as
dependent on the system and/or its interface. In this study we
focus  on  an  attempt  to  obtain  the  user’s  evaluative  view  of  the 
search engine, based on criteria which may be affected by the
developments for new and richer interactive designs. It is
assumed that this would be part of an assessment which when
taken  with  others  will  build  a  picture  of  the  ‘user  experience’  of 
the system used in exploratory search.</p>
          <p>USER EVALUATION
In developing an instrument to collect the user assessment effort
goes into ensuring that the evaluation is made in the task context.</p>
          <p>It means little to know that the user is ‘satisfied’ with the interface 
without gaining insight into why this assessment has been formed.</p>
          <p>
            A variety of questionnaires have been developed for assessing
usability of interactive systems, such as search engines. Two well
known are the SUS (System Usability Scale) developed at the
Digital Equipment Corporation [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ] and the QUIS (Questionnaire
for User Interaction Satisfaction) from the University of Maryland
[
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ]. Both assess usability from the user perspective with 10
statements and rating scales in the SUS and the QUIS with 27
questions. The QUIS asks the user to respond on a rating scale to
statements which address specific usability aspects of the system,
such as “use of the terms were consistent throughout the website”. 
The  SUS  on  the  other  hand  focuses  on  collecting  the  users’ 
overall reaction to the site/system on statements, such as “I found 
the website unnecessarily complex”.  Arguably the QU IS focuses
on the concerns that a developer might have when assessing
usability  whilst  the  SUS  assumes  that  the  user’s  overall 
assessment is a reflection on the extent to which their goal
directed tasks were facilitated by the system and its design.
          </p>
          <p>
            Questionnaires, such as SUS, are used in an experimental set up
when an  explanation  of  the  user’s  overall  assessment  is  sought. 
However, the limitations of the questionnaire to capture and
provide  insight  into  the  complexity  of  the  user’s  assessment  has
lead to alternative tools,  for  example  Microsoft’s  Product 
Reaction Cards in the "Desirability Toolkit". This invites
participants on a usability test to select as many, or as few, words
from a list of 118 which best describe their reaction and/or
interaction with the system they have just used. Benedek and
Miner [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ] includes a list of the words used and point out that the
approach helps elicit negative comments as well as positive, thus
overcoming a problem with questionnaires biased towards
positive responses.
          </p>
          <p>Given the potential scope  of  the  users’  response  (represented  in
the reaction cards with some 100+ terms) this study sets out to
investigate the value in assembling these into a framework (of
sorts)  for  the  collection  of  the  users’  evaluative  judgment  of  an
interactive system based on the technique known  as  ‘semantic 
differentials’.  Specifically the aim of this small preliminary
investigation was to begin to determine the extent to which users
hold  an  evaluative  view  of  a  ‘search  engine’  and,  what  are  the 
dimensions (traits or criteria) on which we form this view. If it
can be found that this view is strongly held (that is, an attitude is
formed which may influence how we behave and interact with the
search engine) then it may be feasible to investigate the influence,
if any, of a design for information interaction on the evaluative
view. In this study the technique of semantic differentials is used
to best describe the evaluative view held by its participants. This
is then employed to assess two quite different search engines
following the completion of two query based searches.</p>
          <p>
            SEMANTIC DIFFERENTIALS
Semantic Differentials (SDs) originate from the work of Osgood
[
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ] as a technique for attitude measurement, scaling people on
their responses to adjectives in respect to a concept. Typically
individuals respond to several pairs of bipolar adjectives scored
on a continuum + to – and in doing so differentiate their meaning
of the concept in intensity and in direction (in a ‘semantic space’). 
          </p>
          <p>The assumption made here, in the use of SDs on ‘search engines’ 
is that users hold an evaluative view which is formed when using
the engine to find and/or explore information. The SD is used to
investigate the adjectives that best ‘conceptualise’ the search
engine, from the user perspective. Factorial analysis is also used
to identify the dimensions of the judgment, in a sense the
packaging of the components of the judgment into smaller units of
meaning reflecting what is important when responding to the
concept ‘search engine’.</p>
          <p>The design of the SD aims to allow a degree of abstraction in the
evaluation so that participants can reflect the complexity of their
response. In this study, the adjectives to include on the SD scale
were chosen from Microsoft’s  Product  Reaction  Cards,  these 
having been collected in previous research, usability studies and
in the marketing of web sites and systems. The majority of the
terms formed pairs on some continuum and 40 terms (20 pairs)
were selected to present in the SD. The selection was subject to
the judgment of the researcher. This is a limitation of this
exploratory study, however some steps were taken to formalise the
selection. A loose grouping of the adjective pairs was made as
relating to appearance (such as ‘attractive’ ), judgment (‘relevant’),
emotive (‘bor ing’) and use (‘fast’).  Five  pairs from each of these
groupings were made. The pairs were mixed on the SD to avoid
having all the positive terms on one side of the scale and only
intervals were shown on the scales with the numerical values used
only for data entry. This allowed participants to focus on how an
adjective pair related to the engine and its characteristics, rather
than on ‘scoring’ it in some way. 
3.1</p>
          <p>Implementation
The study was conducted on our undergraduates studying BSc
Web Development and on a postgraduate cohort studying on MA
Library and Information Management or the MSc Information
Management. A total of 89 students participated in the study. At
the start of the class each participant was asked to think about a
search engine, and adjectives they would use to describe the
engine,  (in  other  words,  “what  it  means  to  them” ). Each
participant was then given the SD to complete. This is referred to
as  the  ‘baseline’  and  the  data  we re analysed to gauge user
perceptions of search engines.</p>
          <p>In the following lab sessions (about one hour later) each
participant was required to perform two search tasks on each of
the two search engines - Google, an engine we can assume some
familiarity and, a second clustering engine (Yippy, formerly
Clusty). The two tasks were as follows
These were selected to give the participants experience of using
the engines for a closed question (find symptoms) and on a more
open  ‘informational’  type  of  query  (on  the  ‘nature  nurture’ 
debate). A measure of search success was not taken as the aim
was simply to get the participants using the engines. The order of
use of the two sites was randomized so that approximately half of
the participants worked on Google first and half on the clustering
engine. All were told to spend no longer than 10 minutes
searching on each engine and to complete the SD for each engine
immediately after each use.
4.1</p>
          <p>FINDINGS</p>
          <p>Evaluative views
The responses to the baseline (think of an engine) were entered
into SPSS with the scales coded (7-1) so that the positive
adjectives corresponded to the higher numbers. Descriptive
statistics of mean, mode and standard deviation were calculated
for each of the adjectives. Those with a mean greater than 4 or
less than 3 were taken to suggest the adjective pairs that best
characterise the participants’  view, as follows
attractive
powerful
valuable
relevant
satisfying
fast
predictable
intuitive
easy
unattractive
simplistic
not valuable
irrelevant
frustrating
slow
unpredictable
rigid
difficult
Factor analysis investigates the correlations among subsets of the
responses to the bipolar pairs and groups the correlated variables
such that each group is largely independent of the others.</p>
          <p>
            Exploratory factor analysis was employed to identify the groups
which might explain most of the variance in the data. With 20
pairs of adjectives to perform Principal Components Analysis
(PCA) in SPSS it is recommended that a minimum of 100
responses are obtained, whilst others recommend that a sample
requires approx 5-10 times the number of people as scale pairs
[
            <xref ref-type="bibr" rid="ref18 ref6">6</xref>
            ]. With 89 responses we should use a reduced number of pairs,
however the Kaiser-Meyer-Olkin measure of sampling adequacy
(.616) is greater than the 0.6 needed to indicate that the
correlations matrix may be able to factorise. So with this, PCA
was run (with varimax rotation to force items to ‘load’ with only 
one  factor  group),  to  identify  the  possible  ‘factors’  or  subsets 
derived from patterns of correlation of the adjective pairs. The
following five subsets were obtained (the adjectives from the list
above having a low or high mean are shown in bold). The labels
were assigned to suggest the evaluative dimension.
          </p>
          <p>Factor 1 label USE – Utility
effective, valuable, satisfy, relevant, predictable,</p>
          <p>intimidating, inspiring, stimulating
Factor 2 label QUALITY – Affective</p>
          <p>engaging, fun, connected
Factor 3 label QUALITY - Appearance</p>
          <p>high quality, personal, meaningful, rigid, attractive
Factor 4 label USE – Efficient</p>
          <p>easy, intuitive, fast, powerful
Factor 5 label USE - Control
controllable
4.2 Comparative evaluations
Using the same SDs, participants scaled their responses post
search using Google and the clustering search engine. These were
entered into a worksheet to obtain basic statistics. The mode for
each adjective is shown Figure 1 with a note of those with mode
&gt;4 and &lt; 3 suggesting a positive or negative response.</p>
          <p>Google (mode &gt; 4 or &lt; 3)
&amp; in bold where mean is also &gt; 4
1attractive - , 6valuable - , 8relevant - ,
15satisfying - , 16fast - , 17predictable - ,18controllable -,
and (where mode &lt; 3) 19rigid</p>
          <p>Clustering search engine (mode &gt; 4 or &lt; 3)
&amp; in bold where mean &gt; 4 or &lt; 3
14engaging - , 19intuitive –
and (where mode &lt; 3)
13intimidating - , 17 – unpredictable
Using the suggested dimensions or aspects of the user evaluation
from the factor analysis of the ‘baseline’ data we can compare the
participants’ responses on the high or low scoring adjectives
across the engines. On QUALITY – Appearance Google
was rated rigid and attractive and whereas Google was neutral on
the factor QUALITY- Affective, the clustering search
engine obtained a positive score towards the adjective engaging.</p>
          <p>On the factor labeled USE- Utility Google was scored as
predictable, valuable, relevant and satisfying, whereas the
clustering engine as unpredictable and towards intimidating. On
USE-Efficient Google was rated as fast and the clustering
engine appears more intuitive. Google was also rated as
controllable.</p>
          <p>DISCUSSION
This is an exploratory study and it has its limitations. It is
questionable whether the selection of the adjectives to use in the
SD influenced the results. In particular there is uncertainty in the
results that intuitive to rigid is on some continuum. Also there is
some unease at accepting a factor with 8 out of 20 pairs and one
with only one. Perhaps the sample size was too small to attempt
factoring. The results also raise questions on how some of the
adjectives were interpreted by the participants. These
withstanding, the participants in this study did appear to hold an
evaluative judgment  of  the  concept  ‘search  engine’ and the traits
represented in the scale were grouped to suggest the aspects on
which an assessment may be formed. It is of particular interest
that upon using the search engine Google to conduct a search task
the ratings on the SD, on the whole, altered only in the factors of
‘controllable’  and  USE –efficient (easy, intuitive and
powerful). Perhaps we can assume that Google was the typical
engine when asked to think of an engine in the baseline and, when
it came to use Google, users shifted their perception with regards
to some of the adjectives. Perhaps this is not surprising but it may
suggest that we hold an implicit view of search engines, and that
this view will be influenced by actual use (and the experience).</p>
          <p>Our participants may have had less familiarity with the clustering
engine, and in the evaluation this appears to have prompted an
‘affective’ response in finding the engine to be ‘engaging’ whilst 
also indicating shifts in  the ‘use’ factors (towards an assessment 
of  the engine as ‘unpredic table’).  Again the infallibility of some
of the terms is highlighted where an ‘unpredictable’ system may
be regarded to be a negative judgment, but if the system is also
considered to be engaging the assessment could be highly
desirable depending  on  the  user’s  goals.  This study of the use of
semantic differentials indicates that it is worth running the test
with a new cohort of students to determine the extent to which a
consistent view is obtained. As an exploratory study it also
suggests that further research on user’s perceptions and mental
models of search engines is worthwhile. With regards to the
challenge of providing an evaluation of the exploratory search,
this study falls short as no behavioural data was obtained.</p>
          <p>However, perhaps, with further design of the SD and use in an
experimental set up with honed tasks, a user assessment of the
interface may be obtained as dependent on the search interface
development and design.
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
unattractive
personal
fun
simplistic
connected
not valuable
low quality
relevant
ineffective
meaningful
confusing
inspiring
empowering
engaging
frustrating
slow
unpredictable
uncontrollable
rigid
eas
The fault, dear researchers, is not in Cranfield,
But in our metrics, that they are unrealistic.</p>
          <p>Mark D. Smucker
Department of Management Sciences</p>
          <p>University of Waterloo, Canada
mark.smucker@uwaterloo.ca</p>
          <p>Charles L. A. Clarke</p>
          <p>School of Computer Science
University of Waterloo, Canada
claclark@plg.uwaterloo.ca
1. INTRODUCTION</p>
          <p>As designers of information retrieval (IR) systems, we
need some way to measure the performance of our systems.</p>
          <p>
            An excellent approach to take is to directly measure actual
user performance either in situ or in the laboratory [
            <xref ref-type="bibr" rid="ref12 ref24">12</xref>
            ]. The
downside of live user involvement is the prohibitive cost if
many evaluations are required. For example, it is common
practice to sweep parameter settings for ranking algorithms
in order to optimize retrieval metrics on a test collection.
          </p>
          <p>The Cranfield approach to IR evaluation provides low-cost,
reusable measures of system performance.</p>
          <p>
            Cranfield-style evaluation frequently has been criticized as
being too divorced from the reality of how users search, but
there really is nothing wrong with the approach [
            <xref ref-type="bibr" rid="ref30">18</xref>
            ]. The
Cranfield approach effectively is a simulation of IR system
usage that attempts to make a prediction about the
performance of one system vs. another [
            <xref ref-type="bibr" rid="ref27">15</xref>
            ].
          </p>
          <p>As such, we should really be thinking of the Cranfield
approach as the application of models to make predictions,
which is common practice in science and engineering. For
example, physics has equations of motion. Civil engineering
has models of concrete strength. Epidemiology has models
of disease spread. Etc. In all of these fields, it is well
understood that the models are simplifications of reality, but that
the models provide the ability to make useful predictions.</p>
          <p>Information retrieval’s predictive models are our
evaluation metrics.</p>
          <p>The criticism of system-oriented IR evaluation should be
redirected. The problem is not with Cranfield — which is
just another name for making predictions given a model —
the problem is with the metrics.</p>
          <p>We believe that rather than criticizing Cranfield, the
correct response is to develop better metrics. We should make
metrics that are more predictive of human performance. We
should make metrics that incorporate the user interface and
realistically represent the variation in user behavior. We
should make metrics that encapsulate our best
understanding of search behavior.</p>
          <p>
            In popular parlance, we should bring solutions, not
problems, to the system-oriented IR researcher. To this end,
we have developed a new evaluation metric, time-biased
gain (TBG), that predicts IR system performance in
human terms of the expected number of relevant documents
to be found by a user [
            <xref ref-type="bibr" rid="ref28">16</xref>
            ].
          </p>
          <p>Presented at EuroHCIR2012. Copyright (C) 2012 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
2.</p>
          <p>TIME-BIASED GAIN</p>
          <p>
            HCI has a long history of automated usability
evaluation [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ], and indeed, so does IR. Cleverdon designed the
Cranfield 2 study carefully in terms of a specific type of
user and how this type of user would define relevance [8,
p. 9]. Taken together, a test collection (documents, topics,
relevance judgments) and an evaluation metric allow for the
simulation of a user with different IR systems.
          </p>
          <p>
            Ja¨rvelin and Keka¨l¨ainen produced a significant shift in
evaluation metrics with their introduction of cumulated
gainbased measures [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ]. The cumulated gain measures are
explicitly focused on a model of a user using an IR system. As
long as the user continues to search, the user can continue to
increase their gain. The common notion of gain in IR
evaluation is the relevant document, but gain can be anything
we would like to define it to be.
          </p>
          <p>Cumulated gain can be plotted vs. time to produce a gain
curve and compare systems. The curve that rises higher and
faster than another curve is the preferred curve. While we
can plot gain curves of one system vs. another, it is
wellknown that users do not endlessly search; different users
stop their searches at different points in time for a host of
reasons. Given a probability density function f (t) that gives
the distribution of time spent searching, we can compute the
expected gain as follows:</p>
          <p>E[G(t)] =
! ∞
0</p>
          <p>G(t)f (t)dt,</p>
          <p>(1)
where G(t) is the cumulated gain at time t. Equation 1
represents time-biased gain in its general form, i.e.
timebiased gain is the expected gain for a population of users.</p>
          <p>While it is natural for us to talk about cumulated gain over
time, the traditional cumulated gain measures have
substituted document rank for time and implicitly model a user
that takes the same amount of time to evaluate each and
every document. By making time a central part of our
metric, we gain the ability to more accurately model behavior.</p>
          <p>
            For example, in a document retrieval system, longer
documents will in general take users longer to evaluate, and if the
retrieval system presents results with document summaries
(snippets), we know that users can use summaries to speed
the rate at which they find relevant information [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ].
          </p>
          <p>Another significant advantage of using time directly in our
retrieval metric is that we now make testable predictions of
human performance. Our predictions are in the same units
as would be obtained as part of a user study. To our
knowledge, this alignment between the units of Cranfield-style
metrics and user study metrics has not previously existed.</p>
          <p>
            Time-biased gain in the form of Equation 1 makes no
mention of ranked lists of documents, for it is a general purpose
description of users using an IR system over time. To
produce a metric suitable for use in evaluating ranked lists, we
followed a process common to development of new
simulations [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ]:
1. Creation of model.
2. Calibration of model.
          </p>
          <p>3. Validation of model.</p>
          <p>Our first step in model creation was to adopt the standard
model of a user that works down a result list and move
Equation 1 to a form common to cumulated gain measures:
∞
! gkD(T (k)),
k=1</p>
          <p>(2)
where gk is the gain at rank k, T (k) is the expected time it
takes a user to reach rank k, and D(t) is the fraction of the
population that survives to time t and is called the decay
function.</p>
          <p>Our model for the time it takes a user to reach rank k,
T (k), takes into consideration a hypothetical user interface
that presents results to the user in the form of document
summaries. A click on a document summary takes the user
to the full document. We model both the probabilities of
clicking on summaries given their NIST relevance and the
probability of then judging a viewed full document as
relevant. We separately model the time to view summaries and
full documents. For the time spent on a full document, we
modeled longer documents taking longer with an additional
constant amount of spent. We treated duplicate documents
as zero length documents. We then calibrated T (k) using
data from a user study, and finally we validated that our
T (k) provided a reasonable fit to the user study data.
Likewise, we modeled D(t) as exponential decay fit to a search
engine’s log data.</p>
          <p>
            In contrast, older evaluation metrics such as mean average
precision [19, p. 59] cannot be calibrated and have only
been validated after their creation. For example, the work
of Hersh and Turpin [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ] is likely the first attempt to validate
a metric (average precision). Many recent metrics can be
calibrated to actual user behavior [
            <xref ref-type="bibr" rid="ref16 ref17 ref19 ref29 ref4 ref5 ref7">4, 5, 7, 17, 20, 21</xref>
            ], but
their calibration and validation often come after their release
and adoption.
3. CONCLUSION
          </p>
          <p>
            The Cranfield approach to IR evaluation is merely
another name for the development and use of predictive
models, which is a fundamental part all science and engineering
fields. In particular, IR evaluation fits nicely into the
framework of simulation where models are created, calibrated, and
validated before being used to make predictions. We have
presented time-biased gain as an example of what we believe
the correct direction is for IR system evaluation. We are not
the only ones to be working on better metrics or taking a
simulation based approach [
            <xref ref-type="bibr" rid="ref14 ref2 ref25">2, 13</xref>
            ], and others also consider
time an important part of evaluation [
            <xref ref-type="bibr" rid="ref1 ref13 ref18 ref6">1, 6</xref>
            ].
          </p>
          <p>Our position is that system-oriented IR research is
useroriented IR research given its use of evaluation metrics that
model users. If HCIR researchers can produce better
models than exist today — by better, we mean more predictive
of human performance — then we can help system
development to focus on changes that help users better search.
4. ACKNOWLEDGMENTS</p>
          <p>This work was supported in part by the NSERC, in part
by GRAND NCE, in part by Google, in part by Amazon,
in part by the facilities of SHARCNET, and in part by the
University of Waterloo. Any opinions, findings and
conclusions or recommendations expressed in this material are the
authors’ and do not necessarily reflect those of the sponsors.
5. REFERENCES</p>
          <p>
            A Model of Consumer Search Behaviour
In order to design better search experiences, we need to
understand the complexities of human information-seeking
behaviour. In previous work [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ], we proposed a model of
information behavior based on an analysis of the information
needs of knowledge workers within an enterprise search context.
          </p>
          <p>In this paper, we extend this work to the site search context,
examining the needs and behaviours of users of
consumeroriented websites and search applications.</p>
          <p>We found that site search users presented significantly different
information needs to those of enterprise search, implying some
key differences in the information behaviours required to satisfy
those needs. In particular, the site search users focused more on
simple  “lookup”  activities,  contrasting  with  the  more  complex ,
problem-solving behaviours associated with enterprise search. We
also found repeating  patterns  or  ‘chains’  of  search  behaviour in
the site search context, but in contrast to the previous study these
were shorter and less complex. These patterns can be used as a
framework for understanding information seeking behaviour that
can be adopted by other researchers  who  want  to  take  a  ‘needs 
first’ approach to understanding information behaviour .</p>
          <p>Categories and Subject Descriptors
H.3.3 [Information Search and Retrieval]: Search process;
H.3.5 [Online Information Services]: Web-based services
General Terms
Human Factors.</p>
          <p>
            Keywords
Site search, enterprise search, information seeking, user
behaviour, search modes, information discovery, user experience
design.
1. INTRODUCTION
Classic IR (information retrieval) is predicated on the notion of
users searching for information in order to satisfy a particular
'information need'. However, it is now accepted that much of what
we recognize as search behaviour is often not informational per
se. For example, Broder [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ] has shown that the need underlying a
given web search could in fact be navigational (e.g. to find a
particular site) or transactional (e.g. through online shopping,
social media, etc.). Similarly, Rose &amp; Levinson [
            <xref ref-type="bibr" rid="ref12 ref24">12</xref>
            ] have
identified the consumption of online resources as a further
common category of search behaviour.
          </p>
          <p>
            Presented at EuroHCIR2012. Copyright © 2012 for the individual papers
by the papers' authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
In this paper, we examine the needs and behaviours of individuals
across a range of site search scenarios. These are based on an
analysis of user needs derived from a series of customer
engagements involving the development of customised site search
applications. In so doing, we extend and validate a model of
information behaviours derived from a previous study of
enterprise search users [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ].
          </p>
          <p>The model is based on a set of ‘search modes’ that users employ 
to satisfy their information search and discovery goals. It extends
the IR concept of information-seeking to embrace a broader
notion of discovery-oriented problem solving, addressing a wider
range of information interaction and information use behaviours.</p>
          <p>
            The overall structure of the model reflects Marchionini’s  [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ]
framework, and consists of three lower-level  ‘lookup’  modes 
(locate, verify and monitor),  three  “learn”  modes  ( compare,
comprehend and explore) and three higher-level  “investigate” 
modes (analyze, evaluate and synthesize).
          </p>
          <p>
            We investigate the degree to which the model extends to
accommodate the domain of site search (i.e. consumer-oriented
websites and search applications) and discuss some of the
differences between the needs and goals of enterprise search users
versus those of site search. We conclude by exploring the ways in
which these modes combine to form distinct chains or patterns,
and reflect on the value this offers as a framework for expressing
complex patterns of behaviour.
2. MODELS OF INFORMATION SEEKING
The framework investigated in this study is influenced by a
number of existing models. For example, Bates [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ] identified a set
of  29  search  ‘tactics’  which  she  organised into four broad
categories, including monitoring (“to  keep  a  search  on  track”). 
Likewise, O’Day &amp; Jeffries  [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ] examined the use of information
search results by clients of professional information intermediaries
and identified three categories of behaviour, including monitoring
a known topic or set of variables over time and exploring a topic
in an undirected fashion. They also observed that a given search
scenario would often evolve into a series of interconnected
searches, delimited by triggers and stop conditions that signalled
transitions between modes within an overall scenario.
          </p>
          <p>
            Cool &amp; Belkin [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ] proposed a classification of interaction with
information which included evaluate and comprehend. They also
proposed create and modify, which together reflect aspects of our
synthesize mode.
          </p>
          <p>
            Ellis and his colleagues [
            <xref ref-type="bibr" rid="ref16 ref17 ref18 ref4 ref5 ref6">4, 5, 6</xref>
            ] developed a model consisting of
a number of broad information seeking behaviours, including
monitoring and verifying (“checking  the  information  and sources 
found for accuracy  and  errors”) . In addition, his browsing mode
(“semi-directed  searching  in  an  area  of  potential  interest”)  aligns
with our definition of explore. He also noted that it is possible to
display more than one behaviour at any given time. In revisiting
Ellis’s  findings  among  social  scientists, Meho and Tibbo [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ]
identified analysing (although they did not elaborate on it in
detail). More recently, Makri et al [
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ] proposed searching
(“formulating  a  query  in  order  to  locate  information”),  whic h
reflects to our own definition of locate.
          </p>
          <p>
            In addition to the research-oriented models outlined above, we
should also consider practitioner-oriented views. Spencer [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ]
suggests four modes of information seeking, including
knownitem (a subset of our locate mode) and exploratory (which mirrors
our definition of explore). Lamantia [
            <xref ref-type="bibr" rid="ref19 ref7">7</xref>
            ] also identifies four
modes, including monitoring.
          </p>
          <p>
            In this paper, we use the characteristics of the models above as a
lens to interpret the behaviours found in a new source of empirical
site search data. We also explore the combinatorial nature of the
modes,  extending  Ellis’s  [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ] concept of mode co-occurrence to
identify and define a set of repeating patterns and sequences.
3. CONSUMER SEARCH BEHAVIOUR
3.1 Data Acquisition
The primary source of data in this study is a set of 277
information needs captured during client engagements involving
the development of a number of custom site search applications.
          </p>
          <p>These information needs take the form of ‘micro-scenarios’, i.e.  a
brief narrative that illustrates the end user’s goal and the  primary
task or action they take to achieve it, for example:



</p>
          <p>Find best offers before the others do so I can have a
high margin.</p>
          <p>Get help and guidance on how to sell my car safely so
that I can achieve a good price.</p>
          <p>Understand what is selling by area/region so I can
source the correct stock.</p>
          <p>See year-on-year ad spend trends for TV and online to
supply to the Head of Global Media.</p>
          <p>The scenarios were collected as part of a series of requirements
workshops involving stakeholders and customer-facing staff from
the respective client organisations. They were generated by
participants in individual breakout sessions, and then moderated
by the workshop facilitator in a group session to maximise
consistency and minimise redundancy or ambiguity. They were
also prioritised by the group to identify those that represented the
highest value both to the end user and to the client organisation.</p>
          <p>
            This data possesses a number of unique properties. In previous
studies of information seeking behaviour (e.g. [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ], [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ]), the
primary source of data has traditionally been interview transcripts
that provide an indirect, verbal account of end user information
behaviours. By contrast, the current data source represents a
selfreported account of information needs, generated directly by end
users (although a proportion were captured via proxy, e.g. through
customer facing staff speaking on behalf of the end users). This
change of perspective means that instead of using information
behaviours to infer information needs and design insights, we can
adopt the converse approach and use the stated needs to infer
information behaviours and the interactions required to support
them.
          </p>
          <p>
            Moreover, the scope and focus of these scenarios represents a
further point of differentiation. In previous studies, (e.g. [
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ]),
measures have been taken to address the limitations of using
interview data by combining it with direct observation of
information seeking behaviour in naturalistic settings. However,
the behaviours that this approach reveals are bounded by the
functionality currently supported by existing systems and working
practices, and as such do not reflect the full range of aspirational
or unmet user needs encompassed by the scenarios in this study.
          </p>
          <p>
            Finally, the data is unique in that is constitutes a genuine
practitioner-oriented deliverable, generated expressly for the
purpose of designing and delivering professional site search
systems. As such, it reflects a degree of realism that interview data
or other research-based interventions might struggle to replicate.
3.2 Data Analysis
These scenarios were analyzed using the model derived previously
for the domain of enterprise search [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ]. In this respect, the
process was partially deductive, applying the model in a top-down
fashion to classify the data. But it was also partially inductive,
applying a bottom-up, grounded analysis to identify new types of
behaviour not present in the original model or to suggest revised
definitions of the existing categories.
          </p>
          <p>Although the original study involved three separate analysts, the
behaviours this time were identified by the first author alone. The
current analysis approach is therefore much more subjective.</p>
          <p>However, the first author was also the facilitator at each of the
requirements workshops at which the scenarios were generated,
and was able to again a deep insight into the needs, goals and
motivations of the participants. This allowed him to be as
confident as possible  in  his  understanding  of  the  users’ 
information needs and consistent in his interpretation of the
information behaviours required to satisfy a particular need.</p>
          <p>
            A number of the scenarios focused on needs that did not involve
any explicit information seeking or use behaviour, e.g. “Achieve a
good price for my current car”.  These were excluded from the
analysis. A further number were incomplete or ambiguous, or
were essentially  feature  requests  (e.g.  “ Have flexible navigation
within the page”), and were also excluded. This process resulted
in further confirmation and validation of the nine search modes
identified in the original study, but with revised definitions to
reflect a broader scope:
1. Locate: To find a specific (possibly known) item, e.g. “ Find my
reading list items quickly”.  This mode encapsulates the
stereotypical ‘findability’  task that is so commonly associated
with site search, consistent with (but a superset of) Spencer’s  [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ]
known item search mode. This was the most frequent mode in the
site search scenarios (120 instances).
2. Verify: To confirm that an item meets some specific, objective
criterion, e.g. “S ee the correct price for singles and deals”. Often
found in combination with locating, this mode is concerned with
validating the accuracy of some data item, comparable to that
proposed by Ellis et al. [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ] (39 instances).
3. Monitor: Maintain awareness of the status of an item for
purposes of management or control,  e.g.  “ Alert me to new
resources in my area”. This activity focuses on the state of
asynchronous responsiveness and is consistent with that of Bates
[
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ],  O’Day  and  Jeffries [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ], Ellis [
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ], and Lamantia [
            <xref ref-type="bibr" rid="ref19 ref7">7</xref>
            ] (13
instances).
4. Compare: To identify similarities &amp; differences within a set of
items, e.g.  “ Compare cars that are my possible candidates in
detail”. This mode has not featured prominently in previous
models (with the possible exception  of  Marchionini’s),  but  was
found to be a significant component of enterprise search
behaviour [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ]. Moreover, it is a common feature of product
search and navigation on many ecommerce sites. However, it
occurred relatively infrequently in the site search scenarios (2
instances).
5. Comprehend: To generate independent insight by interpreting
patterns within a data set, e.g. “Understand what my competitors
are selling”.  Like compare, this mode was found to be a key
element of the enterprise search scenarios, and also features in the
models of Cool &amp; Belkin [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ] and Marchionini [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ]. It occurred
relatively frequently in site search (50 instances).
6. Explore: To investigate an item or data set for the purpose of
knowledge discovery, e.g. “ Find useful stuff on my subject topic”.
          </p>
          <p>
            In some ways the boundaries of this mode are somewhat less
prescribed than the others, but what the instances share is the
characteristic of open ended, opportunistic search and browsing in
the  spirit  of  O’Day  and  Jeffries  [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ] exploring a topic in an
undirected fashion and Spencer’s  [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ] exploratory. This mode
was the second most common in site search (110 instances).
7. Analyze: To examine an item or data set to identify patterns &amp;
relationships, e.g. Analyze the market so I know where my
strengths and weaknesses are”.  This mode features less
prominently in previous models, appearing as a sub-component of
the processing stage  in  Meho  &amp;  Tibbo’s  [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ] model, and
overlapping  somewhat  with  Cool  &amp;  Belkin’s  [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ] organize. This
definition is also consistent with that of Makri et al. [
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ], who
identified analysing as an important aspect of lawyers’ interactive 
information  behaviour  and  defined  it  as  “examining  in  detail  the 
elements or structure of the content found during
informationseeking.”  (p.  630).  Although the most common element of the
enterprise search scenarios, it was less prevalent in site search (59
instances).
8. Evaluate: To use judgement to determine the value of an item
with respect to a specific goal, e.g. “ I want to know whether my
agency is delivering best value”.  This mode is similar in spirit to
verify, in that it is concerned with validation of the data. However,
while verify focuses on simple, objective fact checking, our
conception of evaluate involves more subjective,
knowledgebased judgement, similar to that proposed by Cool &amp; Belkin [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ]
(61 instances).
9. Synthesize: To create a novel or composite artefact from
diverse inputs, e.g.  “ I need to create a reading list on celebrity
sponsorship”.  This mode also appears as a sub-component of the
processing stage  in  Meho  &amp;  Tibbo’s  [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ] model, and involves
elements of Cool &amp; Belkin’s  [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ] create and use. Of all the modes,
this one is the most commonly associated with information use in
its broadest sense (as opposed to information seeking). It was
relatively rare within site search (5 instances).
4. MODE SEQUENCES AND PATTERNS
Applying the modes described above provides a framework for
understanding the needs of site search users, and an insight into
their likely behaviours. But as with the previous study [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ], their
real value lies not so much in the individual instance data but in
the patterns of co-occurrence they reveals. In most scenarios,
modes combine to form distinct chains and patterns, echoing the
transitions observed  by  O’Day  and  Jeffries  [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ] and the
combinatorial behaviour alluded to by Ellis [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ], who suggested
that information behaviours can often be nested or displayed in
parallel.
          </p>
          <p>Just as new definitions were needed to accommodate the new
domain, new patterns of occurrence were identified in the data.</p>
          <p>Typically these consisted of chains of length two or three, of
which the following were most frequent:</p>
          <p>Insight-driven search: (Explore-&gt;Analyze-&gt;
Comprehend): This patterns represents an exploratory
search for insight to resolve an explicit information
need, e.g. “ Assess the proper market value for my car”
(45 instances)
Opportunity-driven search:
(Explore-LocateEvaluate): In contrast to the explicit focus of the pattern
above, this sequence represents a less directed
exploration in the prospect of serendipitous discovery
e.g.  “ Find useful stuff on my subject topic”(31
instances)
Qualified search (Locate-Verify) This pattern
represents a variant of the stereotypical findability task
in which some element of immediate verification is
required, e.g.  “ Find trucks that I am eligible to drive”
(29 instances)
A deeper insight into these patterns can be obtained by presenting
them in diagrammatic form, as a network (Figure 1). This diagram
illustrates the three sequences outlined above plus other
commonly found patterns. It also reflects an outcome of the
pervious study, in that certain modes tend to function as
“terminal” nodes, i.e. entry points or exit points to a scenario. For 
example, Explore typically functions as an opening, while
Comprehend and Evaluate function in closing a scenario. Analyze
typically appears as a bridge between an opening and closing
mode.
4.1 Site search vs. Enterprise Search
The sequences described above also allow us to reflect on some of
the differences between the needs of site search users and those of
enterprise search. One of the most fundamental differences is an
emphasis  on  simpler  “lookup”  modes  such  as  Locate and Verify:
these were relatively rare in the enterprise search data, but
prominent in site search (120 and 39 instances respectively).</p>
          <p>Enterprise search, by contrast, emphasised higher-level
“investigate”  behaviours such as Analyze and Evaluate (modes
which also appeared frequently in site search, but not as
prominently: 58 and 61 instances respectively). However, in
neither case was the stereotype of ‘search as findability’  borne
out: even in site search (where it was the most common mode),
Locate was accountable for no more than a quarter of all
instances.</p>
          <p>But perhaps the biggest difference was in the composition of the
chains: while enterprise search was characterised by a wide
variety of heterogeneous chains, site searched focused on a small
number of common trigrams and bigrams. Moreover, these chains
displayed little evidence of the composite nature observed in
enterprise search, in which certain chains were seen to be
embedded within others to create larger, more complex sequences
of behaviour.
5. DISCUSSION
A key feature of the current model is its emphasis on the
combinatorial nature of search modes, and the value this offers as
a framework for expressing complex patterns of behaviour. Such
an approach is not unique: the second author, for example, has
also previously explored the concept of mode chains to describe
information seeking behaviours observed in naturalistic settings.</p>
          <p>However, his approach was based on the analysis of complex
tasks observed in real time, and as such was less effective in
revealing consistent patterns of atomic behaviour such as those
found in the current study.</p>
          <p>Conversely, this virtue can also be a shortcoming: the fact that
simple repeating patterns can be extracted from the data may be as
much an artefact of the medium as it is of the information needs it
contains. These scenarios were expressly designed to be a concise,
self-contained deliverable in their own right, and applied as a
simple but effective tool in the planning and prioritisation of
software development activities. This places a limit on the length
and sophistication of the information needs they encapsulate, and
hence a natural boundary on the scope and extent of the patterns
they represent. Their format also allows the analyst to apply
perhaps an unrealistic degree of top-down judgement and iteration
in aligning the relative granularity of the information needs to
existing modes; a benefit that is less readily available to those
whose approach involves real-time, observational data.</p>
          <p>A further caveat is that in order to progress from understanding an
information need to identifying the information behaviors
required to satisfy those needs, it is necessary to speculate on the
behaviours that a user might perform when undertaking a task to
satisfy the need. It may transpire that users actually perform
different behaviours which achieve the same end, or perform the
expected behavior but through a combination of other nested
behaviours, or may simply satisfy the need in a way that had not
been envisaged at all.</p>
          <p>Finally, the process of inferring information behaviour from
selfreported needs can never be wholly deterministic, regardless of
the consistency measures discussed earlier. In this respect, further
steps should be taken to operationalize the application of the
framework and apply some independent measure of stability or
objectivity in its usage.
6. CONCLUSIONS
In this study we have investigated a model of information seeking
behaviour derived from the domain of enterprise search, and
validated its extensibility to users of consumer-oriented websites
and search applications. In so doing, we explored a novel,
goaldriven approach to eliciting user needs, and identified some key
differences in user behaviour between the two domains.</p>
          <p>In addition, we have demonstrated the value of the model as a
framework for expressing complex patterns of behaviour,
extending the IR concept of information-seeking to embrace a
broader range of composite information interaction and use
behaviours. Moreover, we propose that our method can be
adopted by other researchers  who  want  to  take  a  ‘needs  first’ 
approach to understanding information behaviour.</p>
          <p>Revisiting User Information Needs in Aggregated Search</p>
          <p>Shanu Sushmita</p>
          <p>University of California LA
shanusushmita@ucla.edu</p>
          <p>Robert Villa</p>
          <p>University of Sheffield
r.villa@sheffield.ac.uk</p>
          <p>Martin Halvey
Glasgow Caledonian</p>
          <p>University
Martin.Halvey@gcu.ac.uk</p>
          <p>
            Mounia Lalmas
Yahoo! Labs Barcelona
mounia@acm.org
ABSTRACT
Aggregated search interfaces are a common way to present
web search results, mixing di↵ erent types of results into one
single result page. Although numerous e↵ orts have been
made to infer users’ information needs in “standard” search,
we know little about users’ information needs within the
context of aggregated search. This paper presents the outcomes
of a survey of 117 respondents, investigating users’
preferences for their type of search result (image, news, video)
and their type of information need (informational,
navigational and transactional). The survey reveals that users’
result preferences di↵ er based on their underlying information
needs, suggesting that the taxonomy provided by Broder [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ]
requires updating to reflect user information needs in the
context of aggregated search. For instance, respondents
indicated a preference for diverse results (news and reviews
about a particular software product) for navigational and
transactional queries rather than a single result (the web
page to download that software product).
1. INTRODUCTION AND BACKGROUND
Aggregated search is the technique of integrating search
results from di↵ erent verticals (e.g., web, image, video, news)
on a single search result page so that users can access the
increasingly diverse content available on the web.
Aggregated search systems aim to facilitate users’ access to
“nonstandard” web results without having to perform separate
searches in the respective verticals, which are source specific
sub-collections provided by search engines [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ].
          </p>
          <p>
            Throughout the evolution of web search, users’ interaction
with search results has been studied by many to improve the
quality of the search results and the search experience.
Efforts were (and are still being) made to understand users’
information seeking process, based upon which several
taxonomies describing users’ behaviours have been proposed [
            <xref ref-type="bibr" rid="ref1 ref10 ref11 ref13 ref17 ref18 ref21 ref22 ref23 ref28 ref5 ref6 ref9">1,
5, 6, 9, 10, 11, 16</xref>
            ].
          </p>
          <p>
            For instance, in 2002, Broder [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ] created a taxonomy of
Presented at EuroHCIR2012. Copyright c 2012 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
web search, classifying users’ information needs into three
categories, namely, informational, navigational and
transactional. For navigational search, the immediate intent is
to reach a particular site (e.g., BBC Homepage); for
informational search, the intent is to acquire some information
likely to be contained in one or more web pages (e.g., global
warming); and finally, for transactional search, the intent
is to perform some web-mediated activity (e.g., download,
purchase).
          </p>
          <p>
            Others such as Lindley et al. [
            <xref ref-type="bibr" rid="ref28">16</xref>
            ] looked at why
people search or go online and identified five main web
activities: respite, orienting, opportunistic use, purposeful use and
lean-back internet. An example of a respite activity is when
people use the web to take a break at work, or through a
mobile phone to occupy themselves while waiting. Similarly,
Chew et al. [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ] explored the contextual and behavioural
details of users’ interaction with web-based images as they
occur in the course of everyday life, showing that users
interact with image results as these help creating connections
to other people and remote places, or reflecting on the past.
          </p>
          <p>While there is a substantial body of work on
understanding users’ information needs and browsing activities in
“standard” search, far less is known about these within the context
of aggregated search. For instance, it is not clear if the
existing taxonomies on information needs for “standard” search
hold in an aggregated search scenario. In aggregated search,
search results may originate from di↵ erent media (e.g.,
images, maps) or may be of di↵ erent genres (e.g., news, blogs).</p>
          <p>
            This may have an e↵ ect on the way users interact with the
results, and a↵ ect their preferences for the types of results. A
study in [
            <xref ref-type="bibr" rid="ref27">15</xref>
            ] investigated the former, but the latter remains
largely unexplored. For instance, it is not known whether for
navigational queries, users prefer to view a specific website,
as would be implied by [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ]. A negative answer would mean
that a revisit of Broder’s three-main-categories of
information needs is needed. Also, building an awareness of web
activities in aggregated search, which cut across domains,
media types and applications, can highlight important
details when designing for interactions with the web [
            <xref ref-type="bibr" rid="ref28">16</xref>
            ].
          </p>
          <p>
            The focus of this short paper is, therefore, two-fold: (1)
to investigate the preference of search results sought by the
users; and (2) to investigate the existing frameworks of web
activities within the context of aggregated search. For this
purpose, users’ preferences for results of several media types
and genres are investigated. Furthermore, since Broder’s
taxonomy has been heavily used (e.g. [
            <xref ref-type="bibr" rid="ref15 ref19 ref21 ref27 ref3 ref7 ref9">3, 7, 9, 15</xref>
            ]) we focus
on the now classic informational, navigational and
transactional categories. We nonetheless aim to extend this work
with other taxonomies (e.g., ODP1) in future work. This
paper makes the following contributions: (1) Investigates
users’ preference for search results (media and genres) for
informational, navigational and transactional search tasks;
and (2) Provides empirical evidence to support the need for
updating the above three categories within the context of
aggregated search.
          </p>
          <p>
            We present the results of a survey that investigated users’
preferences for results of di↵ erent media types and genres,
as answers to informational, navigational and transactional
queries.
2. STUDY
A survey containing sixteen questions (4 background
questions and 12 search task questions) was distributed on
various social networks. The survey allowed us to reach a large
and diverse enough number of users, and is a common way to
elicit user perceptions and preferences [
            <xref ref-type="bibr" rid="ref16 ref20 ref4 ref8">4, 8</xref>
            ]. A total of 117
respondents completed the survey, of which 60 were female
and 54 male; the remaining 3 did not disclose their gender.
          </p>
          <p>The respondents’ age varied between 20-59 years (mean 29).</p>
          <p>Geographically, respondents were distributed across the US
and Canada (3%), Europe (34%), Asia (62%) and Africa
(1%). Most respondents were familiar with search engines
and used them frequently.
2.1</p>
          <p>Task
The aim of the survey was to elicit users’ preferences for
the types (media, genres) of search results for informational,
navigational and transactional search tasks. To this end,
we designed four search topics2 for each of these three
categories. The list of topics for each category is listed in Table 1.</p>
          <p>In total, there were twelve questions for each respondent to
answer. The orders of the questions were rotated to
minimise ordering bias.</p>
          <p>We designed topics that could be understood universally
(e.g, global warming, checking emails, buying dvd,
software download). Furthermore, the topics were devised to fit
the informational, navigational and transactional categories.</p>
          <p>Therefore, we did not manipulate topics to suit specific
media or genre. For instance, for the topic global warming,
some people may want to read the latest news about global
warming, some others may want to view pictures of melting
icebergs, while some others may want to watch a
documentary on global warming. Therefore this topic does not have
an implicit type intent (e.g. image) but requires the
gathering of information (informational search task) from many
web pages; it is expected that users will look for multiple
results to satisfy the corresponding information need.
However, it will depend on users which result types (image, news,
video, etc) they prefer to view – only news articles, few
pictures, or a combination of both.
2.2</p>
          <p>
            Procedure
For each search topic, the respondents were given five choices,
namely, web, news, image, video and other results3. The
re1http://www.dmoz.org/
2A search topic describes a search task scenario. The
concept of a search task scenario was inspired from [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ].
3The definitions of these categories were not specified in the
instructions and were left open to respondents’
interpretation.
spondents were allowed to select as many options as they
desired. That is, they were allowed to select just ‘one’ or
‘all’ options, and therefore were not forced to provide a
preference for all the choices listed. This allowed a more natural
selection of choices, and hence reduced any design bias. In
cases when the respondents selected more than one option,
they were asked to rank the choices, by providing “1st”, “2nd”
......,“5th” preference for each choice. For instance, if image,
news and others were selected as choices, these had to be
ranked in order of preference (e.g., 1st preference - news,
2nd preference – image, 3rd preference – others).
          </p>
          <p>Figure 1, shows the screenshot of an example question
with the preference options. Next, the outcomes of the
survey are presented.
As the data obtained from the survey was non-parametric,
we report medians and the interquartile range for the
preference scores. The results are reported in Table 2, which shows
the median rank of each vertical by information need.
Friedman tests were performed to estimate the significance of
preference for the results types, among and across the three
categories (navigational, informational and transactional).</p>
          <p>Finally, multiple Wilcoxon-tests were run in the post-hoc
analyses while adjusting the p-values using the Bonferroni
method. The outcomes from the post-hoc pair wise
comparisons for navigational, informational and transactional
categories are shown in Tables 3, 4 and 5 respectively. Each
row in these tables indicates whether a particular result type
was preferred over each of the other result types.</p>
          <p>As can be seen in Table 2, most respondents indicated
the ‘web page’ as the most preferred type of results, when</p>
          <p>Navigational Topics
1. When you wish to book tickets with British Airways, which results would be useful for you?
2. When you wish to find an address from yellow pages, which results would be useful for you?
3. When you wish to check courses of a University, which results would be useful for you?
4. When you wish to check your email (e.g, gmail, hotmail, msn, etc), which results would be useful for you?
Informational Topics
5. When you wish to learn about salsa dance, which results would be useful for you?
6. When you wish to gather information about global warming, which results would be useful for you?
7. When you wish to learn on how to make a pancake, which results would be useful for you?
8. When you wish to know about 2011 budget, and how it e↵ ected farmers, which results would be useful for you?
Transactional Topics
9. When you wish to download a free software, which results would be useful for you?
10. When you wish to download a song for your iTunes library, which results would be useful for you?
11. When you wish to file a property complaint, which results would be useful for you?
12. When you wish to buy a DVD online, which results would be useful for you?
compared to the other four types (image, video, news and
others). The di↵ erence was found to be significant for
navigational, informational, and transactional cases (rows 1-4
in Tables 3, 4 and 5 ); thus suggesting that “standard” web
results are the prime source of information sought by most
users. After web results, news was the second most
preferred type of results when compared to image, video and
others (6th row in Table 2). For the navigational category,
news results were significantly preferred over image, video
and others results (rows 6, 8 and 9 in Table 3). However,
video was equally preferred to news for informational and
transactional categories (row 8 in Tables 4 and 5).</p>
          <p>Finally, there is a trend for image and video results to
come third in preference from respondents for most
categories (4th and 5th rows in Table 2). However, post-hoc
analyses suggest a significant di↵ erence of preference for
video and image over ‘other results’ for all three categories
(rows 7 and 10 in Tables 3, 4 and 5). In addition, video
results were significantly preferred to image results for
informational and transactional cases (row 5 in Tables 4 and
5), while no significant di↵ erence was observed for the
navigational case (row 5 in Table 3 ). Therefore, it is possible
that users may prefer image results instead of video results
in some cases, and video results in other cases. In addition,
image and video being the third preference indicates that
providing image and video results for all queries may not be
appreciated by users.</p>
          <p>In Tables 3 to 5, in only two occasions were the ranking
of result types not significantly di↵ erent: image-video for
navigational, and news-video for informational information
needs. This indicates that for navigational needs, neither
image or video results are judged as important to users,
backing up the results in Table 2, where both are ranked
bottom. For informational information needs, both news
and video were judged equally important to the search tasks,
second only to web (Table 2).</p>
          <p>DISCUSSION
The aim of our study was to investigate, via a survey, users’
results preference for navigational, informational, and
transactional search topics.</p>
          <p>
            Overall, three key observations can be made from this
survey. First, for all query categories, web results continue to
be the prime source of information sought by users – 90%
for navigational, 54% for informational and 85% for
transactional – suggesting that for an aggregated search result
page, web results should always be provided. This echoes
the findings of [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ] where the importance of web results for
aggregated result pages was demonstrated through the
mining of query logs.
          </p>
          <p>Second, there appears to be a di↵ erence between the
result preferences for navigational and transactional queries.</p>
          <p>
            From Broder [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ], the corresponding information needs for
these categories were identified to be focused (i.e., specific
website, download, etc). In contrast, our study suggests that
users also prefer to view other results, and not just one (“to
the point”) result, or one type of result. More precisely, for
the navigational search topics, in addition to web results,
respondents also indicated a preference for news and video
results. This may be due to the fact that, since an
aggregated result page is often provided for most queries by
modern search engines4, users are exposed to diverse results and
as a consequence, results other than web have now gained
prominence. However, whether providing diverse results for
informational and transactional information needs facilitates
task completion, and/or increases user satisfaction, requires
further investigation.
          </p>
          <p>
            Third, users’ preferences for the ‘type’ of results vary with
the query category. For instance, for navigational and
transactional search topics, web and news results seem to be
preferred. The preference is more mixed for informational
search topics, with image results least preferred. In itself,
it is not surprising that users’ preferences vary with query
categories. However, concrete knowledge regarding which
‘types’ of sought results are preferred would allow for more
appropriate aggregation of the di↵ erent verticals under
consideration. Similar investigations were carried out in [
            <xref ref-type="bibr" rid="ref12 ref24">12</xref>
            ]
by Sushmita et al. where, associations between query
classifications (e.g., arts, health, etc) and result types were
indeed identified. Such knowledge may then be used by search
systems, to present particular types of result for di↵ erent
queries, for example, a system may not present (or demote
in importance) image results in response to an informational
query.
          </p>
          <p>CONCLUSION AND FUTURE WORK
4http://www.slideshare.net/rankabove/com-scorerankabove-final</p>
          <p>We presented the analysis of a survey of 117 respondents’
preferences regarding the di↵ erent types of results for
navigational, informational, and transactional information needs.</p>
          <p>
            Although small in terms of the number of users and
acknowledging the limitation of an online survey, interesting insights
emerged from our investigation. The outcomes of the
survey support the aggregated search paradigm, showing that
users’ preferences are for a diverse range of result types. The
analysis also indicates a need to revisit the definition of the
three categories of information needs [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ], within the context
of aggregated search. This work initiates two future research
questions: (1) What information needs exist within the
context of aggregated search? and (2) How to identify suitable
results satisfying those information needs?
          </p>
          <p>Figure 4a. Visual aids
(rectangle: action area, here:
“not like”, circle: gaze cursor).</p>
          <p>Figure 4b. No visual aids.
4. EVALUATION DESIGN
In order to evaluate EyeGrab, 24 subjects (7 female) played the
game. The subjects’ age was between 15 and 32 years (mean: 24,
SD: 3.9). 19 subjects were students, 2 research assistant, one pupil
and 2 had other professions. Most of the players had experience in
gaming (mean: 3.5, SD: 1.31). Only a few were familiar with eye
tracking, this is indicated by 19 subjects rating the question
concerning their eye tracking experience with one (mean: 1.63,
SD: 1.38). The subjects were randomly divided into two groups
A and B. Group A had no visual aids during the game, whereas
group B did. 8 users were wearer of glasses or contact lenses (4 in
each group). There were no problems using the eye tracker for
those subjects.</p>
          <p>To avoid distractions, the game was played in our eye-tracking
lab providing a chair, a desk with an eye tracker, and a standard
monitor. The first step was a calibration of the eye tracker. After
this was done, the game was started without further instructions
and was played with 30 images in each of the three rounds. The
data from the first round is not used in the later analysis, because
it has only served for getting the subjects acquainted with the
usage of the eye tracker as input device. At the end of the
experiment, every user filled out a questionaire, including
personal information and questions about the performance of the
game. The answers were given on a 5-point Likert scale.
5. EVALUATION RESULTS
5.1 Satisfaction
The questionnaires show that the subjects enjoyed playing the
game. On average, the statement “It was fun playing the game.” is
rated 3.46 (SD: 0.93) considering all 24 subjects. 20 of the 24
users agreed to this statement. One of the following questions was
if the participants felt like the interaction with the eye tracker
increases the fun of the game. 14 subjects agreed or strongly
agreed to this statement (mean: 3.5, SD: 1.25). Also, most of the
subjects did not feel disturbed by the eye tracker (mean: 2.25,
SD: 1.5).
5.2 Effectiveness and Efficiency
One round of the game comprises 30 images and it takes about
two minutes including the introduction and the input form. Each
level has a different pace at which the images fall down the
screen. Thus, the classification per image takes between 2.6 and 4
seconds.</p>
          <p>In total, 1440 pictures were shown to the subjects within the game
(30 pictures per category “house” and “mouse” times 24 subjects).</p>
          <p>Only in 42 cases, the image passed without classification,
resulting in a total of 1398 classified images. 1162 images were
correctly classified (83%). Thus, only 236 images were
incorrectly classified. Overall we had 897 true-positive
classifications, 128 false-negative and 108 false-positive
classification, which leads to a precision of 89% and a recall of
88% over all users. For the group with the better results (the
group without visual aid, see next section) we obtain a precision
of 92%.
5.3 Visual aid
The subjective perception of the users in group B (the group that
was provided with visual aids) was that the visual aids supported
them in the classification tasks. The question regarding the visual
highlighting of the active areas was rated as very helpful with an
average of 4.67 (SD: 0.49). The subjects also answered that
displaying the gaze point was very helpful and scored this
question on average with 4.5 (SD: 0.67). However, to our
surprise, the following statistical analysis of the data shows that
group B with the visual aids misclassified significantly more
images than group A did.</p>
          <p>Group A correctly classified 296 images for category “house”
whereas group B correctly assigned 264 images for this category.</p>
          <p>Regarding the category “mouse”, in total 317 correct assignments
were made by group A whereas group B correctly assigned 287
images. Regarding the misclassified images, group A
misclassified 59 images for category “house”, whereas group B
wrongly assigned the image category in 81 cases. For the
category “mouse”, the number of incorrect assignments is 37 for
group A and 57 for group B. We compared the values for correct
and incorrect assignments for group A and B in a 2x2 Chi-square
test for both categories. The differences are significant regarding
a significance level of = 0.05 with 2 (1, N = 700) = 5.14, p =
0.023 for the category “house” and 2 (1, N = 698) = 5.6, p =
0.018 for “mouse”. In group B 31, images passed without
classification, in group A only 11 images were not classified.</p>
          <p>These results indicate that the visual support is not improving the
classification. Despite the good impression of the visual support
that group B expressed, the following question might be an
indicator that this group felt less comfortable with the eye
trackerbased interaction than group A did: we asked the subjects to state
if they preferred a mouse-based interaction instead of the eye
tracker-based one. On average, subjects of group A scored this
question with 2.17 (SD: 1.47) and group B with 3.25 (SD: 1.48).</p>
          <p>Using a two-tailed Mann-Whitney U-Test, a weakly significant
difference was determined stating a preference of group B over
group A to use the mouse to play EyeGrab (U = 43, z = -1.719, n1
= n2 = 12, p = .085).
6. FUTURE WORK
For the current version of our EyeGrab game, we have used
preclassified images in order to verify the classification performance
of the subjects. We plan to use images without annotations in
future extensions of the game.</p>
          <p>Also the detailed analysis of the gaze information will be
performed in a next step. In a small sample of 5 images classified
by one user, we received 231 gaze points on the images. An
example of a gaze path visualization is shown in Figure 5. We
expect a sufficient number of fixations and correct classification
to allow a detailed analysis.</p>
          <p>
            We received 897 ratings for the shown images. 556 of them were
positive. The quality of these ratings has to be investigated in a
future experiment, e.g., by repeating the ratings in another context
with the same users or by using a ground truth set with images,
often liked by a big number of other users. However it has to be
clear, that a subjective rating can never be “correct” or not. These
investigations can only provide an indication of the worth of the
rating. Overall, this detailed analysis will allow us to identify the
regions that correspond to the category given in the EyeGrab
game. Such region-based annotations will allow for a better
retrieval of the images in the future.
7. SUMMARY
We have introduced the gaze-based game with a purpose
EyeGrab to classify images using an eye tracker. We have shown
that the game has the potential to entertain the players and that the
classification results are good enough to advance beyond the gaze
analysis. This analysis is the first step in the direction of
extending image context information with information gained in
an eye tracking game. The next step will be the analysis and
evaluation of the gained information and to use it for improving
image search tasks.
8. REFERENCES
[
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ] Hornof, A.J. and Cavender, A. 2005. EyeDraw: enabling
children with severe motor impairments to draw with their
eyes. Proceedings of the SIGCHI conference on Human
factors in computing systems, 170.
[
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ]
          </p>
          <p>Walber, T. and Scherp, A. and Staab, S. 2012. Identifying
Objects in Images from Analyzing the Users’ Gaze
Movements for Provided Tags. Advances in Multimedia</p>
          <p>
            Modeling, 138-148.
[
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ] Carson, C., Thomas, M., Belongie, S., Hellerstein, J. and
          </p>
          <p>Malik, J. 1999. Blobworld: A system for region-based image
indexing and retrieval. Visual Information and Information</p>
          <p>Systems.
Using Wordclouds to Navigate and Summarize Twitter Search</p>
          <p>Results
Rianne Kaptein</p>
          <p>Oxyme
Amsterdam, The Netherlands
This paper describes an application in which wordclouds are used to
navigate and summarize Twitter search results. A search on Twitter
can return thousands of relevant tweets. By just looking at the first
few result pages you will not get an overview of what is discussed
in all search results. Our application summarizes sets of tweets into
wordclouds, which can be used to get a first idea of the contents of
the tweets. Also the application provides the option to zoom in on
a certain part of the search results to inspect them in more detail.</p>
          <p>The application has not been formally evaluated, but we do provide
some insights and points for discussion.
1</p>
          <p>Introduction
One of the most common problems in Information Retrieval is
information overload: there is simply too much relevant information
available for the users to process. Therefore applications are needed
to help users deal with large amounts of data. In this paper we
describe an application which was developed for this purpose. The
use of wordclouds in the application serves two purposes:
1. To summarize
2. To aid navigation
This application was developed with the following two user
scenarios in mind:
1. General Twitter search</p>
          <p>Nowadays many people express their opinions about
products, services and companies on Twitter. When you want to
get a broad overview of what people are tweeting in general
about a company or event, it does not suffice to read the first
few pages of search results. You want to get a feeling for the
most frequently discussed topics overall, and dive into
particular subtopics of special interest, such as product
recommendations.</p>
          <p>Presented at EuroHCIR2012. Copyright c 2012 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
2. Searching fragments of categorized data</p>
          <p>Besides Twitter there are many more places on the Web where
people express their opinions. These opinions can be
collected and annotated with labels such as sentiment, source,
market etcetera. When you have a large amount of annotated
data available, it is interesting to see for example what are
the different topics discussed in positive and in negative
messages.</p>
          <p>In this paper we will focus on the first user scenario: General
Twitter search, since Twitter data is abundant and publicly available.</p>
          <p>Humans have a great capacity to notice terms which are out of the
ordinary. When looking at a wordcloud there will always be some
unexpected terms which catch your attention and are good pointers
for further investigation. In tweets about public transport you can
expect for example tweets about delays, but you might not expect
certain tweets about recent events such as a new colour of the trains.</p>
          <p>What we try to do in the wordclouds is to emphasize the words that
are noteworthy from a statistical point of view, and leave it up to
the user to decide which messages to explore further.</p>
          <p>
            Although the usefulness of tagclouds for navigation is still a topic of
debate [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ], exploratory applications which make use of wordclouds
for summarization and navigation of search results have been
moderately successful on specific domains such as web documents [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ]
and PubMed publications in biomedical literature [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ].
          </p>
          <p>The search results that we are investigating in this paper have three
characteristics:
• A search result is a short textual message. By design a Twitter</p>
          <p>message cannot contain more than 140 characters.
• The number of search results is large. If this would not be the
case, since the results are short texts, you could simply read
through all of them.
• There are many, equally relevant search results. In web search
there are usually not more than a handful highly relevant
search results. Many of the search results contain copied or
redundant information, or only mention the search words
occasionally. Although Twitter search results also contain
redundant information, i.e. repeated tweets and retweets, the
set of relevant tweets can still consist of thousands of equally
highly relevant tweets.</p>
          <p>In the next sections of this paper we will present our approach
(Section 2), a case study (Section 3), and finally our conclusions
(Section 4).
The application consists of two screens. The first screen handles the
input, the second screen displays the results based on your input.</p>
          <p>On the first screen the system offers a number of selections that
can be made to make sure you generate the wordclouds that are
best representing your data and your analysis purpose. The input
is collected using textfields, radiobuttons and checkboxes. The first
part of the inputscreen is shown in Figure 1.</p>
          <p>The following selections can be made:
• File selection, a tab separated text file is required as input.
• Text selection, which column in the dataset to use as textual</p>
          <p>input for the wordcloud generation.
• Category selection, based on a value in any column of your
dataset your data can be categorized. It is also possible to
create categories based on the presence of words in the contents
of your data, e.g. to create a category for all tweets containing
the term ‘happy’.
• Language, used for the removal of standard stopwords.
• Optionally, additional stopwords can be specified. These</p>
          <p>words will not occur in any of the wordclouds.
• Stemming, currently available only for English. The Krovetz
stemmer is used, because this stemmer always stems words
into other valid English words.
• Exclude numbers, when your data includes many numbers
such as product prices it can be desirable to exclude these
numbers from the wordcloud.
• Exclude retweets / repeated posts, when your data contains
a tweet that is retweeted very frequently, this one tweet will
dominate the wordcloud which can be undesirable.
• Include only usernames, for Twitter data only, keep only the</p>
          <p>usernames, i.e. all the words starting with @.
• Include only hashtags, for Twitter data only, i.e. all the words</p>
          <p>starting with #.</p>
          <p>The second screen shows the output, which consists of wordclouds
for the categories you have specified, as well as a wordcloud for all
the search results.</p>
          <p>
            Wordclouds for categories are generated using a parsimonious
language model. This model compares the frequency of words in a
set of documents to the average term probability in a background
collection containing similar documents to extract the most
noteworthy terms. In this case the background collection are all the
retrieved search results. Terms that are only mentioned occasionally
in the set of documents and terms which have a similar or higher
probability of occurrence in the background collection will not be
included in the parsimonious language model [
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ].
          </p>
          <p>
            The parsimonious language model [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ] is an extension to the
standard language model based on maximum likelihood estimation, and
is created using an Expectation-Maximization algorithm.
Maximum likelihood estimation is used to make an initial estimate of
the probabilities of words occurring in the set of documents.
          </p>
          <p>(1)
(2)
Pmle(ti|S) =
t f (ti, S)
Â t t f (t, S)
where S is the set of documents, and t f (t, S) is the text frequency,
i.e. the number of occurrences of term t in set of documents
S. Subsequently, parsimonious probabilities are estimated using
Expectation-Maximisation:</p>
          <p>E-step:
M-step:</p>
          <p>(1 l )P(t|S)
et = t f (t, S) · (1 l )P(t|S) + l P(t|C)</p>
          <p>e
Ppars(t|S) = t , i.e. normalize</p>
          <p>Â t et
where C is the background collection model. In the initial E-step,
maximum likelihood estimates are used for P(t|S). We set the
smoothing parameter l to 0.9. In the M-step the words that
receive a probability below a threshold of 0.001 are removed from
the model. The iteration process stops after a fixed number of
iterations.</p>
          <p>In the next section we present a case in which the generated output
of the application is presented.
3</p>
          <p>Case
Using an example search we will demonstrate how we use
wordclouds in our application to navigate and summarize the search
results. We executed a search on Twitter using the Twitter search
API1 for the query ‘#london2012’ over the last 5 days, saved all
the 30,504 search results in a .csv file and load this file into our
application. Looking at the wordcloud over all the results that is
shown in Figure 2, we see the term ‘torch’ is frequently used, and
we zoom in on this aspect of the ‘#london2012’ search. By
clicking on the word ‘torch’ a list of messages is shown that all contain
the term ‘torch’, so these messages can be inspected in more detail.</p>
          <p>This list of messages is still quite long however, consisting of 1,046
tweets. We can zoom in further on these tweets by going back to
the input screen and specifying ‘torch’ as a category. Now, a
parsimonious wordcloud is created from the 1,046 tweets that contain
the term ‘torch’. The resulting wordcloud is shown in Figure 3.</p>
          <p>The figure is a screenshot of the screen that is displayed when the
word ‘Sheffield’ is clicked, showing the tweets containing the word
‘Sheffield’.</p>
          <p>Words which occur frequently in all of the ‘#london2012’
messages, such as ‘#london2012’, ‘2012’, and ‘olympics’, receive a
lower score from the parsimonious model, and almost none of these</p>
          <p>1https://dev.twitter.com/docs/api/1/get/search
words occur in the ‘torch’ wordcloud. Also general words that
occur frequently in all of the messages, such as ‘get’, and ‘will’ are
filtered out. Instead the cloud contains words that occur more
frequently in the subset of messages that contain the word ‘torch’, for
example some of the cities that the torch passes through such as
Sheffield, Leeds and Manchester. Every result in this cloud by
definition contains the word ‘torch’, therefore it takes a prominent place
in the wordcloud. You can choose to not show the word ‘torch’ in
the wordcloud by specifying it as a stopword on the input screen.</p>
          <p>Clicking on a term in the wordcloud has the same effect as query
expansion, i.e. adding that term to your query and retrieve another
set of results. When you use the Twitter API to search Twitter
without query operators, only results will be returned that contain all of
the search terms in the Tweet, username or hyperlink. This means
adding a term to your query will not lead to more search results.</p>
          <p>Only if you remove the original query terms, other results will be
returned.</p>
          <p>Observations
We have not had the chance to evaluate our application through
means of a user study. However, we do want to point out the
following observations. Given the nature of our data, i.e. a collection
of tweets, there might be some improvements possible that exploit
this particular type of data. Tweets can contain special elements
in the text, namely usernames, hashtags, links, and emoticons. We
make the following observations:
• Usernames and hashtags are currently considered in the sense
that we remove all punctuation except the characters ‘@’ and
‘#’ which are the indicators of usernames and hashtags
respectively. There is an option to generate wordclouds containing
only usernames, or only hashtags. In the default settings
usernames and hashtags are included as is in the wordcloud. For
future work we want to discuss and investigate two open
issues:
1. Can a word with a hashtag be considered as the same
word without the hashtag? While a hashtag term does
not always have to be a real word, e.g. #london2012,
in many cases it is, e.g. #london. For the wordcloud
should the terms ‘london’ and ‘#london’ be merged?
Sometimes usernames are used in a similar way as
hashtags to address companies, e.g. in this tweet:
‘Ambush marketing at the Olympics! Well played, @Nike.</p>
          <p>bit.ly/N4zAUc #London2012’.
2. A related issue is the importance or term weights of
usernames and hashtags. Is a hashtag a stronger signal,
and should it therefore be featured more prominently in
the wordcloud? Similarly for usernames, but usernames
could also be considered a weaker signal, so should they
be featured less prominently?
Both of these questions can also be considered when you want
to optimize a retrieval algorithm.
• Besides the ‘@’, and ‘#’ all other punctuation is removed
during text preprocessing. This means all emoticons like ‘:)’
are removed. Sometimes these emoticons are used as
indicators of sentiment, i.e. tweets containing ‘:)’ are classified
as positive messages, and tweets containing ‘:(’ as negative
messages. In this sense the emoticons do indeed represent
valuable information that could be included in the wordcloud.</p>
          <p>When an emoticon appears in the wordcloud, clicking on it
can give you all the messages associated with for example a
positive emoticon.</p>
          <p>Feedback from users is required to determine the most useful
improvements for the application.
4</p>
          <p>
            Conclusions
results. Wordclouds are a quick way to summarize and get a first
overview of large amounts of data. Using human observation skills
it is easy to zoom in on a group of messages in which you are
interested, i.e. all messages that contain a specific term from the
wordcloud. In future work we would like to evaluate the usefulness of
wordclouds for navigation and summarization of search results in a
user study.
5
[
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ] B. Y.-L. Kuo, T. Hentrich, B. M. Good, and M. D. Wilkinson.
          </p>
          <p>Tag clouds for summarizing web search results. Proceedings
of the 16th international conference on World Wide Web WWW
07, 196:1203, 2007.</p>
          <p>Do users benefit from
controlled vocabularies in search interfaces?</p>
          <p>Ying-Hsang Liu
School of Information Studies</p>
          <p>Charles Sturt University
Wagga Wagga, NSW 2678,</p>
          <p>Australia
yingliu@csu.edu.au</p>
          <p>Paul Thomas</p>
          <p>CSIRO</p>
          <p>GPO Box 664
Canberra, ACT 2601, Australia
paul.thomas@csiro.au</p>
          <p>Tom Gedeon
Research School of Computer</p>
          <p>Science
Australian National University
Canberra, ACT 2601, Australia
tom@cs.anu.edu.au</p>
          <p>Jan-Felix Schmakeit
Research School of Computer</p>
          <p>Science
Australian National University
Canberra, ACT 2601, Australia</p>
          <p>janfelix.schmakeit@anu.edu.au
ABSTRACT
Search providers in domains from medicine to news have
long labelled documents with controlled vocabularies, to
help users explore their collections. These vocabularies are
expensive to build and use, however, and seem to be useful
mostly for domain experts.</p>
          <p>This paper describes an on-going gaze-tracking study which
asks whether users notice controlled vocabularies when they
are exposed in a search interface; whether they make use of
them; and whether this improves search. We also hope to
learn what e↵ ect several standard search interfaces have on
the use of controlled vocabularies.</p>
          <p>Categories and Subject Descriptors
H.3.3 [Information Search and Retrieval]: Search
process; H.5.2 [User Interfaces]: User-centered
design—performance measures
General Terms
Experimentation, Human Factors
Keywords
Search results presentation, individual di↵ erences, gaze
behaviour, MeSH terms
1. INTRODUCTION</p>
          <p>
            It has been recognised that people engage with di↵
erent kinds of searching behaviours, but current information
Presented at EuroHCIR2012. Copyright c 2012 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
retrieval (IR) systems are primarily designed for specified
search [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ]. The simple search box is still the dominant
interaction mode in modern search engines. However, a user-centred
approach to interface design that takes into account
individual di↵ erences, search goals and tasks, has the potential to
support users interacting with IR systems more e ciently
and e↵ ectively.
          </p>
          <p>
            To this end researchers have advocated “natural” search
user interfaces, arguing they are easier to use and require less
user training [e.g. 9, 20]. It is however challenging to design
natural interfaces because of the complexity of information
problems and associated searching behaviours. For instance,
user studies have demonstrated that user queries are typically
very short representations of complex information needs [
            <xref ref-type="bibr" rid="ref11 ref15 ref23 ref3">3,
11</xref>
            ], and users have di culty formulating queries to represent
information problems. User interaction with IR systems is
inherently interactive and exploratory [e.g. 2, 17], so usable
interfaces for query formulation are important in support
of natural search interactions. (See Wilson [24] for a recent
comprehensive review of search interfaces, and Wacholder [22]
for a review of interactive query formulation.)
          </p>
          <p>One way to support query formulation is with a controlled
indexing language, where each document is assigned terms
from an predefined list or hierarchy of indexing terms.
Examples include Medical Subject Headings (MeSH) terms and
Library of Congress Subject Headings (LCSH). The
usefulness of MeSH terms in biomedical searching is especially
important because of the extreme popularity of the PubMed
database1, the publicly accessible version of MEDLINE on
the web.</p>
          <p>
            Controlled vocabularies are expensive to build, use, and
maintain, and they may contribute to clutter in a search
interface. There is some evidence that domain experts benefit
from controlled vocabularies, but results have been mixed
for ordinary users (e.g., [
            <xref ref-type="bibr" rid="ref10 ref22 ref27">10, 15, 19</xref>
            ]). Given these costs, and
the unclear benefits for most searchers, we are interested in
whether and how users make use of controlled vocabularies
when they are available.
          </p>
          <p>1http://www.ncbi.nlm.nih.gov/pubmed</p>
          <p>This paper describes an on-going eye-tracking study of
user gaze and search behaviours searching clinical search
topics, with particular reference to the user’s attention to
and use of the document surrogates (i.e., MeSH terms, title
and abstract). The specific research questions are:
1. What components of document surrogates do searchers
look at when reformulating their queries? Do searchers
even notice MeSH terms in standard search interfaces?
2. If they do notice them, how do searchers use the
dis</p>
          <p>played MeSH terms in their search processes?
3. If they are used at all, do MeSH terms lead to better</p>
          <p>search performance and e ciency?</p>
          <p>RELATED WORK</p>
          <p>
            Past work has considered system designs to support query
reformulation. From a system perspective, researchers have
proposed visualizing document inter-relationships [21],
explicit term distribution information [
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ] and search interfaces
in support of search results navigation [
            <xref ref-type="bibr" rid="ref30">18</xref>
            ] to help users
refine their queries. From a user perspective, research has
revealed that searchers prefer to use such search interfaces
for reformulating their queries and to have some degree of
control over the search process [e.g. 12, 8, 13, 23]. In a recent
study of search interfaces in support of interactive query
expansion [
            <xref ref-type="bibr" rid="ref19 ref7">7</xref>
            ], it was found that displaying expanded terms and
corresponding changes in summaries of search results was
useful for the decision-making process in query reformulation;
particularly for di cult search topics. However, it is still
unclear whether users pay attention to these system features,
and whether the use of these features contributes to better
search performance and e ciency.
          </p>
          <p>
            Recent HCI and IR research has focused on users’ cognitive
aspects in search interactions by measuring the gaze patterns,
an indicator of searcher attention (see e.g. Dumais et al. [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ] or
Logio et al. [
            <xref ref-type="bibr" rid="ref28">16</xref>
            ]). The use of eye-tracking equipment for
capturing searchers’ fixation patterns provides a rich set of data
to understand whether searchers read document surrogates
(e.g. summary and metadata) and more importantly, how
searchers attend to di↵ erent components of search results or
search interfaces [
            <xref ref-type="bibr" rid="ref16 ref26 ref4">4, 14</xref>
            ]. We are adopting similar techniques
in our study.
          </p>
          <p>We are conducting a user experiment to assess the e↵ ect of
displayed MeSH terms on search behaviors and performance.</p>
          <p>The search task is to perform searches on clinical
information for other patients, and find the best query to obtain
as many relevant documents as possible. Our recruits are
undergraduate and postgraduate students with search engine
experience but without advanced academic background in
the biomedical domain. Each user searches 8 topics in total,
with a 7-minute limit for each topic, and the experiment
takes about 90 minutes in total.</p>
          <p>Participants are given brief instructions about the search
task and system features, followed by a practice topic and
then the searches proper. User interaction data is recorded:
we are noting all queries, mouse clicks, retrieved documents,
time spent, and eye movements. Electroencephalogram
(EEG) readings are also captured.
(a) Screenshot of Interface “B”, suggestions per-query and
displayed at top.
(b) Screenshot of Interface “D”, suggestions per-document and
displayed with the document.</p>
          <p>Participants search on four di↵ erent search interfaces
using a single search system. The four search interfaces are
distinguished by whether the MeSH terms are presented and
how the displayed MeSH terms are generated:
Interface “A” mimics web search and other search systems
with no controlled vocabulary. This interface has a
brief task description at top; a conventional search box
and button; and each result is represented with its
title, authors, publication details, and abstract where
available.</p>
          <p>Full text is not available, so the results are not
clickable. Users must judge their success on the titles and
abstracts alone.</p>
          <p>Interface “B” (Figure 1(a)) adds MeSH terms to the
interface. After the user’s query is run, MeSH terms from all
results are collated; the most frequent ten are displayed
at the top of the screen. This mimics the per-query
suggestions produced by systems like ProQuest2.</p>
          <p>MeSH terms are introduced with “Try:” and are
clickable: if a user clicks a term, their query is refined to
2For example, see http://www.proquest.co.uk/en-UK/
products/brands/pl_pq.shtml
Imagine that you are 63-year-old male with acute renal
failure probably 2nd to aminoglycosides/contrast dye.</p>
          <p>You would like to find information about acute tubular
necrosis due to aminoglycosides, contrast dye, outcome
and treatment.</p>
          <p>include the MeSH term and then re-run. We hope
that the label, and the fact they work as links, will
encourage users to interact with them.</p>
          <p>Interface “C” uses the same MeSH terms as “B” but
displays them alongside each document, where they may
be more (or less) visible. It is a hybrid of interfaces “B”
and “D”.</p>
          <p>Interface “D” mimics EBSCOhost3 and similar systems
that provide indexing terms alongside each document.</p>
          <p>As well as the standard elements from interface “A”,
interface “D” displays the MeSH terms associated with
each document, as part of that document’s surrogate
(Figure 1(b)).</p>
          <p>Again, terms are introduced with “Try:” and are
clickable.</p>
          <p>Each interface is labelled with a simple figure—a square,
circle, diamond, or triangle—which we refer to in our exit
questionnaire.
3.2</p>
          <p>Design</p>
          <p>
            This experiment is a 4 ⇥ 4 factorial design with four search
interfaces and four topic pairs. We are using a 4 ⇥ 4
GraecoLatin square design [
            <xref ref-type="bibr" rid="ref18 ref6">6</xref>
            ] to arrange the experimental conditions.
          </p>
          <p>We expect to enroll 32 participants from the campus of a
large university, which will give good statistical power (when
N = 32, ANOVA &lt; 0.01 for “medium” e↵ ect of = 0.75).</p>
          <p>Entry and exit questionnaires are collecting demographic
information and information on participants’ cognitive styles
and their perception of the search process. We also ask
participants’ opinions of the tasks and the interfaces.
3.3</p>
          <p>Topics</p>
          <p>
            Search topics used here are a subset of the clinical topics
from OHSUMED [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ], originally created for batch-mode IR
system evaluation. We have re-written the topics slightly so
they read as instructions to our participants (see Figure 2
for an example).
          </p>
          <p>We selected topics to cover a range of di culties: we
sorted the topics according to the number of judged relevant
documents and selected two topics, at random, from each
quartile. These eight topics were then randomly paired o↵
to produce four pairs of topics. A final topic, the same for
all participants, is used for training.
3.4</p>
          <p>Software and hardware</p>
          <p>The search system is built on Solr4, with the search results
ranked by default relevance score. The MeSH terms are not
specifically weighted.
3http://www.ebscohost.com/
4http://lucene.apache.org/solr/</p>
          <p>Gaze tracking uses FaceLab5 software and hardware. We
use Eyeworks software6 for recording and basic analysis.</p>
          <p>EEG data is recorded with an Emotiv headset7.</p>
          <p>Where do people look?. Recordings will be analysed to
see how often there are fixations in di↵ erent parts of
document surrogates, and therefore how often people have looked
at each part. In particular, for interfaces B, C and D we
will consider how often participants look at the controlled
vocabularies (“Try:. . . ”). Any e↵ ect on gaze patterns due
to interface would tell us which interfaces make the extra
information easiest to discover.</p>
          <p>Our exit questionnaire also asks whether users noticed the
controlled vocabularies: we would not be surprised if there
were di↵ erences between the self-reported data and the gaze
data, for example if participants were trying to please us.</p>
          <p>Do they use the controlled vocabulary?. Our software
records all clicks on terms from the controlled vocabulary,
so it will be easy to note how often it is used and whether
there is any correlation with interface, task, sequence, or
user. Again, an e↵ ect due to interface would suggest which
style of interface makes features like the controlled terms
most attractive.</p>
          <p>Participants who merely read and re-type the controlled
vocabulary may be picked up in query logs.</p>
          <p>Again, we intend comparing these recordings with
selfreports.</p>
          <p>If so, does it help?. Assuming some participants do make
use of the MeSH terms, we anticipate four ways to address
this question. First, as before, we will consider self-reports
of task di culty to see whether these correlate with the use
of controlled vocabulary features. Second, since participants’
final queries on each topic should be the ones the like best,
we can check how many of these use MeSH terms. Third,
the judgements associated with OSHUMED topics will allow
us to measure the actual e↵ ectiveness of queries with and
without controlled terms. Finally, if participants do not use
all their allocated time for each task, variations in completion
time may be interesting.</p>
          <p>FIRST RESULTS AND NEXT STEPS</p>
          <p>We have conducted a small-scale pilot to test our design
and instruments.</p>
          <p>Our participants did glance at MeSH terms: 8% of fixations
were on MeSH terms in interfaces B to D, which compares to
6% on document titles and 12% on abstracts. However, they
were very seldom used – only one query, of 44 queries issued
on these interfaces, used any MeSH terms at all. There are
also some indications of a per-interface e↵ ect, with the MeSH
terms at the top of interface D receiving little attention
We will shortly be recruiting for the full-scale experiment.</p>
          <p>We hope this will o↵ er some insight into the relationship
5http://www.seeingmachines.com/product/facelab/
6http://www.eyetracking.com/Software/EyeWorks
7http://www.emotiv.com/
between interface, reading patterns, search behaviour, and
search e↵ ectiveness.</p>
          <p>ACKNOWLEDGMENTS</p>
          <p>Ying-Hsang Liu has been supported by the School of
Information Studies Research Fellowship from Charles Sturt
University and working as Visiting Fellow at Research School
of Computer Science, The Australian National University.
User-Centred Design to Support Exploration and Path</p>
          <p>Creation in Cultural Heritage Collections
1
1
1
1
1
Paula Goodale, Paul Clough, Nigel Ford, Mark Hall, Mark Stevenson, Samuel Fernando,
1Nikolaos Aletras, 2Kate Fernie, 3Phil Archer, 4Andrea de Polo
1</p>
          <p>1University of Sheffield, Sheffield, United Kingdom
p.goodale│p.d.clough│n.ford│m.mhall│m.stevenson│s.fernando│naletras1@sheffield.ac.uk
2MDR Partners, United Kingdom; 3iSieve Technologies, Greece; 4Alinari 24 ORE Florence, Italy
2kate.fernie@mdrpartners.com, 3phil@philarcher.org, 4andrea@alinari.it
ABSTRACT
In this paper, we present the results of the user requirements and
interface design phase for a prototype system, designed to
enhance interaction with cultural heritage collections online
through means of a pathway metaphor. We present a single user
interaction model that supports various work and information
seeking tasks undertaken by both expert and non-expert users
within the context of collection exploration and path creation. The
user interaction model is shown to enable seamless movement
between interaction modes, with the potential over time to
encourage deeper engagement and learning.</p>
          <p>
            Categories and Subject Descriptors
H.5.m [Miscellaneous]: Interaction framework
General Terms
Design, Human Factors, Theory.
1. INTRODUCTION
Large-scale projects for the digitisation of cultural heritage (CH)
have become commonplace in recent years, and yet complex
issues arise with regard to information access. Specialist metadata
and the often variable quantity and quality of object descriptions
make it difficult for users to navigate vast, structured and often
very scholarly collections. It is therefore difficult to locate
resources of interest, especially for those without advanced levels
of subject and domain knowledge [
            <xref ref-type="bibr" rid="ref18 ref6">6</xref>
            ]. User experience online is
thus far removed from that of visiting a museum or gallery in
person, where guidance through a much smaller selection of
carefully curated objects is the norm, for example, via the medium
of visitor-friendly object labels, guide books, audio tours and
activity trails. Exhibit information is designed for general rather
than academic audiences, with additional materials tailored for
family groups and learners, amongst others.
          </p>
          <p>
            Guided tours and activity trails are commonplace offerings to aid
visitor orientation at physical cultural heritage sites, and offer a
Presented at EuroHCIR2012. Copyright © 2012 for the individual
papers by the papers' authors. Copying permitted only for private and
academic purposes. This volume is published and copyrighted by its
editors.
range of opportunities for immersive and more highly engaged
visitor experiences [
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ], often utilising technological solutions, and
even extending to the latest mobile devices [
            <xref ref-type="bibr" rid="ref27">15</xref>
            ]. They are though
much less in evidence online, despite the fact that the idea of
documents or other items linked together in the form of hypertext
trails is considerably older than the web itself [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ].
          </p>
          <p>
            Online paths and trails are seen as a means of aiding navigation,
exploration and learning [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ] in general and educational online
environments, and there are many examples of research [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ] [
            <xref ref-type="bibr" rid="ref12 ref24">12</xref>
            ]
[
            <xref ref-type="bibr" rid="ref25">13</xref>
            ] and commercial activity in offering tools to develop paths
from web pages (e.g. www.trailmeme.com) and social media
content (e.g. www.storify.com). However, very few examples are
domain-specific and/or pertain to digital library collections, and in
consequence, it is rare for all of the associated exploration,
authoring and use activities to be integrated within the same
space. Through our current research we therefore aim to exploit
opportunities to utilize paths to support diverse groups of users in
the complete cycle of information seeking, exploration, path
creation and interaction within CH digital collections, opening up
their use to more widespread educational and leisure audiences.
2. RELATED WORK
Research on information user behavior in CH digital collections is
scarce, especially when considering the needs on non-expert
users, i.e. those without detailed subject and domain knowledge.
          </p>
          <p>
            Expert users regularly engage in both simple fact-finding and
more complex information gathering tasks, amongst others, with
the latter having multiple variations and components such as topic
searches, exploration, collecting/combining [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ], all of which are
relevant to our current study. Similarly, non-expert users [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ] also
engage in known-item searching and exploration. Visual
representations of artefacts are highly important in this context,
and the process of meaning-making through contextual
information and the derivation of personal inferences and
connections is also strongly evidenced [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ].
          </p>
          <p>
            For known-item or fact-finding searches, some knowledge of the
metadata and collection structure is imperative, but such
knowledge is much less likely to be used effectively, if at all by
non-expert users than expert users [
            <xref ref-type="bibr" rid="ref19 ref7">7</xref>
            ]. In addition, information
retrieval tools in CH collections, and the web more generally, are
much less likely to effectively support the needs of users more in
more open-ended exploratory tasks.
          </p>
          <p>
            Exploratory search extends the idea of basic lookup into the areas
of learning and investigation, which in turn incorporate extended
information processing, evaluation and annotation [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ]. Aligned
with these variations of exploratory search are the concepts of
serendipity [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ], where the user encounters information that they
were not actively looking for, and berry-picking [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ], which is an
extended, iterative and adaptive search process that also
incorporates the idea of collecting information objects as the
search progresses over time. Solutions for these more complex
user needs are yet to fully exploited, with greatest potential in
adaptive systems that take account of patterns of user behaviour
[
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ] [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ], and the use of paths or trails as a means of capturing items
of interest [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ].
3. METHODS
In the absence of an existing system, extensive requirements
gathering [
            <xref ref-type="bibr" rid="ref18 ref6">6</xref>
            ] was conducted with potential users, as the first stage
in a user-centered design process. The goals of this research were
to:
          </p>
          <p>Develop a detailed understanding of the characteristics
and needs of potential users across four primary
domains: heritage, education, professional/commercial,
and general/leisure.</p>
          <p>Explore the meanings and potential applications of the
path metaphor in the context of digital CH.</p>
          <p>Gain an understanding of the path-creation process and
the types of paths that might be created.</p>
          <p>Determine the current availability and functionality of
path-creation tools in CH collections.</p>
          <p>In order to achieve these goals, mixed methods were employed,
gathering a variety of complementary qualitative and quantitative
data. First, an online user survey was used to collect data from 79
expert and non-expert users, comprising questions about their
personal and cultural participation characteristics, and information
behavior and use in the CH context. This was complemented by
in-depth semi-structured interviews conducted with 22 expert
users, which focused on exploring the meaning of the path
metaphor in CH environments, and understanding the process of
development and use of paths in this context.</p>
          <p>Secondary data was used to scrutinise the features of published
paths from various sources, to ascertain their core elements.</p>
          <p>Similarly, a comparative analysis of general and cultural-heritage
specific systems offering path-creation functionality was
conducted, to identify common features and standard approaches
to the proposed core functionality. These findings were validated
via user participation in path-creation tasks, utilsing low- and
medium-fidelity techniques.</p>
          <p>Analysis of these various complementary data enabled the
development of detailed domain and role-specific information
user profiles; a user interaction model supporting four key modes
of interaction; and, use cases illustrating some of the primary user
interaction scenarios. From these we extrapolated detailed user
requirements, and in turn, interface designs and functionality for
the first PATHS prototype. The resulting system is intended to
support all elements of the interaction model, allowing users to
move seamlessly between modes of use.
4. RESULTS
Given the breadth and depth of data, this paper focuses on the
findings relating to paths and their uses in CH, and in turn, how
related user tasks are incorporated within a single user interaction
model, to be implemented in the prototype system.
4.1 Existing Path Forms
Analysis of existing paths and trails found that online and offline
paths both have similar characteristics. Nodes are the essential
building blocks of all paths, representing collection objects. Each
node has associated metadata and primary content (e.g.
descriptions, images) relating to the object. Connections between
nodes enable navigation through the path and often represent
meaningful relationships between objects. In the online
environment, additional features of paths included navigation
tools (e.g. path overviews and back/forward arrows), annotations
added by the path creator to give context and guidance for use,
and occasionally links to other related content, both within the
same collection, and/or in external web sites. These findings
largely support the initial vision for PATHS and can all be seen in
the first prototype design.</p>
          <p>In addition, it was found that most existing online paths are static
and pre-published by an author, linear in form, rather than a more
complex map or network structure, and standalone, without
interconnections with other paths. These findings fall somewhat short
of the PATHS vision, limiting the possibilities for exploration and
discovery, although for pragmatic reasons, they form the core
functionality of the first prototype, with more advanced variations
of paths coming later.
4.2 The Path Metaphor in Cultural Heritage
Interviews with potential expert users in the heritage, education
and professional domains found a strong affinity with the path
metaphor, revealing a range of different interpretations of what it
means in the CH context, and similarly about what form paths
might take, and how they could be employed in an online
environment to engage with key audiences. Eight interpretations
of the path metaphor emerged:</p>
          <p>Path as information literacy journey / learning process</p>
          <p>
            Path as transaction process
The first three of these are closest to the idea of hypertext trails
[
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ], with trails defined by user interaction in 1 and 2, and trails
defined automatically, by the system in 3. Variations 4-6 are more
creative interpretations, all suggesting opportunities for guiding
the user into and through collections, encouraging exploration
and/or offering an immersive experience. In addition to
expertdefined routes, 5 also incorporates the idea of users being able to
see  and  follow  “well -trodden  paths”  defined  by  the  cumulative 
interactions of other users, thus extending the opportunities for
utilizing search histories. Lastly, 7 and 8 are both process
oriented, although 7 is experiential, user-defined,
learningoriented, typified by trial and error and unique to the individual,
whilst 8 is a rigid process designed to escort all users consistently
through a standard process of pre-defined steps.
4.3 Desired Characteristics of Paths
Expected characteristics of paths were explored, and views
contrasted markedly with the existing path formats enabled by
          </p>
          <p>Path as search history
Path as information seeking journey
Path as linked metadata
Path as a starting point or way in
Path as a route through
Path as augmented reality
path-creation tools currently available. Linearity is rarely seen as
the best option for maximizing the potential of paths as
exploration devices. Allied to this is the belief that starting and
end points for paths should be mutable rather than fixed, allowing
different users to explore a path in different ways according to
their preferences and needs.</p>
          <p>In the absence of linearity, some form of organization is still
required to aid the accessibility and navigation of the path. The
most popular option is for path content to be aligned to themes,
with other alternatives including date, location, narrative and
author, where the latter might present multi-layered paths offering
the differing perspectives of several path-creators on the same
topic. An over-arching conceptual framework for the path is also
desirable, in order to tie together the themes and other ideas.</p>
          <p>As a way-finding or navigational aid, paths are seen to support
both guided and exploratory behavior, with the latter seen as the
more desirable goal for user interaction. Features that are needed
to enable way-finding include path overviews, navigational
context in the form of next/last and nearby nodes, branching
opportunities where paths converge and diverge, visualization,
e.g. in the form of timelines or maps, and some degree of object
level information at the node and overview display.</p>
          <p>Path  content  must  be  carefully  selected  or  ‘curated’  by  the  path
creator, with the addition of context and interpretation so that the
objects within the path convey a narrative or meaning. Content
may be derived from one collection, but there are significant
benefits from including objects from diverse collections, along
with other materials from external web sites. It may also be
beneficial for interpretation of the path content to be extended by
user-generated content and/or annotations of various kinds.</p>
          <p>Many of these characteristics are seen in existing path systems,
but limitations arise from the linearity that is commonplace.</p>
          <p>Exploration and deeper levels of engagement within collections
requires more complex path structures, carefully curated content,
interpretation and narrative, and interconnectedness of paths and
other content within and outside of the system. The fact that most
of these more advanced characteristics are rare, and that linearity
prevails also suggests that these are complex issues yet to be
adequately resolved.
4.4 Potential Applications of Paths
Many opportunities for the use of paths in CH were suggested.</p>
          <p>Two major themes emerging from these are the use of paths to
achieve learning, and to support exploration and browsing. For
learning to occur there needs to be strong contextual information,
along with questions and other exercises to structure the learning
process. Exploration and browsing activities implicitly enable
meaning making and learning to take place, as users become more
familiar with a topic and select or interpret the objects they
encounter.</p>
          <p>Specific instances of learning activities that may be delivered via
paths are collection or subject familiarization, story-telling,
individual or collaborative inquiry-based learning utilising path
creation, modeling the research process, and comparative analysis
of differing view-points on a topic of interest.</p>
          <p>
            In addition to learning,, paths may also serve to deliver
entertainment and an enjoyable interaction experience for more
general audiences. In practical terms, paths may simply be used as
a means of introducing people to a collection and its stories, and
in due course, encouraging them to venture further in a more
independent fashion. Paths facilitate topic-based information
retrieval typified by the berry-picking mode of interaction [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ],
rather than known item searching. Furthermore, paths may be a
useful tool for personal information management in both formal
and informal research scenarios, enabling the user to record, reuse
and share their research activity, or helping them to organize their
ideas. Creativity is also encouraged, as user-generated paths
provide the means to repurpose CH objects into users’ own
narratives for private or public consumption.
5. USER INTERACTIONS WITH PATHS
By consolidating findings across the various data collection
methods, we were able to discern five core elements of interaction
with CH collections relating to activities that encompass creating,
using and sharing paths as a means of exploration and
engagement.
          </p>
          <p>Findings from the qualitative data collected via interviews and
path-creation tasks revealed a set of five core activities relating to
the creation, use and sharing of paths; developing a concept for a
path; collecting items in include in a path; creating a path from
items collected; communicating about paths found and about
paths created; and, consuming (following or exploring) paths
created by others. All elements of the model may be undertaken
by expert and non-expert users, in any sequence, and with varying
degrees  of  iteration,  according  to  the  user’s  preferences  and 
behavioural traits.</p>
          <p>Figure 1. PATHS user interaction model.</p>
          <p>Initially, we expect users to begin by Consuming paths created by
others, using them as a means of exploration and familiarization
with the collection and the system. Collecting items of interest
when exploring a collection is a natural behaviour in
berrypicking mode, and is implicit in the process of creating a path, or
as a by-product of a user’s information seeking history. When the
path creation activity is purposeful, it is likely that an
overarching Concept is devised, which may come from activities
undertaken outside of the system, but also may be developed via a
process of exploration within the collection and any pre-existing
paths. The concept may also evolve alongside the collection and
path creation activities, through a process of iteration and
meaning-making. A path is Created once a number of appropriate
items have been collected, and this activity may include ordering
the items into a narrative, and adding contextual information
and/or metadata. In a web 2.0 environment, it is also important to
allow for Communication activities in support of the interaction
experience. These may include sharing paths that have been
created or discovered, both within and outside of the system (e.g.
via social media), commenting on and rating content, and adding
narrative to personal paths as a means of making meaning.</p>
          <p>It is imperative in an adaptive web environment that systems do
not prescribe modes of interaction or enforce sequences of
activities. During the design of PATHS we have uncovered four
primary interaction modes, all of which are supported by the user
interaction model, but each with a somewhat different typical
interaction flow.</p>
          <p>Path consumers are the most passive users, and likely to be in the
majority. By using paths as a guided tour or means of simple
exploration of the collection and its content, we expect users to
become more interested in communicating their discoveries with
others and exploring further within the main collection. Over time
we would expect some of them to move onto collecting and
creating paths of their own, as they develop into more
independent and active users of the system.</p>
          <p>Path creators will likely be in a minority in the early stages, and
primarily expert users such as curators and educators, and perhaps
a few more independent non-expert users. In expert path-creation
mode we believe interaction will be purposeful and systematic,
with a goal of creating a path about a defined topic. Topics and
styles of paths may vary by domain, and we expect that educators
are more likely to adapt ideas from existing paths, whilst CH
experts will try to develop something novel, showcasing elements
of a collection or subject expertise. In contrast, non-expert path
creators are more likely to develop their concept as they explore
the collection, and their paths may be more idiosyncratic,
evolving over time, or in the education domain, may even be
directed in the task by an expert in a path facilitator role.</p>
          <p>
            Path facilitators are most likely to be found within educational
settings, where inquiry-based learning is prevalent. These users
may not create paths themselves, but may curate a broad
collection of objects from which a group of non-expert users are
encouraged to create their own paths. (for instance, as a
homework project). Facilitators are more interested in enabling
deeper engagement with CH materials, and in fostering
communication and reflection on the activity and the content of
the paths created in this way.
6. CONCLUSIONS
We have presented the findings of our user requirements study on
the creation and use of paths as a means of aiding information
access and exploration in CH digital collections. It has been
shown that paths support many of the needs for exploratory
information behavior, and have applications for diverse users
across multiple domains. Users interactions with paths comprise
five core elements, integrated into a single user interaction model
and can be used in varying sequences, illustrated by four primary
modes of interaction.. An initial prototype has been developed
from the user interaction model, which is currently being
evaluated within a task-based user-centred evaluation setting.
7. ACKNOWLEDGMENTS
The research leading to these results has received funding from
the  European  Community’s  Seventh  Framework  Programme 
(FP7/2007-2013) under grant agreement no. 270082. We
acknowledge the contribution of all project partners involved in
PATHS (see: http://www.paths-project.eu).
8. REFERENCES
[
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ] Amin, A., et al. 2008. Understanding cultural heritage
experts' information seeking needs. In Proc. 8th
ACM/IEEECS Joint Conference on Digital Libraries (Pittsburgh, PA,
June 16-20, 2008) JCDL’08,  ACM New York, NY. DOI=
http://dx.doi.org/10.1145/1378889.1378897
[
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ] Bates, M. 1989. The design of browsing and berry picking
techniques for the online search interface. Online Review,
13:5, 407-431.
[
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ] Camhi, J. 2008. Pathways for communicating about objects
on guided tours. Curator: The Museum Journal 51, 3,
275294.
          </p>
          <p>Supporting Serendipitous and Focused Search</p>
          <p>Junte Zhang
Meertens Institute, Royal Netherlands Academy of Arts and Sciences</p>
          <p>Amsterdam, the Netherlands
ABSTRACT
People with complex information needs are for example
Humanities researchers, who need advanced search engines to
investigate their research questions. Much can be gained
by combining research datasets, reusing tools and
serendipitously discovering new insights for further research.
Humanities researchers have di↵ erent (large-scale) research datasets
and tools, which are described di↵ erently with metadata.</p>
          <p>We present a highly interactive advanced search engine for
Humanities researchers that semantically converges di↵
erently structured metadata records from di↵ erent collections
and institutions. It has features that support serendipitous
and focused search in context based on the structure of the
metadata used. This single system serves Humanities
researchers by allowing them to search interactively across yet
unexplored (research) data, discover patterns, locate
relevant data for new insights, and find existing tools that could
provide novel use cases.</p>
          <p>Categories and Subject Descriptors
H.3.3 [Information Search and Retrieval]: Search
process; H.3.7 [Digital Libraries]: Systems issues, user
issues; H.5.2 [Information interfaces and presentation]:
Graphical user interfaces (GUI)
General Terms
Design, Human Factors
Keywords
information retrieval, metadata, user interfaces, ehumanities
1. INTRODUCTION</p>
          <p>
            The Common Language Resources and Technology
Infrastructure (CLARIN) initiative seeks to establish an
integrated and interoperable research infrastructure of language
Presented at EuroHCIR2012. Copyright c 2012 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
resources and its technology.1 Descriptive metadata is used
to characterize large number of (legacy) research data
resources (collections) and tools (e.g. Web services) to
facilitate their management and discovery. The Search &amp; Develop
(S&amp;D) project within CLARIN in the Netherlands uses the
Component MetaData Infrastructure (CMDI; [
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ]) with
ISOcat [
            <xref ref-type="bibr" rid="ref12 ref18 ref24 ref6">6, 12</xref>
            ] to open up the sharing of resources and Web
services for people and machines first within the collections of
a single institution, then across institutions in the
Netherlands and eventually across Europe as whole. This
infrastructure enables new research methods in language research
and stimulates the Digital Humanities, where new insights
can be gained by combining and reusing resources from
different institutions and domains, and existing tools can be
more e↵ ectively found and reused based on new insights.
          </p>
          <p>
            How to use the CMDI framework with ISOcat to search
for data and services, which can be understood by both
people from varying disciplines and machines? The challenge is
that the data is heterogenous both in content and
structure, and can be massive in amount. In [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ], we show how
to deal with such heterogeneously structured data in the
CMDI MI Search Engine. Users of the CMDI framework
are mostly Humanities researchers. What type of system is
needed driven by CMDI that matches with the search
behavior of these users? This paper presents a proposition that
has been implemented on a live system.
          </p>
          <p>USING CMDI FOR FOCUSED AND
SE</p>
          <p>MANTIC ACCESS</p>
          <p>
            CMDI has grown out of the need to facilitate access,
reuse, and interoperability using metadata [
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ]. A CMDI file
in XML consists of a &lt;Header&gt;, &lt;Resources&gt;, and
&lt;Components&gt;. The former two are fixed in structure, while the
content and structure within &lt;Components&gt; is flexible and
can encapsulate any data in any structured form. An XML
schema can be used to make CMDI files coherent in
structure for a (sub)collection and it contains references to ISOcat
data categories (DC) stored in the Registry (DCR; [
            <xref ref-type="bibr" rid="ref18 ref19 ref6 ref7">7, 6</xref>
            ]).
          </p>
          <p>
            The DCR was established by the ISO Technical Committee
37, Terminology and other language and content resources
based on the ISO 12620:2009 standard. Because multiple
elements may refer to the same DC, semantic interoperability
can be achieved across di↵ erent datasets. A specification
using the DCR and projected for example in an XML schema
is called a metadata profile and can be (re)used for
describ1See
http://www.clarin.eu/external/index.php?page=aboutclarin
(a) Query autocompletion based on the count that a query (b) The selection widget that allows users to keep overview of
occurs in a tag within the result set. By default the query box the search trail and change it, while updating the result list.
is content-centric, but searching directly in a tag is possible Here, the query stored is “periode” (period) within the tag
with Advanced Search (can be collapsed with a click). Users time coverage!description. Interesting terms are suggested
can express queries using the metadata or only the fulltext by presenting the top TF⇤ IDF terms, which people can use
of the document by discarding autocompletion. to start a parallel search episode.
(c) To further support query expansion and serendipitous in- (d) The distribution of retrieved time-referenced documents
formation seeking, a dynamic tag cloud is generated based (given the tags Century of Publication and Year of
Publicaon the last retrieved result list and used metadata label with tion) are visualized in bar or line charts. Users can click in
keyword highlighting. Moreover, retrieved geo-referenced the charts to narrow down the result set. The distribution of
documents are projected on a map and clustered by markers. results in tags collection and schema profile always appear.
(a) Retrieved list of results with the display of the list of
results with ‘fixed’ contextual information, snippets and
keywords in context within the last searched metadata label and
the presentation of all used keywords in context given the
fulltext. There is links to the fulltext of the metadata record
and the actual resource in the digital archive.
(b) For each retrieved result in the list, there is a
recommendation (when available) of related results based on the
content similarity of the last used metadata label. A
recommendation consist of a link to the record, the collection it
belongs to, and a snippet (can be collapsed with a click).
ing datasets and for eventual access. Moreover, RELcat [
            <xref ref-type="bibr" rid="ref10 ref22">10</xref>
            ]
goes a step further by allowing for the storage of arbitrary
relationships between data categories to assist crosswalks
and to specify ontological relationships for further semantic
search, which in the future can be used in the CMDI MI
Search Engine using field collapsing.
          </p>
          <p>
            We have indexed 246,728 CMDI files from 18 di↵ erent
profiles consisting of 143 di↵ erent types of elements in a single
stream, which shows our indexing method for CMDI files is
robust enough to deal with complex data [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ]. By indexing
metadata in CMDI on the XML element level, the search
engine can provide focused access [
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ]. We use straight-forward
information retrieval techniques only. The ‘Liederenbank’
(Dutch Song Database) alone has 9 di↵ erent profiles (XML
schemas), which is equivalent to a sub-collection, ranging
from very di↵ erently structured descriptions about songs to
singers. How to provide interactive access to such
heterogeneously structured data for Humanities researchers?
3. SERENDIPITY IN CONTEXT
          </p>
          <p>
            When a user with no a priori intentions interacts with a
node of information and acquires useful information, then
serendipitous information retrieval occurs [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ]. The success
of serendipitous discovery is not just the find itself, but
being able or willing to do something with it, so that users get
more insight and can enhance the domain expertise [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ].
Humanities researchers are the type of users who can be greatly
supported in their research tasks with serendipitous IR,
because their information-seeking behavior can be described
as an idiosyncratic process of constant reading, “digging,”
searching, and following leads [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ]. This confirms with the
Berrypicking model of [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ], such as that queries are not static,
but rather evolve, and users “gather information in bits and
pieces instead of in one grand best retrieved set.”
          </p>
          <p>Since the CMDI MI Search Engine should serve
Humanities researchers, we design it to support serendipitous search
and be highly interactive. The system has been designed to
maximize the user’s ability to explore. This is our focus.</p>
          <p>
            The user interface of the system is depicted in Fig 1. It uses
the JavaScript library AJAX Solr2, which has been
heavily modified and extended by us with JQuery. It allows for
faceted search [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ] as we treat the indexed elements of the
CMDI files as one large category hierarchy.
          </p>
          <p>A user can improving the search episode (session) by
effectively reducing the information space step by step. These
steps are stored as part of the search trail, so the overview
is kept. There are di↵ erent search strategies possible. Users
can search by fulltext by entering a query. This makes sure
users can always search in everything. The query get
highlighted in context given the fulltext, but the dynamic tag
cloud widget that supports query expansion is not activated,
see Fig.1(a). Users can also do a focused search request by
using structure, i.e. within the content of a specified tag,
and get the content of these tags returned. This can be
content-centered, as users enter a keyword and the
autocompletion widget returns a list consisting of keyword plus
field name and hit count. It can also be structure-centered
(using the Advanced Search option) by looking up a tag and
then entering a keyword also with the autocompletion
feature. When the last two options are used, then the keyword
highlighting also occurs within the context of the retrieved
snippets of the searched tag, see Fig.2(a).</p>
          <p>A challenge is how we can support serendipitous search
given the diversely structured metadata in CMDI. Hence, we
introduce and propose the concept of serendipitous search in
context. We can use the heterogeneous structure of di↵ erent
collections to provide context to the user in a single search
engine. We propose the following contextual system features
that aim to support serendipitous and focused search.</p>
          <p>• Help users by automatically completing the query that
the user is entering while simultaneously and directly
giving the hit count for the suggested queries in
conjunction with a tag, see Fig.1(a).
• Provide inline suggestions (Did you mean...) based on</p>
          <p>a spell checker whenever applicable.
• Suggest a new parallel search episode (You could also
look for...) by presenting interesting terms based on
the content of the first few retrieved results after each
used query, see Fig.1(b). This increments and becomes
more focused as a search episode gets more queries.
• O↵ er di↵ erent overviews of the retrieved results and
allow for query expansion by directly presenting a
dynamic tag cloud of the aggregated content within the
metadata label used and highlighting the query entered
in this context, see Fig.1(c).
• Preserve the overview of a search episode by storing
the search selection (see Fig.1(b)), and the overview on
collection level by the result type, e.g. the metadata
profile ‘lied’ (song) in the Dutch Song Database, and
the collection a document belongs to (see Fig.1(d)).
• Aggregate and visualize collection-specific search
features in extra widgets, such as projecting and
clustering the list of retrieved geo-referenced resources on a
map (see Fig. 1(c)), and displaying the date ranges of
the documents in charts that can be clicked to narrow
down a result set (see Fig. 1(d)).
• Entice users to explore further by recommending
related resources using the content similarity by
presenting a link to the metadata record and a snippet of a
recommendation, see Fig.2(b).</p>
          <p>So the context consists of di↵ erent modalities and features
existing in the structure of the metadata of a collection, and
used in the retrieval and visualization of information. This
can be displayed on a aggregated level based on the set of
retrieved results. And it can be displayed with di↵ erent
displays of the result types given the metadata profile.
Eventually, the user finds the links to the resources in the digital
archive using the metadata, and can use the found resources
for further research or development. However, there is no
real definite end of the search episode as people still can
continue searching using the above proposed system features.</p>
          <p>We have presented a working proposition for
serendipitous and focused search by describing the CMDI MI search
engine. The novelty is that it provides semantic access to
diversely structured language and digital heritage resources
with di↵ erent metadata schemas for users such as researchers
with very specific and complex information (research) needs.</p>
          <p>The search engine provides faceted search and has
serendipitous features that maximize the user’s ability to explore any
metadata in CMDI in context, such as query
autocompletion, tag clouds, and recommendation of related resources,
while keeping track of the search trail. It is a tool that
provides interactive and focused access to heterogeneous
metadata, gives new perspectives on legacy (research) data and
tools, and provides new insights for research and
development. It has been released as live, and can be used at
www.meertens.knaw.nl/cmdi/search.</p>
          <p>ACKNOWLEDGMENTS</p>
          <p>This work is part of the Search &amp; Develop project at the
Meertens Institute, and funded by CLARIN-NL.</p>
          <p>Marcus Nitsche
Faculty of Computer Science,
Otto-von-Guericke-University,</p>
          <p>Germany
marcus.nitsche@ovgu.de</p>
          <p>Andreas Nürnberger
Faculty of Computer Science,
Otto-von-Guericke-University,</p>
          <p>Germany
andreas.nuernberger@ovgu.de
When users search for information in domains they are not familiar
with, they usually struggle to formulate an adequate (textual) query.</p>
          <p>Often users end up with repeating re-formulations and query
refinements without necessarily achieving their actual goals. In this paper
we propose a user interface that is capable to offer users flexible and
ergonomic interaction elements to formulate even complex queries
simple and direct. We call this principle vague query formulation
by design. By this formulation we like to point out its design-driven
origin. The proposed radial user interface supports phrasing and
interactive visual refinement of vague queries to search and explore
large document sets. The main idea is to provide an integrated
view of queries and related results, where both queries and results
can be interactively manipulated and influence each other. Changes
will be immediately visualized. The concept was implemented on
a tablet computer and the usability was stepwise evaluated during
a formative and a summative user study. The results reveal high
usability ratings, even if the concept was completely unknown to
our test users.</p>
          <p>Keywords
Search User Interface, Query Reformulation, Query Refinement,
Information Retrieval.</p>
          <p>Categories and Subject Descriptors
H.3.3 [Information Storage and Retrieval]: Information Search
and Retrieval.; H.5.2 [Information Interfaces and Presentation]:
User Interfaces.</p>
          <p>General Terms
Design, Human Factors, Management.</p>
          <p>INTRODUCTION</p>
          <p>
            When users try to handle complex information needs they often
end up in conducting exploratory searches [
            <xref ref-type="bibr" rid="ref11 ref23">11</xref>
            ]. One of the main
characteristics of exploratory searches is that users often do not
Presented at EuroHCIR2012. Copyright c 2012 for the individual papers
by the papers’ authors. Copying permitted only for private and academic
purposes. This volume is published and copyrighted by its editors.
know how to formulate their information need. Often this problem
coexists with an unfamiliarity with the domain they search in [
            <xref ref-type="bibr" rid="ref29">17</xref>
            ].
          </p>
          <p>
            In this work we like to tackle this problem of querying
appropriate queries by offering users dynamic user interface (UI) elements
that they can manipulate directly by touch gestures to give them
a feeling for a certain query configuration that matches a certain
result set. Thereby learning and exploring aspects will covered as
well [
            <xref ref-type="bibr" rid="ref11 ref23 ref29">17, 11</xref>
            ]. This concept of interactive visual filtering of
relevant information in a more natural way enables data processing in
cases, where standard algorithms can not be applied since these
algorithms might filter out relevant data. We introduced the concept
of this paper back in 2011 [
            <xref ref-type="bibr" rid="ref27">15</xref>
            ], where we described the basic idea
and did some pre-studies with a digital mockup prototype. In this
paper, we first introduce a running implementation and a more
detailed user study towards this concept. Therefore we present some
related work aspects in Section 2, followed by a presentation of the
UI concept in 3 and the description of the implementation,
evaluation concept and results of the final user study in Section 4. Finally,
we conclude and discuss possible future work in 5.
          </p>
          <p>User-specific context aware data filtering is not a new challenge.</p>
          <p>
            In the following we show two tools, that can also be used for these
application domains. The VIBE-system [
            <xref ref-type="bibr" rid="ref10 ref22 ref28">10, 16</xref>
            ] supports users in
finding relevant information using magnets to attract relevant
documents to specific screen points (Fig. 1).
          </p>
          <p>
            This system follows the principle of dust-and-magnet [
            <xref ref-type="bibr" rid="ref30">18</xref>
            ]. Our
proposed concept uses this principle also as one aspect of
interaction. In contrast to VIBE we offer users of our system an
interactive visualization without any classical WIMP-interface elements
(Windows, Icons, Menus, Pointer). By this, no virtual mapping of
functions is necessary and users might be able to use the interface
in a more firm and reliable way. Cousins et al. [
            <xref ref-type="bibr" rid="ref17 ref5">5</xref>
            ] developed a
system that follows a direct manipulation approach like done here.
          </p>
          <p>But in contrast to our proposed solution it is divided into different
UI elements and different views. It is less integrated in a single
view. Therefore user’s work load might be higher since he needs to
face various mode switches. Commercial systems, like the Vis4you
concept1, are more focused on visualization than on interaction via
direct manipulation. Furthermore, this system is designed to be
used on desktop computers with a mouse (single point and
clickprinciple), no multi-touch-support. In the next section we like to
present our concept in more detail.</p>
          <p>CONCEPT &amp; DESIGN</p>
          <p>
            Due to the increasing amount of data and complexity, it is
necessary to apply and improve the concepts of visual information
filtering and retrieval. This goes along with the underlying methods and
tools. Considering clustering algorithms (e.g., k-means [
            <xref ref-type="bibr" rid="ref15 ref3">3</xref>
            ]), we
thought about the concept of vague query formulation: Since users
sometimes do not know what they are searching for, we like to
support them by the opportunity to formulate vague queries. Here the
user is asked to narrow the search results by dragging user interface
(UI) elements, so called widgets, with query terms, see also Fig. 2.
          </p>
          <p>The concept follows the idea that more relevant data are
centred. Note, this is equivalent to filtering an overcrowded desktop,
cf. Fig. 3 (left picture)2, where the more centralized documents are
possibly more important (highlighted in the right picture).</p>
          <p>
            The system was designed to be a multi-user system. Therefore
a number of multiple users need to be supported at the same time,
also considering security aspects [
            <xref ref-type="bibr" rid="ref26">14</xref>
            ]. To offer each user the same
possibility to interact with the system we use for the interface a
radial form. Furthermore, an underlying multi-touch device is a
hardware requirement, that enhances the combination of tool and
application domain significantly. Another appealing advantage is,
that multi-touch also supports users in a more natural way of
interaction [
            <xref ref-type="bibr" rid="ref21 ref9">9</xref>
            ]. Other radial user interfaces for selecting or filtering
often offers fixed places for items. In contrast to this our system is
supposed to be more flexible since users are allowed to position UI
elements where they like.
          </p>
          <p>We offer users a dimension merging according specified weights,
similar to the result listing of search engines, where also different
1http://www.vis4you.com/vis4you/ (accessed on 04.07.2012)
2http://lawprofessors.typepad.com/ (accessed on 04.07.2012)
weights can be linked to specific query terms (Fig. 2). Data points
represent the data space. Query objects, so called widgets, can be
entered via a virtual keyboard and can also be dragged by the user
to formulate more complex or vague queries. Selecting a specific
data point supports the user with additional information on this data
point and highlights all related data points.</p>
          <p>
            The distance of a certain term is directly connected to its
importance for the user. In other words, if a user thinks a specific term is
more relevant to its actual filter-/search-task, she or he positions the
corresponding UI-element more to the center, which influences the
weight of this term when computing its Term Frequency / Inverted
Document Frequency (TF/IDF)-value [
            <xref ref-type="bibr" rid="ref14 ref2">2</xref>
            ], which in fact is a
calculated weight to influence the ranking of the data space and this in
return the visualization (Fig. 6). Thereby, users do not need to
specify a concrete position of UI elements on the screen, we support this
by a non-determined precision. The widget-induced relevance of a
query term is calculated according to the formula in Fig. 4.
          </p>
          <p>Result elements are placed near to corresponding query elements.</p>
          <p>The formula for calculating the relevance of a SearchResult object
(result dot) is shown in Fig. 5.</p>
          <p>Figure 5: Relevance of a SearchResult object.</p>
          <p>The calculated relevance determines the distance to the center,
considering further result objects.</p>
          <p>To address various types of end devices such as multi-touch
desktops or mobile interfaces with large displays, we use direct
manipulation as a central interaction paradigm. Only the relative distance
of an UI element to the center is relevant for the system. Thus,
we provide users with a direct linking to the data they like to filter.</p>
          <p>By this interaction concept, we propose to achieve more precise
results. Additionally, we support users with the concept of
Whatif -queries, which supports a fault-tolerant interaction system, using
a ghosting technique: Dragging an element and holding it on a
specific position triggers the system to show the user how many items
are in the center point of interest (POI) after releasing the element.</p>
          <p>Thereby, users are able to explore the impact of possible next steps.</p>
          <p>
            Changes of the query configuration also effect the data points
to provide the user with a direct link to the data (interactive
visualization). By the underlying metaphor of magnets, we offer an
integrated feedback, comparable to Dust-and-Magnet [
            <xref ref-type="bibr" rid="ref30">18</xref>
            ]: When
users drag a specific UI element to a certain point, relevant data
points follow this UI element. Data points that have the same
TFIDF value (equal relevance to a query configuration) are drafted
with a minimal distance to each other to prevent occlusions.
3.1
          </p>
          <p>Features</p>
          <p>The UI supports direct feedback since the relevance value is
simultaneously shown while users interact with the widget (Fig. 7).</p>
          <p>Results, corresponding to a specific query object are visually
highlighted and grouped to each other (Fig. 8).</p>
          <p>Detailed information on particular result objects, like a website
preview, is provided after clicking on the result dot (Fig. 9).</p>
          <p>
            Since this contribution is basically driven by fields of human
factors and user interface design, we are using common methods
from these research areas. Such as user centred design (UCD)
processes [
            <xref ref-type="bibr" rid="ref19 ref7">7</xref>
            ], formative evaluation methods [
            <xref ref-type="bibr" rid="ref12 ref24">12</xref>
            ], questionnaires [
            <xref ref-type="bibr" rid="ref18 ref6">6</xref>
            ],
think-aloud-protocols [
            <xref ref-type="bibr" rid="ref20 ref8">8</xref>
            ], and cognitive walkthroughs [
            <xref ref-type="bibr" rid="ref16 ref4">4</xref>
            ].
          </p>
          <p>
            To proof the concept of the proposed user interface, a prototype
was implemented. This was done by using an Apple iPad.
ThereFigure 9: Prototypical search result popover as a website
preview feature, here a result for ’Labrador Retriever’.
fore the application was written in ObjectiveC using the xCode3
environment. The backend architecture is the CARSA system [
            <xref ref-type="bibr" rid="ref1 ref13">1</xref>
            ],
an information retrieval framework for research purposes. For a
detailed overview about the system’s architecture see Fig. 10.
          </p>
          <p>
            The evaluation concept followed a formative evaluation
principle where several usability testings were conducted. Also in
parallel to the development process: To identify at least 85% of all
usability issues this mock-up was evaluated according to Nielsen
and Landauer [
            <xref ref-type="bibr" rid="ref25">13</xref>
            ] with only a small number of test users since
most usability issues will be mentioned repeatedly by users. The
sixth tested user would report new usability issues in only 15%
3developer.apple.com/xcode/ (accessed on 04.07.2012)
of all cases. Therefore we decided to ask only eight users. The
results of this first user test seem to be promising that this
concept works as desired. Users were introduced in the main features
and were asked afterwards to formulate a filter query consisting of
three terms to find all relevant documents while visualizing most
important relations to other potential interesting data. After
going through a cognitive walk-through of a movie filtering task our
eight test users (six male, two female, average age: 23.4) answered
seven usability questions by filling out a 7-step Likert scale from 1
(very bad) to 7 (very good). Next to cognitive walk-throughs, we
used think-aloud-protocols and questionnaires. The usefulness of
the prototype was rated high, the functionality was praised by test
users, performing tasks were rated as very easy and test users were
satisfied with this prototype. Terminology, attractiveness, and
consistency were rated lower. Our final evaluation revealed the results
you can see in Fig. 11. Even if there is room for improvement the
results reveal overall a good usability, several test users mentioned
that it was fun to use it, which might is reflected by a high rating of
joy of use measurings.
5.
          </p>
          <p>DISCUSSION &amp; OUTLOOK</p>
          <p>We described a newly designed user interface for filtering,
exploring and managing data via direct manipulation supporting
multiple reference systems to support context sensitive interaction
techniques. We proposed a UI concept for visual filtering, that is
• flexible: parameters can be adapted or enhanced by users
• context-sensitive: initial parameters are extracted from the</p>
          <p>current use case
• easy to learn: through work environment metaphor and direct</p>
          <p>manipulation
In near the future a more detailed and larger user study will be
conducted to identify further improvements of our tool and the overall
concept. Also a plan to re-design slightly is already in place.
6.</p>
          <p>ACKNOWLEDGEMENT</p>
          <p>Part of the work is funded by the German Ministry of
Education and Science (BMBF) within the ViERforES II project (no.
01IM10002B). We also thank Martin Schemmer for the
implementation of the presented concept during his diploma thesis.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Andr</surname>
          </string-name>
          ´e, m. schraefel, J. Teevan, and
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Dumais</surname>
          </string-name>
          .
          <article-title>Discovery is never by chance: designing for (un)serendipity</article-title>
          .
          <source>In Proceedings of the seventh ACM conference on Creativity and cognition</source>
          , C&amp;C '09, pages
          <fpage>305</fpage>
          -
          <lpage>314</lpage>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Barrett</surname>
          </string-name>
          .
          <article-title>The information-seeking habits of graduate student researchers in the humanities</article-title>
          .
          <source>The Journal of Academic Librarianship</source>
          ,
          <volume>31</volume>
          (
          <issue>4</issue>
          ):
          <fpage>324</fpage>
          -
          <lpage>331</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Bates</surname>
          </string-name>
          .
          <article-title>The design of browsing and berrypicking techniques for the online search interface</article-title>
          .
          <source>Online Review</source>
          ,
          <volume>13</volume>
          (
          <issue>5</issue>
          ):
          <fpage>407</fpage>
          -
          <lpage>424</lpage>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Broeder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kemps-Snijders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Uytvanck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Windhouwer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Withers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wittenburg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Zinn</surname>
          </string-name>
          .
          <article-title>A data category registry- and component-based metadata framework</article-title>
          .
          <source>In LREC</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Karadi</surname>
          </string-name>
          .
          <article-title>Cat-a-cone: an interactive interface for specifying searches and viewing retrieval results using a large category hierarchy</article-title>
          .
          <source>In SIGIR</source>
          , pages
          <fpage>246</fpage>
          -
          <lpage>255</lpage>
          , New York, NY, USA,
          <year>1997</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kemps-Snijders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Windhouwer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wittenburg</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. E. Wright.</surname>
          </string-name>
          <article-title>ISOcat: remodelling metadata for language resources</article-title>
          .
          <source>IJMSO</source>
          ,
          <volume>4</volume>
          (
          <issue>4</issue>
          ):
          <fpage>261</fpage>
          -
          <lpage>276</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kemps-Snijders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zinn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ringersma</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Windhouwer</surname>
          </string-name>
          .
          <article-title>Ensuring semantic interoperability on lexical resources</article-title>
          .
          <source>In LREC</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lalmas</surname>
          </string-name>
          .
          <source>XML Retrieval. Synthesis Lectures on Information Concepts</source>
          , Retrieval, and Services. Morgan &amp; Claypool Publishers,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E. G.</given-names>
            <surname>Toms</surname>
          </string-name>
          .
          <article-title>Serendipitous information retrieval</article-title>
          .
          <source>In DELOS Workshop: Information Seeking, Searching and Querying in Digital Libraries</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Windhouwer</surname>
          </string-name>
          .
          <article-title>RELcat: a relation registry for isocat data categories</article-title>
          .
          <source>In LREC</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kemps-Snijders</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>H.</given-names>
            <surname>Bennis. The CMDI MI Search</surname>
          </string-name>
          <article-title>Engine: Access to language resources and tools using heterogeneous metadata schemas</article-title>
          .
          <source>In TPDL</source>
          , volume
          <volume>7489</volume>
          of Lecture Notes in Computer Science. Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zinn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hoppermann</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Trippel</surname>
          </string-name>
          .
          <article-title>The isocat registry reloaded</article-title>
          .
          <source>In The Semantic Web: Research and Applications</source>
          , volume
          <volume>7295</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>285</fpage>
          -
          <lpage>299</lpage>
          . Springer Berlin / Heidelberg,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bade</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Luca</surname>
            ,
            <given-names>E. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nürnberger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stober</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>CARSA - an architecture for the development of context adaptive retrieval systems</article-title>
          .
          <source>In: Proceedings of Adaptive Multimedia Retrieval: User</source>
          , Context, and
          <string-name>
            <surname>Feedback</surname>
          </string-name>
          , Volume
          <volume>3877</volume>
          /
          <year>2006</year>
          , Lecture notes in computer science, pp.
          <fpage>91</fpage>
          -
          <lpage>101</lpage>
          , Springer Berlin / Heidelberg (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ribeiro-Neto</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          : Modern Information Retrieval, pp.
          <fpage>29</fpage>
          -
          <lpage>30</lpage>
          . Addison Wesley / ACM Press, NY (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Bradski</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kaehler</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Learning OpenCV Computer Vision with the OpenCV Library</article-title>
          .
          <source>O'Reilly</source>
          , p.
          <volume>479</volume>
          (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Busemeyer</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          :
          <article-title>Choice behavior in a sequential decision-making task</article-title>
          .
          <source>In: Organizational Behavior and Human Performance</source>
          <volume>29</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>175</fpage>
          -
          <lpage>207</lpage>
          (
          <year>1982</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Cousins</surname>
            ,
            <given-names>S. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paepcke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winograd</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bier</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pier</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The digital library integrated task environment (DLITE)</article-title>
          .
          <source>In: Proceedings of the second ACM international conference on Digital libraries (DL '97)</source>
          . ACM, New York, NY, USA, pp.
          <fpage>142</fpage>
          -
          <lpage>151</lpage>
          (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Czaja</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Blair</surname>
            ,
            <given-names>J.: Designing</given-names>
          </string-name>
          <string-name>
            <surname>Surveys</surname>
          </string-name>
          . Pine Forge Press.
          <article-title>A useful resource for factual-style surveys, including material on interviews as well as mail surveys (</article-title>
          <year>1996</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Eason</surname>
            ,
            <given-names>K. D.:</given-names>
          </string-name>
          <article-title>User centred design for information technology systems</article-title>
          .
          <source>In: Physics in Technology</source>
          <volume>14</volume>
          (
          <issue>5</issue>
          ), p.
          <volume>219</volume>
          (
          <year>1983</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Ericsson</surname>
            ,
            <given-names>K. A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Simon</surname>
            ,
            <given-names>H. A.</given-names>
          </string-name>
          :
          <article-title>Verbal reports as data</article-title>
          .
          <source>In: Psychological Review</source>
          <volume>87</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>215</fpage>
          -
          <lpage>241</lpage>
          (
          <year>1980</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          . Y.:
          <article-title>Multi-touch interaction wall</article-title>
          .
          <source>In: Proceedings of ACM SIGGRAPH</source>
          <year>2006</year>
          <article-title>Emerging technologies (</article-title>
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Koshman</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          :
          <source>VIBE User Study Technical Report LS062/IS97001</source>
          , University of Pittsburgh (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Marchionini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Exploratory search: from finding to understanding</article-title>
          .
          <source>In: Communications of the ACM</source>
          <volume>49</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Moxley</surname>
            <given-names>Jr.</given-names>
          </string-name>
          , R. A.:
          <article-title>Formative and non-formative evaluation</article-title>
          .
          <source>In Instructional Science</source>
          <volume>3</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>243</fpage>
          -
          <lpage>283</lpage>
          (
          <year>1974</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Nielsen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Landauer</surname>
          </string-name>
          , T. K.:
          <article-title>A mathematical model of the finding of usability problems</article-title>
          .
          <source>In: Proceedings of ACM INTERCHI'93 Conference</source>
          , pp.
          <fpage>206</fpage>
          -
          <lpage>213</lpage>
          , Amsterdam, The Netherlands (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Nitsche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dittmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nürnberger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vielhauer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buchholz</surname>
          </string-name>
          , R.:
          <article-title>Security-relevant Challenges of selected Systems for Multi-User Interaction</article-title>
          .
          <source>In Proceedings of the 7th International Workshop on Adaptive Multimedia Retrieval</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Nitsche</surname>
          </string-name>
          , Marcus and Nürnberger, Andreas:
          <article-title>Supporting vague query formulation by using visual filtering</article-title>
          .
          <source>In Proceedings of Lernen</source>
          , Wissen, Adaption (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Olsen</surname>
            ,
            <given-names>K. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korfhage</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sochats</surname>
            ,
            <given-names>K. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spring</surname>
            ,
            <given-names>M. B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J. G.</given-names>
          </string-name>
          :
          <article-title>Visualization of a Document Collection: the VIBE System</article-title>
          .
          <source>In: Information Processing &amp; Management</source>
          ,
          <volume>29</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>69</fpage>
          -
          <lpage>81</lpage>
          (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [17]
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R. W.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>R. A.</given-names>
          </string-name>
          :
          <article-title>Exploratory search: Beyond the Query-Response paradigm</article-title>
          .
          <source>In: Synthesis Lectures on Information Concepts</source>
          , Retrieval, and Services, Ed.: G. Marchionini, Morgan &amp; Claypool Publishers (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Yi</surname>
            ,
            <given-names>L. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stasko</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Jacko</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Dust &amp; Magnet: multivariate information visualization using a magnet metaphor</article-title>
          .
          <source>In: Information Visualization</source>
          , pp.
          <fpage>239</fpage>
          -
          <lpage>256</lpage>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>