<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Value of Multistage Search Systems for Book Search</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hugo Huurdeman</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaap Kamps</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marijn Koolen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sanna Kumpulainen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Archives and Information Studies, Faculty of Humanities, University of Amsterdam</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ISLA, Faculty of Science, University of Amsterdam</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Logic, Language and Computation, University of Amsterdam</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Often, our exploratory quests for books are highly complex endeavors which feature activities such as exploration, searching, selecting and comparing various books. Current systems for book search may not provide optimal support for this wide range of activities. The interactive Social Book Search Track investigates how users utilize di erent access interfaces in the context of two types of tasks, and evaluates a streamlined baseline interface and a rich multistage interface, potentially supporting di erent stages of search. In this paper, we analyze how these two types of interfaces in uence user behavior, in terms of task duration, book selection and interaction patterns. Furthermore, we characterize the use of the di erent panels of the experimental multistage interface, as well as user engagement. We nd initial evidence for the additional value of providing stage-based search support in the context of open-ended and focused book search tasks.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The interactive Social Book Search (iSBS) Track studies how searchers use
professional and user-generated metadata during di erent stages of complex search
tasks. The iSBS track uses two experimental interfaces (a baseline and
multistage interfaces), combined with open-ended and focused search tasks. Research
groups participating in the iSBS track had to recruit at least 20 users for the
shared study to gain access to the collected data from the experiment. In 2015,
the second iteration of the iSBS track took place, and 7 teams recruited 192
participants for the study, resulting in a rich dataset.</p>
      <p>This paper describes the University of Amsterdam's participation in this
track, and we analyze the in uence of task and interface on user behavior, in
terms of task duration, book selection and interaction patterns. In addition, we
characterize user engagement with both experimental interfaces.
Previous work related to the iSBS track has been carried out in the INEX
Interactive Retrieval Experiments (2004-2010) [6], the Cultural Heritage in CLEF
(CHiC) Interactive Track 2013 [8], and the INEX 2014 Interactive Social Book
Search (ISBS) Track [3]. In these tracks, a standard procedure for collecting
data was being used by participating research groups, including common topics
and tasks, standardized search systems, document corpora and procedures. The
system used for the iSBS track is a modi ed version of the one used for CHiC
and is based on the Interactive IR evaluation platform developed by Hall and
Toms [1], where di erent search engines and interfaces can be plugged into fully
developed IIR framework that runs the entire user study [2].
3</p>
    </sec>
    <sec id="sec-2">
      <title>Experimental Setup</title>
      <p>In this section we describe the tasks, system interfaces and our pre-processing
of the data generated by the experiment.
3.1</p>
      <sec id="sec-2-1">
        <title>Tasks</title>
        <p>The experiment includes two search tasks, a focused task and an open task, and
each participant performs both. During the focused task, users were asked to
compile a list of books, each matching a speci ed criterion. The focused task
contains ve sub-tasks, some of which are speci c and some are more open:
Imagine you participate in an experiment at a desert-island for one
month. There will be no people, no TV, radio or other distraction. The
only things you are allowed to take with you are 5 books. Please search
for and add 5 books to your book-bag that you would want to read
during your stay at the desert-island:
{ Select one book about surviving on a desert island
{ Select one book that will teach you something new
{ Select one book about one of your personal hobbies or interests
{ Select one book that is highly recommended by other users (based
on user ratings and reviews)
{ Select one book for fun
Please add a note (in the book-bag) explaining why you selected each of
the ve books.</p>
        <p>The open task is derived from the non-goal task used in the iCHiC task at CLEF
2013 [9], which allows participants to come up with their own goals and
subtasks. During the open task users could explore the collection based on their
own interests, for as long as they wanted:</p>
        <p>Imagine you are waiting to meet a friend in a co ee shop or pub or the
airport or your o ce. While waiting, you come across this website and
explore it looking for any book that you nd interesting, or engaging
or relevant. Explore anything you wish until you are completely and
utterly bored. When you nd something interesting, add it to the
bookbag. Please add a note (in the book-bag) explaining why you selected
each of the books.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Interfaces</title>
        <p>Two interfaces were developed for this study. The baseline interface is a standard
search interface with a single screen, containing a query box, a left column with
facet lters and a right column with a book bag in which users could store books
they had selected. The multistage interface contains three screens, each with its
own functionality to support di erent stages in the search process. It is inspired
by various models of the information seeking process [5, 10], which indicate that
users experience evolving stages of search in the context of complex tasks.</p>
        <p>The Browse screen only allows browsing through predetermined categories
(based on Amazon book categories), where the middle panel shows lists of book
titles which users can click on to get detailed information on that book and the
ability to save it to the bookbag.</p>
        <p>
          The Search screen has a search box and search results in the middle panel,
search lters based on the Amazon book categories and on user-supplied tags
from LibraryThing users. By default, the detail-view of a search result shows
a thumbnail and a publisher-supplied description of each book and four tabs
that allow the user to switch between (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) the publisher-supplied description, (2)
publication metadata, (3) Amazon user reviews and (4) LibraryThing user tags.
        </p>
        <p>The Book-bag interface shows the bookbag in the left panel and per book
a number of buttons allowing the user to search for similar books. There is a
separate button for books by the same author, one for books with similar titles
and one for books with the same subject categories. When clicking one of these
buttons, the right panel shows the search results.
3.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Data</title>
        <p>The transaction log contains transactions of 192 users who completed both tasks,
with 97 users using the baseline interface for both tasks, and 95 using the
multistage interface.</p>
        <p>The dataset collected during the iSBS study consists of questionnaire and
logging data. In this paper, we focus on the system logs in section 4, while we
look at the questionnaire data in section 5.</p>
        <p>The log data includes the duration of each task, which we used to
calculate task duration for each task and experimental interface. In the multistage
interface, the user starts in the browse panel of the interface. Each time a user
switches between interface stages (explore, search and book-bag), this is logged
as an action. Using these switches, we can reconstruct all actions per interface
stage. We found a very small number of impossible combinations (168 out of
22,152, such as adding a search lter in the explore stage, which has no lters).
We surmise this is either because some stage switches were not logged or because
a switch was logged but did not take place.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Analysis of Results</title>
      <p>We compare the two interfaces based on a number of aspects: 1) task duration,
2) di erence in book bag content between the two interfaces, and 3) di erence
in the types of actions performed. Finally, we zoom in further on the multistage
interface and look at the use of the di erent panels in the multistage interface.
4.1</p>
      <sec id="sec-3-1">
        <title>Task Duration</title>
        <p>First, we examine the duration of the included tasks. Di erences can be expected
at both the task and interface level. The focused task is more complex, so users
might take longer to complete that task than the open task.</p>
        <p>The distribution of task lengths in seconds is shown in Table 1. The
majority of users spent less than 15 minutes on the task|the median is just under
12 minutes for the focused task (696.5 seconds) and just over 6 minutes (369
seconds) for the open task|but a few spent an hour or more (1 for the focused
task and 3 for the open task). Due to such outliers the mean is higher than the
median. The higher median and mean for the focused task is probably due to
the higher complexity of the task, as it consists of 5 sub-tasks.</p>
        <p>Also, di erences in the task time between the baseline and multistage
interface can be observed: the median task time for both the focused and the open
task is higher for the multistage interface. Hence, participants spend more time
in the multistage interface, regardless of the nature of the task.
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Bookbag</title>
        <p>Next, we analyse the content of the bookbag at the end of each task, to determine
whether users show di erent book selection behaviour across the tasks and the
two interfaces. Given the very di erent natures of the tasks, we expect to see
clear di erences between the bookbags after each task. The focused task asks
the user to select books given a list of ve criteria, which may steer the user to
select ve books. In the open task, users are instructed that they can select as
many or as few books as they like. Therefore, we expect the number of books in
the bookbag in the open task to be more widely distributed.</p>
        <p>Indeed, Table 2 indicates that the number of books in the book-bags are
higher for the focused task: on average, participants choose 4.75 books in the
focused task, and 3.45 for the open task. Also, the standard deviation is
substantially higher for the open task, so there is more variation in the number of
books that participants selected.</p>
        <p>Some of the sub-tasks of the focused task may be interpreted as similar to
the open task, e.g. books about one of your personal hobbies or interests (third
sub-task) and books for fun ( fth sub-task). If that is the case, user may simply
add some of the same books to the book-bag in both tasks. We checked the
overlap between the books in the book-bag for the focused and the open task
and found that only 9 users have some overlap in the book-bags, with 7 only
having a single book in both bags. From this we conclude that user treat the
sub-tasks of the focused task as di erent from the open task.</p>
        <p>An additional question is whether the supplied interface makes a di erence.
Our analysis shows that the number of gathered books is slightly higher for the
multistage interface, especially in the case of the open task.</p>
        <p>We also look at the overlap between the book bags of users, that is, whether
di erent users nd and select the same books or di erent books. The ratio
between the size of all book bags combined as a bag (with repetition) and as a
set (without repetition). The ratio of number of overall books selected over the
number of distinct books selected. For the open task, the overlap ratio of the
baseline interface (309 book selections of 290 distinct books) is 1.07 and the
overlap ratio of the multistage interface is also 7% (358 book selections of 335
distinct books). The overlap is low, which is not surprising given the open
nature of the task, and the type of interface seems to have little e ect. For the
focused task, the overlap ratio of the baseline interface is 1.23 (442 selections of
360 distinct books) and that of the multistage interface is 1.14 (470 selections
of 412 distinct books). The overlap for the focused task is thus higher than for
the open task, probably because all users are constrained in their selection by
the more speci c sub-tasks. Here the type of interface has a larger e ect. The
users of the baseline interface more often select the same books. Perhaps the
multistage interface encourages users to explore the collection in more di erent
ways than the baseline interface.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Users' actions in the di erent interfaces</title>
        <p>Each interface allows users to perform certain actions, with some overlap between
the available actions across the baseline interface and the three stages of the
multistage interface. The mean number of actions of each action type per user
is shown in Table 3, split between the focused task (top half) and the open task
(bottom half). Certain actions are only available in the multistage interface, like
show-layout, which corresponds to a switch between stages in the interface, and
browse, which allows users to browse through the Amazon hierarchy of book
categories without providing a search box.</p>
        <p>The Table outlines some di erences between the baseline and multistage
interface. First of all, the users of the multistage interface utilize `paginate' more,
suggesting that the interface encourages users to explore a larger part of the
collection. On the other hand, they use fewer lters and queries (both available
in the search panel of the multistage interface). This is perhaps due to the fact
that the users had more elaborate options to explore (via the browse panel) and
to review results (via the book-bag panel) in the multistage interface, hence did
not have to rely on querying and ltering alone, as in the case of the baseline
interface. Finally, a di erence can be observed in the use of book metadata:
in the open task, participants view more book metadata using the multistage
interface than via the baseline interface. It is possible that participants are
triggered to check more books by the distinct functionality of the di erent panels
of the multistage interface, especially since the open task allows users to explore
freely. The same di erence cannot be observed for the focused task, however,
where users review slightly more metadata via the baseline interface than via
the multistage interface.
4.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>The use of interface panels in the multistage system</title>
        <p>In the case of the multistage interface, users had the ability to switch between
interface screens, or stages in the interface. In this section, we look at the time
spent in each screen, the number of switches between screens and the transition
probabilities for switching between screens.</p>
        <p>First of all, Table 4 shows the number of screens a participant viewed, as
the participant has the possibility to switch multiple times between the browse,
search and book-bag screens. For comparison purposes, we initially look at the
training task, in which we expect users to test out all three interface screens.
This is re ected in the table, as the mean and median of switches is close to
three. Of the 95 users of the multistage interface, 75 (79%) went through the
screens in order|i.e. from browse to search and nally to book-bag to nish
the training task. During the open task, participants may use more di erent
interface units. This is re ected in the mean number of screens used, and a
higher standard deviation in the second row of the table. Finally, in the focused
task, participants have to carry out ve sub-tasks, so we would expect that
participants switch interface units more frequently. Indeed, the focused task
results in a substantially higher median, mean and standard deviation. Hence,
we have found evidence for frequent switching between interface units.</p>
        <p>Next, we analyze the probabilities of switching between these interface units.
Figure 1 shows the transition probabilities for the focused task, i.e. the
probability that a user switches from a certain interface screen to another (or ends the
task). Participants frequently switch between the browse, search and book-bag
screens, and most commonly end the task from the book-bag screen. The higher
probability of switching between search and explore in the focused task can be
explained again by the task properties: having ve sub-tasks to complete, the
participants frequently move from one interface screen to another. The gure
also shows the transition probabilities for the open task. Here, we see that in the
open task, users more frequently switch between the browse and the book-bag
stage, and less often from the browse to the search stage. Hence, the browse
screen may be more important than the search screen in the open task, while
the search screen is more important in the focused task.</p>
        <p>To derive more insights into the importance of each screen in both focused
and open tasks, we look at the time spent in each interface unit. We measured
the time spent in an interface screen after initiating the task or switching to the
browse, search or book-bag screen. Table 5 provides a summary. It shows that
for the focused task the search screen is used for the longest total duration by
far, re ected in the highest median and mean duration, followed by the book-bag
and browse screen, which are used substantially shorter. The open task features
a di erent emphasis: the browse screen is used most frequently, as shown by the
median and mean duration. The search and book-bag screen are comparatively
used less often, both having similar values, but the di erence for the median is
less clear as in the case of the focused task. Finally, the standard deviation of
total usage duration of the browse screen is a lot higher. Even without taking
one outlier into account (most likely caused by a speci c user's long period of
inactivity), the variation in the use of this screen in the open task is the highest.</p>
        <p>Summarizing, we found evidence for frequent switching between interface
units, especially in the focused task. The willingness of users to switch between
screens does provide positive indications for the usefulness of novel multistage
interfaces and the enrichment of existing search options.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Participants' perceptions of multistage interfaces</title>
      <p>The User Engagement Scale (UES) [7] is a multidimensional scale that contains
six sub-scales: Aesthetics, Novelty, Felt Involvement, Focused Attention,
Perceived Usability, and Endurability. Its purpose is to assist researchers in
reaching a holistic understanding of users perceptions of technology. According to
O'Brien and Toms [7] the scale seeks to measure multiple aspects of engagement
and understand their relationships to one another.</p>
      <p>To prepare the data for analysis some items were reverse coded. An initial
examination of the data showed that there were no missing variables for any
of the items. The 31 items were comprised into the 6 sub-scales. Table 6 shows
the sub-scale means with both interfaces. The multistage interface seems more
engaging in all sub-scales. However, we tested the di erences using the
MannWhitney test and found that only the di erences for Endurability (p.= 0.006)
and Felt Involvement (p. = 0.041) were statistically signi cant.</p>
      <p>We also grouped the participants in three age groups (Group 1: age range
18-25, N=80; Group: 2: age range 26-35, N=80; Group: age range 35+, N=40)
and examined whether the engagement varied between the groups, but the
differences were not signi cant (Krutska-Wallis test). Also engagement did not vary
signi cantly whether the open task or the focused task was performed rst.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>The analyses performed in this paper lead to various insights into the value of
multistage interfaces in the context of complex book search tasks. First of all, the
task duration in the multistage interface is substantially higher for both focused
and open-ended tasks, suggesting that users are more involved in searching for
books. This is also re ected in the signi cant di erences for user engagement in
the multistage interface, in terms of Endurability and Felt Involvement. Second,
users viewed more result pages and collected more books in the multistage
interface as compared to the baseline interface. In addition, the collected books
had less overlap between participants. Hence, the longer task time also seems to
result in a larger and more varied set of collected books. Finally, the frequent
screen switching in both tasks suggests that the di erent screens encourage
different types of activities. This can also be seen in the time spent in each screen:
the browse screen is used more in the open-ended task, while the search screen
is used for a longer total duration in the focused task.</p>
      <p>The results suggest that the multistage interface encourages users to explore
the collection in more di erent ways than the baseline interface. Further
analysis is needed, however, since there may be personal di erences between users,
for example in terms of common patterns of interactions with the multistage
interface. We plan to analyze these aspects in future work. Similar to [4], we also
plan to look at the di erences at di erent points in time of the tasks.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This research was supported by the Netherlands Organization for Scienti c
Research (WebART project, NWO CATCH # 640.005.001).
8138 of Lecture Notes in Computer Science, pages 17{28. Springer Berlin
Heidelberg, 2013. ISBN 978-3-642-40801-4. doi: 10.1007/978-3-642-40802-1
3.
2. M. Hall, S. Katsaris, and E. Toms. A pluggable interactive ir evaluation
work-bench. In European Workshop on Human-Computer Interaction and
Information Retrieval, pages 35{38, 2013.
3. M. Hall, H. Huurdeman, M. Koolen, M. Skov, and D. Walsh. Overview of the
INEX 2014 interactive social book search track. In L. Cappellato, N. Ferro,
M. Halvey, and W. Kraaij, editors, CLEF 2014 Labs and Workshops,
Notebook Papers, CEUR Workshop Proceedings (CEUR-WS.org), 2014.
4. H. C. Huurdeman and J. Kamps. From Multistage Information-seeking
Models to Multistage Search Systems. In Proceedings of the 5th Information
Interaction in Context Symposium, IIiX '14, pages 145{154, New York, NY,
USA, 2014. ACM. ISBN 978-1-4503-2976-7. doi: 10.1145/2637002.2637020.</p>
      <p>
        URL http://doi.acm.org/10.1145/2637002.2637020.
5. C. C. Kuhlthau. Inside the search process: Information
seeking from the user's perspective. Journal of the American
Society for Information Science, 42(5):361{371, 1991. ISSN
10974571. doi: 10.1002/(SICI)1097-4571(199106)42:5h361::AID-ASI6i3.0.CO;
2-#. URL http://dx.doi.org/10.1002/(SICI)1097-4571(199106)42:
5&lt;361::AID-ASI6&gt;3.0.CO;2-#.
6. R. Nordlie and N. Pharo. Seven years of inex interactive retrieval
experiments - lessons and challenges. In T. Catarci, P. Forner, D. Hiemstra,
A. Pen~as, and G. Santucci, editors, CLEF, volume 7488 of Lecture Notes in
Computer Science, pages 13{23. Springer, 2012. ISBN 978-3-642-33246-3.
7. H. L. O'Brien and E. G. Toms. The development and evaluation of a survey
to measure user engagement. Journal of the American Society for
Information Science and Technology, 61(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ):50{69, 2010.
8. V. Petras, T. Bogers, E. Toms, M. Hall, J. Savoy, P. Malak, A. Pawowski,
N. Ferro, and I. Masiero. Cultural heritage in clef (chic) 2013. In P. Forner,
H. Muller, R. Paredes, P. Rosso, and B. Stein, editors, Information Access
Evaluation. Multilinguality, Multimodality, and Visualization, volume 8138
of Lecture Notes in Computer Science, pages 192{211. Springer Berlin
Heidelberg, 2013. ISBN 978-3-642-40801-4. doi: 10.1007/978-3-642-40802-1 23.
9. E. Toms and M. M. Hall. The chic interactive task (chici) at
clef2013.
http://www.clef-initiative.eu/documents/71612/1713e643-27c34d76-9a6f-926cdb1db0f4, 2013.
10. P. Vakkari. A theory of the task-based information retrieval process: a
summary and generalisation of a longitudinal study. Journal of documentation,
57(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ):44{60, 2001.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Hall</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Toms</surname>
          </string-name>
          .
          <article-title>Building a common framework for iir evaluation</article-title>
          . In P. Forner, H. Muller, R. Paredes,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , and B. Stein, editors,
          <source>Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Visualization, volume
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>