<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Edge Hill Contribution to the INEX Interactive Social Book Search Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Walsh</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Hall</string-name>
          <email>Mark.Hallg@edgehill.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>David.Walsh</institution>
          ,
          <addr-line>Mark.Hall</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Edge Hill University</institution>
          ,
          <addr-line>St Helens Road, Ormskirk, L39 4QP</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <fpage>549</fpage>
      <lpage>556</lpage>
      <abstract>
        <p>In our contribution we use log-analysis to investigate whether participants in the INEX Interactive Social Book Search are able to use the new multi-stage interface and whether it provides any bene ts over the traditional IR baseline interface. Our initial results show that participants are able to successfully use the new multi-stage interface, with no signi cant learning e ects. Additionally, for the non-goal task, the multi-stage interface actually enables the participants to collect more books than when using the baseline interface.</p>
      </abstract>
      <kwd-group>
        <kwd>human computer information retrieval</kwd>
        <kwd>user study</kwd>
        <kwd>log analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The CLEF1 INEX2 track's Interactive Social Book Search task gathered data
from users using one of two interfaces to complete two tasks. The baseline
interface implemented a standard Information Retrieval (IR) interface [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] consisting
of a search box, a search result list, and an individual item display. The second
interface (multi-stage) attempted an implementation of Kuhlthau's multi-stage
search process[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], ltered through Vakkari's simpli cation of the model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Two
tasks were tested, the rst an non-goal task where participants were asked to
look for any book they might nd interesting, and a goal-oriented task where
participants were asked to nd books for a given topic (\laymen books on
mathematics and physics"). Each participant completed both tasks in one of the two
interfaces. Task order was automatically balanced to avoid ordering bias.
      </p>
      <p>We investigated the following three research questions:
1. RQ1: Does the multi-stage interface enable the participants to explore and
nd a larger number of books?
2. RQ2: Does the multi-stage interface have an additional learning time?
3. RQ3: Do participants make use of all three stages in the multi-stage
multistage interface?
1 Conference and Labs of the Evaluation Forum
2 INitiative for the Evaluation of XML retrieval</p>
    </sec>
    <sec id="sec-2">
      <title>Time Spent in the System</title>
      <p>
        The rst analysis focused on how long the participants spent using the system in
order to determine whether there were any di erences between the two systems
and tasks. The experiment system [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] automatically measured the time taken on
the non-goal and goal-oriented tasks for every participant and the main
analysis is based on this data. For the three stages implemented in the multi-stage
interface, the log data acquired by the IR system [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] was processed to determine
how long each participant spent using the each of the three stages (\explore",
\focus", and \re ne").
      </p>
      <p>The rst step in the analysis was to determine if the task ordering impacted
the time spent on either of the tasks. Wilcox signed rank tests were used to
compare the task times for all interface and task combinations, showing no signi cant
di erences in task duration for any of the combinations. For the multi-stage
interface, the time spent in each of the three stages was also compared using a
Wilcoxon rank-sum test for both ordering conditions, and also showed no
significant di erences in times in the two stages. For the remainder of the analysis,
the task order can thus be ignored and the times for the two order conditions
aggregated.</p>
      <p>Table 1 shows the task times for the two interfaces and tasks. The data
seems to indicate that participants are faster with the multi-stage interface for
both tasks and that within the interfaces, participants are faster to complete
the non-goal task than the goal-oriented task. However, for neither of these
conditions does a Wilcoxon rank-sum test show signi cant di erences. Thus for
the remaining analysis presented here, we can assume that any di erences in
participant performance are due to the task or interface and not due to the time
the participants spent on the task or interface.
In the multi-stage interface participants were able to switch between three stages
(\explore", \focus", and \re ne"). Table 2 shows the time spent in each of the
three stages for the two tasks. Wilcoxon rank-sum tests were used to test for
ordering e ects. There are no ordering e ects for the time spent in the explore
and focus stages, but there is an ordering e ect in the re ne stage. For the
nongoal task, the time spent in the re ne stage is longer, if it is the second task
(p = 0:012). No ordering e ect was shown for the goal-oriented task.</p>
      <p>The times shown in Table 2 follow similar patterns for both the non-goal
and closed tasks. Participants spent slightly less than a minute using the explore
stage, and then spent between one and a half and two minutes on the focus
stage. Only a small fraction of participants used the re ne stage at all and those
that did, did so only very brie y.</p>
      <p>Considering RQ3, participants obviously do not use the nal re ne stage,
either because they did not notice the stage in the user interface or because
the label \Re ne" did not clearly state what functionality would be available.
Without looking at the participants qualitative responses it is impossible to
determine the cause. However, the use pattern for the rst two stages is as
expected, with participants rst spending time in the explore stage gaining an
overview and then using the focus stage.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Books Collected</title>
      <p>To investigate RQ1 we looked at the number of books participants added to
their book-bag and also how quickly they added the rst book. For RQ2 we also
investigated which of the three stages participants added books from.
3.1</p>
      <p>Total Number of Books Added
The total number of books each participant added to their book-bag was
determine using a manual analysis of the log-data. For the baseline interface, the
number of books added to the book-bag was counted and any books that were
subsequently removed from the book-bag subtracted from that count. For the
multi-stage interface, the same process was applied, but book counts were
separated according to which of the stages the books were added from.</p>
      <p>The resulting data-set was checked for task ordering e ects using Wilcoxon
rank sum tests and no signi cant ordering e ects were found for any of the
interface / task combinations. In the multi-stage task, the same checks were
applied to the more detailed data and only for the explore stage in the
goaloriented task, was there a signi cant ordering bias. If the goal-oriented task was
the second task, then signi cantly fewer books were added to the book-bag in
the explore stage (Wilcoxon signed rank, p = 0:035). As there were no overall
ordering e ects, the further analysis did not take task ordering into account.</p>
      <p>Table 3 shows that signi cantly more books were added in the non-goal task
using the multi-stage interface than using the baseline interface (Fig. 1, Wilcoxon
signed rank test, p = 0:011). No such e ect is visible in the goal-oriented task.
This seems to indicate that the multi-stage interface provides signi cant bene t
to the user when they do not yet have an explicit goal that they are searching
for. At the same time, the multi-stage interface does not impact the performance
when the user has an explicit goal in mind.
To investigate RQ1 further we used the dataset from the previous section, but
now looked in detail at the number of books added in the three stages of the
multi-stage interface (Tab. 4). Task ordering e ects were investigated and no
signi cant e ects were found.</p>
      <p>Table 4 clearly shows that the re ne stage was not used to add any books,
which is in line with the timing results that showed that the re ne stage was
essentially not used. Interestingly, although the participants spent more time in
the focus stage, they collected more items in the initial explore stage. While the
data seems to indicate that in the goal-oriented task participants collected more
books in the explore stage, the e ect is not signi cant.
To investigate RQ2, we analysed how quickly participants added their rst book
in each task. The log was manually analysed and the time between the session
start and when the time at which the rst book was added to the book-bag
determined.</p>
      <p>Table 5 shows the median times to collect their rst book. While it looks as
if the modern interface enables the participant to nd the rst book faster, the
di erence is not statistically signi cant.
The nal analysis looked at the interaction patterns, using a user-interaction
bi-gram analysis. To create the interaction bi-gram distributions needed for the
analysis, the log was processed in the following steps:
1. Generate interaction string { in the initial step each user-system
interaction was mapped to a single letter. Using this mapping, for each participant
and each of the participant's tasks a string representation of their
interactions with the system was generated. Repeats of a single letter were reduced
to a single letter;
2. Create participant pattern distribution { based on the interaction
strings, all bi-grams were determined and bi-gram frequency distributions
calculated;
3. Aggregate distributions { the participants' bi-gram distributions were
aggregated into interface and task bi-gram distributions;
4. Filter distributions { the interface and task bi-gram distributions were
ltered. All bi-grams that occurred fewer than three times were aggregated
into a single value. This ensures that a potentially large number of interaction
patterns that only occurred once or twice do not skew the results, while at
the same time not completely loosing that data.</p>
      <p>Before analysing the data in any more depth, potential ordering e ects were
investigated and only for the goal-oriented task with the multi-stage interface
is there a signi cant di erence in the interaction pattern distributions ( 2 test,
p = 0:047). As the signi cance is border-line and there is no signi cant di erences
in any of the other metrics tested, for the purpose of this analysis, the in uence
of task order will be ignored.</p>
      <p>Comparing the two interfaces shows a signi cant di erence in the bi-gram
distributions between the baseline and multi-stage interfaces on the goal-oriented
task ( 2 test, p = 0:036). Looking at the bi-gram distribution (Tab. 6) shows
that the main di erence is at which point participants added books to their
book-bag. In the baseline interface this primarily happened after the participants
had viewed the book's detail (IA), while in the multi-stage interface it happens
directly from the search results list (QI ). The behaviour after adding an book
to the book-bag is also di erent. In the baseline interface, the next action is to
run another query (IQ ), while in the multi-stage interface it is to view another
book (AI ).</p>
      <p>Comparing the two tasks within both interfaces, shows a signi cant di erence
between the non-goal and goal-oriented tasks in the multi-stage interface ( 2,
p &lt; 0:001), but none in the baseline interface. From the bi-gram distribution
of the multi-stage interface (Tab. 7), three di erences stand out. In the
goaloriented task, participants made more use of the pagination functionality to
see more items (PI ) and also viewed more items after selecting a search facet
(FI ). In the non-goal task, participants more frequently viewed di erent bits of
meta-data after viewing an item (MI ).</p>
      <p>The di erence is in line with what would be expected, due to the task
differences. In the goal-oriented task, participants use the facetting and pagination
functionality to dig into the results, a pattern that is not so relevent when the
task is non-goal. At the same time, in the non-goal task, participants interact
more with the books' meta-data, as the participants use the meta-data to develop
the search goal.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>In conclusion the goal of the initial log analysis was to determine how participants
used the multi-stage interface. The initial question was whether there would be
a learning impact into the participants' peformance. The results clearly show
that participants are able to use the new multi-stage interface just as well as
the baseline interface and that there are no learning e ects. For the non-goal
task, the multi-stage interface even outperforms the baseline interface. Clearly,
the initial explore stage, designed to support open-ended exploration, enables
the user to explore better and thus collect more books.</p>
      <p>Finally, both the considering the number of books collected and the time
spent in the three stages of the multi-stage interface, participants clearly only
make use of the rst two stages. For the rst two stages, the behaviour is as
expected, with more time spent in the second focus stage, compared to the rst
explore stage. Interestingly, the majority of books were collected using the rst
explore stage, a result that needs more analysis. The use of the nal re ne stage
requires further analysis, as it is clearly not used much and the reasons for this
need to be investigated.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katsaris</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toms</surname>
          </string-name>
          , E.:
          <article-title>A pluggable interactive ir evaluation workbench</article-title>
          .
          <source>In: European Workshop on Human-Computer Interaction and Information Retrieval</source>
          . pp.
          <volume>35</volume>
          {
          <issue>38</issue>
          (
          <year>2013</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1033</volume>
          /paper4.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toms</surname>
          </string-name>
          , E.:
          <article-title>Building a common framework for iir evaluation</article-title>
          .
          <source>In: CLEF 2013 - Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Visualization. pp.
          <volume>17</volume>
          {
          <issue>28</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Search User Interfaces</article-title>
          . Cambridge University Press (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kuhlthau</surname>
          </string-name>
          , C.C.
          <article-title>: Inside the search process: Information seeking from the user's perspective</article-title>
          .
          <source>JASIS</source>
          <volume>42</volume>
          (
          <issue>5</issue>
          ),
          <volume>361</volume>
          {
          <fpage>371</fpage>
          (
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Vakkari</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A theory of the task-based information retrieval process: a summary and generalisation of a longitudinal study</article-title>
          .
          <source>Journal of documentation 57(1)</source>
          ,
          <volume>44</volume>
          {
          <fpage>60</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>