<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop on Supporting Complex Search Tasks, March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Lessons Learned from the CHiC and SBS Interactive Tracks: A Wishlist for Interactive IR Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Toine Bogers</string-name>
          <email>toine@hum.aau.dk</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Gäde</string-name>
          <email>maria.gaede@ibi.hu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark M. Hall</string-name>
          <email>mark.hall@edgehill.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vivien Petras</string-name>
          <email>vivien.petras@ibi.hu-berlin.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mette Skov</string-name>
          <email>skov@hum.aau.dk</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Berlin School of Library and, Information Science, Humboldt-Universität zu Berlin</institution>
          ,
          <addr-line>Dorotheenstr. 26, 10117, Berlin</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Berlin School of Library and, Information Science, Humboldt-Universität zu Berlin</institution>
          ,
          <addr-line>Dorotheenstr. 26, 10117, Berlin</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Computer Science, Edge Hill University</institution>
          ,
          <addr-line>St Helens Road, L39 4QP</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>E-learning Lab, Department of Communication &amp;</institution>
          ,
          <addr-line>Psychology</addr-line>
          ,
          <institution>Aalborg University</institution>
          ,
          <addr-line>Rendsburggade 14, 9000, Aalborg</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Science and Information Studies, Department of Communication &amp;</institution>
          ,
          <addr-line>Psychology</addr-line>
          ,
          <institution>Aalborg University Copenhagen</institution>
          ,
          <addr-line>A.C. Meyers Vaenge 15, 2300, Copenhagen</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <volume>11</volume>
      <issue>2017</issue>
      <abstract>
        <p>Over the course of the past two decades, the Interactive Tracks at TREC and INEX have contributed greatly to our knowledge of how to run an interactive IR evaluation campaign. In this position paper, we add to this body of knowledge by taking stock of our own experiences and challenges in organizing the CHIC and SBS Interactive Tracks from 2013 to 2016 in the form of a list of important properties of any future IIR evaluation campaigns.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Evaluation has always played an important role in IR research.
Traditional IR evaluation campaigns have typically focused on
measuring system performance in controlled environments, but over
the past two decades, IR research has expanded to include a more
user-based perspective, which considers the interactive nature of
information seeking behavior and the individual contexts
surrounding it [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. However, IIR evaluation campaigns remain relatively
ad-hoc and disjointed and still face signi cant challenges [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        The rst large-scale IIR evaluation campaign was the TREC
Interactive Track (1997-2002) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], followed by the INEX Interactive
Track (2004-2010) [
        <xref ref-type="bibr" rid="ref18 ref20">18, 20</xref>
        ]. These campaigns taught the
community important lessons, not the least of which was the di culty of
maintaining a stable evaluation framework, because of the lessons
learned after each iteration. Participant recruitment was often
CHIIR 2017 Workshop on Supporting Complex Search Tasks, Oslo, Norway.
Copyright for the individual papers remains with the authors. Copying permitted
for private and academic purposes. This volume is published and copyrighted by its
editors. Published on CEUR-WS, Volume 1798, http://ceur-ws.org/Vol-1798/.
di cult, resulting in small convenience samples of students and
colleagues. The campaigns also highlighted the in uence of GUI
design on interaction behavior and engagement with the study.
The Cultural Heritage in CLEF (CHiC) Interactive Task took up the
mantle of IIR evaluation again in 2013, which continued with the
interactive Social Book Search (SBS) task (2014-2016) [
        <xref ref-type="bibr" rid="ref10 ref22 ref8 ref9">8–10, 22</xref>
        ].
      </p>
      <p>In this position paper, we take stock of our experiences in the
CLEF Interactive Tracks and the lessons learned from previous
IIR evaluation campaigns to list important properties for future
campaigns. Complex search tasks—the theme of this workshop—are
at the heart of information seeking studies. We hope to contribute
to the development of an IIR evaluation framework that mirrors
the complex nature of information seeking behavior. We make no
claims about the completeness of this list, but instead see it as a
starting point for iterative re nement and improvement.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>PROPERTIES OF IIR EVALUATION</title>
    </sec>
    <sec id="sec-3">
      <title>Continuity</title>
      <p>
        The success of any IIR evaluation campaign depends rst and
foremost on getting a sustainable number of researchers to participate.
The ideal IIR evaluation campaign should run continuously over
a multi-year period. NewsReel [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] is an example of a campaign
that has moved from a rigid campaign schedule to a continuous
campaign, running throughout the year.
      </p>
      <p>A continuous campaign has several advantages, including
increased researcher participation. Our experience shows that a
common barrier to participation is tting it into a university’s schedule,
both in terms of teaching obligations as well as being able to use
the evaluation campaign as a teaching tool. Clashes and
misalignment with these professional obligations was the most commonly
stated reason for not participating in the di erent CLEF Interactive
Tracks. A continuous campaign would also aid in participant
recruitment as students or other participant groups are not always
available throughout the year. System maintenance could easily be
accounted for by short, xed ‘downtime’ periods.</p>
      <p>Workshops could then be organized to highlight interesting
research and spark discussion. We suspect that the 12-month TREC
or CLEF cycles are too short to achieve meaningful results with IIR
studies, to attract enough participants, and account for the longer
time it takes to run user studies. Data collected up to a particular
point could be presented at workshops in 18-24 month intervals—
ideally attached to di erent conferences to ensure a varied potential
audience. Naturally, it should be possible at any given point in time
to extract the collected data for a particular research group that is
participating, for instance to support its use in teaching.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Complexity</title>
      <p>
        In real life, searching for information takes place in a broader
context. IIR studies have long tried to capture and represent this context
through instruments such as simulated work tasks [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. However,
we believe that future IIR evaluation ought to broaden its scope to
the entire information seeking process, both individual and
collaborative [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Searching is only one aspect of information behavior
and is often combined with browsing, exploration or interaction
with a recommender system. Moreover, information behavior
often takes place across and between di erent devices (desktop vs.
smartphone) and modalities (digital vs. paper).
      </p>
      <p>IIR studies should be representative of the real-life variety in
users types, domains, and system designs. A comprehensive IIR
research program would investigate domain-speci c and professional
information seeking alongside casual leisure searching, combining
specialized systems as well as general web search for resources
such as professional documents, books, music, traveling etc. Not all
contexts may be studied within the same IIR evaluation campaign,
but some could be combined or contrasted. The CHiC and SBS
Interactive Tracks aimed to do this through variety in the simulated
work tasks that were used, but we believe this was only a rst step
in the right direction.</p>
      <p>
        Related phenomena such as serendipity should also be
incorporated into the IIR evaluation wherever possible, drawing upon the
scant lessons that have been learned about evaluating this in a
laboratory setting [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In general, the complex nature of information
behavior should be better represented in IIR evaluation campaigns.
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Flexibility</title>
      <p>While the core of an IIR evaluation campaign should be as stable as
possible to enable comparisons over time, we wish to advocate for
as much exibility as possible around this stable core to allow for
di erent research questions and approaches within a common IIR
framework.</p>
      <p>When it comes to individual research contexts and interests of
the researchers, experiments often need to be exible with respect
to gathering of background and contextual information. Di erent
research teams may want to focus on di erent aspects of search
behavior or interfaces and may therefore need to ask di erent
preand post-task questions of their participants. Ideally, the system
would let researchers easily plug in and swap out di erent
questions depending on the research team’s interests. Ideally, such
question sets could then be stored in the system and used in future
experiments by other teams.</p>
      <p>
        Another example is the exibility in choosing di erent search
interfaces to study the e ects of the GUI on information seeking
behavior. This was used to great e ect in the 2015 and 2016 SBS
Interactive Tracks [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ] to examine how di erent interfaces can
support the di erent search stages.
      </p>
      <p>Increasing exibility is not without its challenges however.
Exposing di erent participant groups to di erent sets (and sizes) of
pre-/post-task questions or interfaces may lead to di erent levels
of fatigue or learning e ects, which could in uence the integrity of
the collected data and any future comparisons. The optimal
tradeo between stability and exibility is far from a solved problem.
2.4</p>
    </sec>
    <sec id="sec-6">
      <title>Realism</title>
      <p>Another essential property of user-based IR evaluation is realism,
which applies to di erent aspects of the evaluation process. One
is the use of a document collection that is realistic in size and
content variety. It ensures that participants do not become frustrated,
because they cannot nd relevant documents, which could a ect
search behavior as well as engagement.</p>
      <p>Ensuring realism also includes participant sampling. Ideally, an
IIR evaluation campaign recruits participants that are a realistic
representation of the general target population to avoid the
introduction of biases [7, p. 241]. Future IIR studies should have a high
variety of participants rather than focusing on the academic sector.
The in uence of the participants’ motivation to be part of an IIR
study also needs to be taken into account.</p>
      <p>
        Another element is the use of realistic representations of
realworld information needs. Typically, simulated work tasks [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] are
used to support the search behavior of participants in the form of
cover stories that include extra context about the background task.
Their formulation based on “the user’s own personal information
need or in relation to a set of simulated needs” [6, p. 229] is essential
for using them to their full potential [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4–6</xref>
        ].
      </p>
      <p>
        The question of how best to generate such realistic information
needs remains largely unanswered. One possibility is to use search
log analysis to identify popular queries and generate a cover story
around them. However, such cover stories would always be a
bestguess justi cation. The 2016 SBS Interactive Track [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] was the
rst edition of the CLEF Interactive Tracks where we successfully
used real-world examples of complex information needs taken from
the LibraryThing forums as additional, optional work tasks. These
tasks were rich in variety and detail, and this sentiment was also
expressed by our participants. This does yield two new questions:
(1) whether participants would be able to tell the di erence between
simulated and real-world work tasks, and (2) whether any resulting
di erences in information seeking behavior would observable and
meaningful.
      </p>
      <p>In addition to realism in the information needs, the evaluation
data set also should represent a realistic and engaging scenario. The
CHiC data set was based on an extract from Europeana and
demonstrated a wide range of topics, however the individual items in the
data-set were often sparse in their information and primarily
metadata based. In the SBS tasks, the data-set used was based on data
aggregated from Amazon, LibraryThing, and the British Library.
This provided a realistic and detailed data-set. However, over the
duration of the SBS tasks, the data became more and more outdated
and users often tried to search for newer books that the data-set
did not yet contain. Thus how to have a realistic, engaging, and
upto-date data-set, while at the same time maintaining comparability
across iterations of the evaluation is an open question.</p>
      <p>
        Simulated system prototypes have the advantage of full system
and content control as well as the possibility to test alternative
features and interactions. On the other side, the user experience
highly depends on the look and feel of the interfaces and interaction
design. Systems need to meet user needs and expectations in order
to reduce their in uence on behavior and learning e ects. Training
tasks can be used to create similar requirements for participants as
well as familiarity with unknown systems [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
2.5
      </p>
    </sec>
    <sec id="sec-7">
      <title>Measurability</title>
      <p>IIR studies use a variety of quantitative measures, but also
qualitative descriptions of information seeking behavior. Both
quantitative and qualitative approaches for result representations in this
research are rarely comparable—not only because of their di erent
study and audience foci, but also because the reported measures
and descriptions are not based on agreed-upon standards. The
comparability of results needs to be ensured for long-term-oriented IIR
evaluation campaigns.</p>
      <p>
        While the majority of studies gather a variety of performance,
interaction, and usability data, no framework for the interpretation
of these “numbers” exists [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Depending on expected behaviors
or user goals, the session duration, number of queries and results
viewed can be interpreted in di erent ways. For example, a long
session duration can be a sign of both a good or bad user experience.
Binary relevance assessments are rarely applicable when dealing
with real users. During an information seeking process, usefulness
and satisfaction seem to play a more important role but need to be
operationalized for IIR research [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Any successful IIR evaluation framework needs to include
standardized measures and description approaches, possibly with
agreedupon interpretations—comparable to the common e ectiveness
measures used in system-based IR evaluation.</p>
      <p>
        Full comparability also requires that IIR experiments can easily
be replicated by other researchers. A number of tools have been
proposed to enable this replication. Toms et al. (2004) designed
and implemented the WiIRE (Web-based Interactive Information
Retrieval) system [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], which was used in the TREC-11 Interactive
Track, an updated version of which was used at INEX 2007 [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
Hall et al. (2013) developed the PyIRE system [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which has been
used in the CHiC and SBS Interactive Tracks. Other systems include
SCAMP [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and CIRSE [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However, none of these have been
used outside of the initial IIR evaluation they were designed for,
primarily due to the complexity of setting up the systems and the
lack of documentation.
3
      </p>
    </sec>
    <sec id="sec-8">
      <title>WHAT NEXT?</title>
      <p>We believe that an essential next step for our wishlist of IIR
evaluation properties is to discuss them with the community. The list
should then be extended and elaborated upon using other related
work with a particular focus on how to approach any challenges.
Another important step is to investigate how di erent properties
interact and interfere with each other (e.g., exibility &amp; continuity).
A new IIR evaluation campaign and related researcher
community could then be designed around these properties to expand our
understanding of information (seeking) behavior.</p>
      <p>Another aspect of IIR evaluation for future work is the question
of real vs simulated users. As the usability and visual design of the
evaluation interface have a signi cant impact on the IIR evaluation
process, the question of how these factors should be integrated into
user simulation is one that will need further attention.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Nicholas J Belkin</surname>
            ,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Cole</surname>
          </string-name>
          , and Jingjing Liu.
          <year>2009</year>
          .
          <article-title>A Model for Evaluation of Interactive Information Retrieval</article-title>
          .
          <source>In Proceedings of the SIGIR 2009 Workshop on the Future of IR Evaluation. 7-8.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Ralf</given-names>
            <surname>Bierig</surname>
          </string-name>
          , Jacek Gwizdka, and
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Cole</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>A User-centered Experiment and Logging Framework for Interactive Information Retrieval</article-title>
          .
          <source>In Proceedings of the SIGIR 2009 Workshop on Understanding the User: Logging and Interpreting User Interactions in Information Search and Retrieval</source>
          . 8-
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Toine</given-names>
            <surname>Bogers</surname>
          </string-name>
          ,
          <source>Rune Rosenborg Rasmussen, and Louis Sebastian Bo Jensen</source>
          .
          <year>2013</year>
          .
          <article-title>Measuring Serendipity in the Lab: The E ects of Priming and Monitoring</article-title>
          .
          <source>In Proceedings of the iConference</source>
          <year>2013</year>
          .
          <fpage>703</fpage>
          -
          <lpage>706</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Pia</given-names>
            <surname>Borlund</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>The IIR Evaluation Model: A Framework for Evaluation of Interactive Information Retrieval Systems</article-title>
          .
          <source>Information Research 8</source>
          ,
          <issue>3</issue>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Pia</given-names>
            <surname>Borlund</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A Study of the Use of Simulated Work Task Situations in Interactive Information Retrieval Evaluations: A Meta-Evaluation</article-title>
          .
          <source>Journal of Documentation 72</source>
          ,
          <issue>3</issue>
          (
          <year>2016</year>
          ),
          <fpage>394</fpage>
          -
          <lpage>413</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Pia</given-names>
            <surname>Borlund</surname>
          </string-name>
          and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Ingwersen</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>The Development of a Method for the Evaluation of Interactive Information Retrieval Systems</article-title>
          .
          <source>Journal of Documentation 53</source>
          ,
          <issue>3</issue>
          (
          <year>1997</year>
          ),
          <fpage>225</fpage>
          -
          <lpage>250</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Donald</surname>
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Case</surname>
          </string-name>
          and
          <string-name>
            <surname>Lisa M. Given</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Looking for Information: A Survey of Research on Information Seeking</article-title>
          , Needs, and
          <string-name>
            <surname>Behavior</surname>
          </string-name>
          (4th ed.). Emerald Group Publishing, Bingley, UK.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Gäde</surname>
          </string-name>
          , Mark Michael Hall, Hugo C. Huurdeman, Jaap Kamps, Marijn Koolen, Mette Skov, Toine Bogers, and
          <string-name>
            <given-names>David</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Overview of the SBS 2016 Interactive Track</article-title>
          .
          <source>In Working Notes of the CLEF 2016 Conference (CEUR Workshop Proceedings)</source>
          , Krisztian Balog, Linda Cappellato,
          <source>Nicola Ferro, and Craig Macdonald (Eds.)</source>
          , Vol.
          <volume>1609</volume>
          . CEUR-WS.org,
          <volume>1024</volume>
          -
          <fpage>1038</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Gäde</surname>
          </string-name>
          , Mark Michael Hall, Hugo C. Huurdeman, Jaap Kamps, Marijn Koolen, Mette Skov, Elaine Toms, and
          <string-name>
            <given-names>David</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Overview of the SBS 2015 Interactive Track</article-title>
          .
          <source>In Working Notes of the CLEF 2015 Conference (CEUR Workshop Proceedings)</source>
          , Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>Gareth J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          , and Eric SanJuan (Eds.), Vol.
          <volume>1391</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Michael</surname>
          </string-name>
          <string-name>
            <surname>Hall</surname>
          </string-name>
          , Hugo C. Huurdeman, Marijn Koolen, Mette Skov, and
          <string-name>
            <given-names>David</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Overview of the INEX 2014 Interactive Social Book Search Track</article-title>
          .
          <source>In Working Notes of the CLEF 2014 Conference (CEUR Workshop Proceedings)</source>
          , Linda Cappellato, Nicola Ferro,
          <source>Martin Halvey, and Wessel Kraaij (Eds.)</source>
          , Vol.
          <volume>1180</volume>
          . CEUR-WS.org,
          <volume>480</volume>
          -
          <fpage>493</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Mark</surname>
            <given-names>M Hall</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Spyros</given-names>
            <surname>Katsaris</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Elaine</given-names>
            <surname>Toms</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A Pluggable Interactive IR Evaluation Work-bench</article-title>
          .
          <source>In European Workshop on Human-Computer Interaction and Information Retrieval</source>
          .
          <fpage>35</fpage>
          -
          <lpage>38</lpage>
          . http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1033</volume>
          /paper4.pdf
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Preben</surname>
            <given-names>Hansen</given-names>
          </string-name>
          , Chirag Shah, and
          <string-name>
            <surname>Claus-Peter Klas</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Collaborative Information Seeking</article-title>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Diane</given-names>
            <surname>Kelly</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Methods for Evaluating Interactive Information Retrieval Systems with Users</article-title>
          .
          <source>Foundations and Trends in Information Retrieval 3</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          (
          <year>2009</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>224</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Diane</given-names>
            <surname>Kelly and Cassidy R Sugimoto</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A Systematic Review of Interactive Information Retrieval Evaluation Studies,</article-title>
          <year>1967</year>
          -
          <fpage>2006</fpage>
          .
          <source>Journal of the American Society for Information Science and Technology 64</source>
          ,
          <issue>4</issue>
          (
          <year>2013</year>
          ),
          <fpage>745</fpage>
          -
          <lpage>770</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Benjamin</surname>
            <given-names>Kille</given-names>
          </string-name>
          , Andreas Lommatzsch, Gebrekirstos G. Gebremeskel, Frank Hopfgartner, Martha Larson, Jonas Seiler, Davide Malagoli, András Serény, Torben Brodt, and
          <string-name>
            <surname>Arjen P. de Vries</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Overview of NewsREEL'16: Multi-dimensional Evaluation of Real-Time Stream-Recommendation Algorithms</article-title>
          .
          <source>In CLEF '16: Pro(Eds.)</source>
          , Vol.
          <volume>9822</volume>
          . Springer,
          <fpage>311</fpage>
          -
          <lpage>331</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Marijn</surname>
            <given-names>Koolen</given-names>
          </string-name>
          , Toine Bogers, Maria Gäde,
          <string-name>
            <surname>Mark A</surname>
          </string-name>
          . Hall, Hugo C. Huurdeman, Jaap Kamps, Mette Skov, Elaine Toms, and
          <string-name>
            <given-names>David</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Overview of the CLEF 2015 Social Book Search Lab</article-title>
          .
          <source>In CLEF '15: Proceedings of the 6th International Conference of the CLEF Association (Lecture Notes in Computer Science)</source>
          , Josiane Mothe, Jacques Savoy, Jaap Kamps, Karen Pinel-Sauvagnat,
          <string-name>
            <given-names>Gareth J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <surname>Eric</surname>
            <given-names>SanJuan</given-names>
          </string-name>
          , Linda Cappellato, and Nicola Ferro (Eds.), Vol.
          <volume>9283</volume>
          . Springer,
          <fpage>545</fpage>
          -
          <lpage>564</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Over</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>The TREC Interactive Track: An Annotated Bibliography</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>37</volume>
          ,
          <issue>3</issue>
          (
          <year>2001</year>
          ),
          <fpage>369</fpage>
          -
          <lpage>381</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Nils</surname>
            <given-names>Pharo</given-names>
          </string-name>
          , Thomas Beckers, Ragnar Nordlie, and
          <string-name>
            <given-names>Norbert</given-names>
            <surname>Fuhr</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Overview of the INEX 2010 Interactive Track</article-title>
          .
          <source>In INEX '10: Proceedings of the Ninth International Workshop of the Initiative for the Evaluation of XML Retrieval</source>
          , Shlomo Geva, Jaap Kamps, Ralf Schenkel, and Andrew Trotman (Eds.). Springer, Berlin, Heidelberg,
          <fpage>227</fpage>
          -
          <lpage>235</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Gareth</given-names>
            <surname>Renaud</surname>
          </string-name>
          and
          <string-name>
            <given-names>Leif</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>SCAMP: A Tool for Conducting Interactive Information Retrieval Experiments</article-title>
          .
          <source>In IIiX '12: Proceedings of the 4th Information Interaction in Context Symposium. ACM</source>
          ,
          <volume>286</volume>
          -
          <fpage>289</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Anastasios</surname>
            <given-names>Tombros</given-names>
          </string-name>
          , Birger Larsen, and
          <string-name>
            <given-names>Saadia</given-names>
            <surname>Malik</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>The Interactive Track at INEX 2004</article-title>
          . In INEX '
          <volume>04</volume>
          :
          <article-title>Proceedings of the Third International Workshop of the Initiative for the Evaluation of XML Retrieval, Norbert Fuhr</article-title>
          , Mounia Lalmas, Saadia Malik, and Zoltán Szlávik (Eds.). Springer, Berlin, Heidelberg,
          <fpage>410</fpage>
          -
          <lpage>423</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Elaine</surname>
            <given-names>G Toms</given-names>
          </string-name>
          ,
          <article-title>Luanne Freund</article-title>
          , and
          <string-name>
            <given-names>Cara</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>WiIRE: The Web Interactive Information Retrieval Experimentation System Prototype</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>40</volume>
          ,
          <issue>4</issue>
          (
          <year>2004</year>
          ),
          <fpage>655</fpage>
          -
          <lpage>675</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Elaine</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Toms and Mark M. Hall</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The CHiC Interactive Task (CHiCi) at CLEF 2013</article-title>
          .
          <source>In Working Notes of the CLEF 2013 Conference (CEUR Workshop Proceedings)</source>
          , Pamela Forner, Roberto Navigli,
          <source>Dan Tu s, and Nicola Ferro (Eds.)</source>
          , Vol.
          <volume>1179</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Elaine</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Toms</surname>
          </string-name>
          ,
          <string-name>
            <surname>Heather O'Brien</surname>
          </string-name>
          , Tayze Mackenzie, Chris
          <string-name>
            <surname>Jordan</surname>
            , Luanne Freund, Sandra Toze, Emilie Dawe, and
            <given-names>Alexandra</given-names>
          </string-name>
          <string-name>
            <surname>Macnutt</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Task E ects on Interactive search: The Query Factor</article-title>
          . In Focused Access to XML Documents. Springer,
          <fpage>359</fpage>
          -
          <lpage>372</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>