<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Proposal for User-Focused Evaluation and Prediction of Information Seeking Process</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chirag Shah</string-name>
          <email>chirags@rutgers.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Communication &amp; Information (SC&amp;I) Rutgers University 4</institution>
          <addr-line>Huntington St, New Brunswick, NJ 08901</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>One of the ways IR systems help searchers is by predicting or assuming what could be useful for their information needs based on analyzing information objects (documents, queries) and finding other related objects that may be relevant. Such approaches often ignore the underlying search process of information seeking, thus forgoing opportunities for making process-based recommendations. To overcome this limitation, we are proposing a new approach that analyzes a searcher's current processes to forecast his likelihood of achieving a certain level of success in the future. Specifically, we propose a machine-learning based method to dynamically evaluate and predict search performance several time-steps ahead at each given time point of the search process during an exploratory search task. Our prediction method uses a collection of features extracted solely from the search process such as dwell time, query entropy and relevance judgment in order to evaluate whether it will lead to low or high performance in the future. Experiments that simulate the effects of switching search paths show a significant number of subpar search processes improving after the recommended switch. In effect, the work reported here provides a new framework for evaluating search processes and predicting search performance. Importantly, this approach is based on user processes, and independent of any IR system allowing for wider applicability that ranges from searching to recommendations.</p>
      </abstract>
      <kwd-group>
        <kwd>Exploratory search</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Performance prediction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.3: INFORMATION STORAGE AND RETRIEVAL H.3.3:
Information Search and Retrieval: Search process; H.3:
INFORMATION STORAGE AND RETRIEVAL H.3.4:
Systems and Software: Performance evaluation (efficiency and
effectiveness)
Presented at EuroHCIR2013.</p>
      <p>Copyright © 2013 for the individual papers by the papers’ authors.
Copying permitted only for private and academic purposes. This volume
is published and copyrighted by its editors.
1</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        IR evaluations are often concerned with explaining factors
relating to user or system performance after the search and
retrieval are conducted [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Most recommender systems,
however, operate with an objective to suggest objects that could
be useful to a user based on his/her or others’ past actions
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. We commenced our investigation by broadly asking
how we could take valuable lessons from both IR evaluations
and recommender systems to not only evaluate an ongoing
search process, but also predict how well it will unfold and
suggest a better path to the searcher if it is likely to
underperform. The motivation behind this investigation was
based on the following assumptions and realizations grounded in
the literature.
1.
2.
      </p>
      <p>
        The underlying rational processes involved in information
search are reflected in the actions users make while
searching. These actions include entering search queries,
skimming the results, as well as selecting and collecting
useful information [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref14">14</xref>
        ][
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        A searcher’s performance is a function of these actions
performed during a search episode [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ][
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>With these assumptions, we propose to quantify a search process
using various user actions, and use it for user performance
(henceforth, ‘search performance’ or ‘performance’) prediction
as well as search process recommendations.
2</p>
    </sec>
    <sec id="sec-3">
      <title>BACKGROUND</title>
      <p>Past research on predictive models that relates to the approach
we describe in this paper can be grouped into two main
categories: (1) behavioral studies and (2) IR approaches. In both
cases; however, the focus has been on end products instead of in
the process required to produce them.</p>
      <p>
        As far as the behavioral studies go, research has been conducted
to explore users models that help anticipating specific aspects of
the search process. One goal in this context has been the
determination of whether a search process will be completed in a
single or multiple sessions. For example, Agichtein et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
investigated different patterns that can be identified in tasks that
require multiple sessions. As a result, the authors devised an
algorithm capable of predicting whether users will continue or
abandon the task. Similar work is described in Diriye et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
which focuses on predicting and understanding of why and
when users abandon Web searches. To address this problem, the
authors studied features such as queries and interactions with
result pages. Based on this approach, the authors were able to
determine reasons for search abandonment such as accidental
causes (e.g. Web browser crashing), satisfaction levels, and
query suggestions, among others.
      </p>
      <p>
        There have been also attempts to understand past users'
behaviors in order to predict future ones in similar conditions.
For example, Adar et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] visually explored behavioral aspects
using large-scale datasets containing queries and other
information objects produced by users. The authors were able to
identify different behavioral patterns that seem to appear
consistently in different datasets. While not directly related to
performance prediction, this work focused on attributes of the
search process instead of in final products derived from it.
Research like the ones described above often relies on historic
data from large populations and the use of trend and seasonal
components, which are used to model long-term direction and
periodicity patterns of time-series [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. For example, some have
explored seasonal aspects in Web search (e.g. weekly, monthly,
or annual behaviors) that provides useful information to predict
and suggest queries [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        From an IR perspective, Radinski et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] explored models to
predict users’ behaviors in a population in order to improve
results from IR systems. The authors also developed a learning
algorithm capable of selecting an appropriate predictive model
depending on the situation and time. As described by the
authors, applications of this approach could go from click
predictions to query-URL predictions. In contrast to this
approach, our method presented in this paper considers both the
population trends and an individual user behavior.
      </p>
      <p>
        In a similar track, several works have been conducted on query
performance prediction, focusing on developing techniques that
help IR system to anticipate whether a query will be effective or
not to provide results that satisfy users’ needs [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][
        <xref ref-type="bibr" rid="ref10">10</xref>
        ][
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. For
example, Gao et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] found that features derived from search
results and interactions features offer better prediction results
than a prediction baseline defined in terms of query features.
Results from this study have direct implications to individual
users by aiding the auto evaluation process of IR systems.
In information search, users may be unaware of their individual
performance when solving an information search task. For
instance, Shah &amp; Marchionini [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] showed how lack of
awareness about different objects involved in searching (queries,
visited pages, bookmarks) could result in mistaken perception
about search performance during an exploratory search task.
Even if an IR system is highly effective, users may run into
multiple query formulation and evaluation of several pages
before finding what they need. This process, which can be
related to search strategies, implies effort and time that is
usually underestimated by the users themselves. In this sense,
instead of predicting end products (i.e., overall performance),
the approach we introduce in this paper is oriented toward
predictions at different times in order to increase the level of
awareness of users about their own search process. Similar to
weather forecast, this information could help users to be aware
of possible trends based on past and current behavior.
For a more recent discussion on IR evaluations and their
shortcomings, see [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. To the best of our knowledge, search
process performance prediction at different times from a user
perspective has not been explored. Similar approaches can be
found in weather and stock market studies. For example, using
machine learning approaches such as Support Vector Machine
(SVM), some models have been implemented to predict the
trends of two different daily stock price indices using NASDAQ
and Korean Stock prices [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ][
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In a similar fashion, our
approach is oriented to forecast users’ search performance
Nsteps ahead with the aim to aid their search process awareness
and performance trends.
      </p>
      <p>Unlike previous works in IR, we are not proposing to use time
series analyses or seasonal components of historic data. Instead,
we investigate predictive models based on machine learning
(ML) techniques; namely: SVM, logistic regression, and Naïve
Bayes which are trained over a set of features such as time,
number of queries, and page dwell time. In contrast to most IR
evaluations, our method focuses on user-processes. Also, unlike
most recommender systems, our approach could output
alternative strategies instead of similar/relevant products to help
the searcher. In essence, the work reported here takes several
lessons from tradition IR evaluations, recommender system
designs, and weather/stock forecasting to come up with a new
approach for evaluating and predicting search performance.
In the next section we provide a detailed description of our
method, feature selection, and the measures we used in order to
create ML-based predictive models.</p>
    </sec>
    <sec id="sec-4">
      <title>3 METHOD</title>
      <p>
        In order to analyze the search processes followed by different
users/teams, we assume that the underlying dynamics of the
search processes are expressed by a collection of activities that
take place from the beginning to the end of the search processes.
The first part of our method is a feature extraction step in which
we extract a wide array of features relating to webpages, queries
and snippets saved from the search processes for each unit of
time t. This step is performed in order to evaluate how well we
could use those features to capture the underlying dynamics
which would lead to recognizing whether a search process is
going to lead to high or low performance in the future time steps
at t+n (n=1,2,….,N), where N is the furthest time step.
The decision to include or exclude a feature was based on
literature (e.g., [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) as well as our past experience [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] with
representing and evaluating search objects and processes. Each
feature is extracted for each user or team, u, up to time t from
the search processes and they are explained in detail as follows.
•
•
•
•
•
      </p>
      <p>
        Total coverage (u,t): The total number of distinct
Webpages visited by a user (u) up to time t. This feature
captures the Webpage based activity performed by a user
and provides a measure to see how much distinct
information has been found by the user up to this time.
Useful coverage (u,t): The total number of distinct
webpages in which a user spent at least 30 seconds, up to
time t. This measure evaluates out of the total pages he/she
has visited how many of them were useful in finding
relevant information leading to satisfaction with their
context in completing the exploratory task [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][
        <xref ref-type="bibr" rid="ref22">22</xref>
        ][
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
Number of queries (u,t): Total number of unique queries
executed by a user up to time t. This feature implicitly
relates to how much effort and cognitive thinking a user has
put in to this task.
      </p>
      <p>Number of saved snippets (u,t): Total number of snippets
saved by user u up to time t. This measures the amount of
information that the user thought that might be relevant in
the future to complete the task and needed to be
remembered. In other words, this feature is an indication of
explicit relevance judgments made by the user.</p>
      <p>Length of Query (u,q,t): Length of each query(q) executed
by a user u based on the character count of the query up to
time t. This feature captures how the user imposed the
•
•
queries and how long they were at different times of the
search process.</p>
      <p>Number of tokens in each query (u,q,t): This is the count of
tokens/words in each query(q) executed by user u up to
time t. This query based measure takes into account how
specific a user was in defining the query. By inspecting the
datasets, we realized that queries with a less number of
tokens tend to get general results. On the other hand,
composed queries with multiple terms are related to more
specific searchers. We also observed that typically the users
started with general queries with few words at the
beginning of the search process but then went into more
detailed queries to find more specific information later. For
all these reasons, we found it to be useful to capture the
number of token used in a query.</p>
      <p>
        Query entropy (u,q,t): This measures the information
content in a given query (q), by finding the expected value
of information contained in a query. We used the widely
recognized notion of Shannon entropy [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] in Information
Theory to calculate the information content of a query. We
calculated the number of unique characters appearing in
each of the queries, which represent the observed counts of
the random variable. This was used as the input to Shannon
entropy calculation and we used to the
maximumlikelihood method to calculate the entropy. Query entropy
feature has been used in the past to predict goodness of a
query for making query expansion decision [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        The method used to assess the search performance of a user is
described below. We define a measure called Efficiency (u,t), for
each user u up to time t in order to predict whether a given
search process is going to yield in high/low performance in the
future We first define Effectiveness of user u up to time t as the
ratio of useful coverage and total coverage (both defined
earlier). A similar measure was used in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>Effectiveness(u,t) =</p>
      <sec id="sec-4-1">
        <title>Useful coverage(u,t)</title>
      </sec>
      <sec id="sec-4-2">
        <title>Total coverage(u,t)</title>
        <p>(1)
(2)
We then calculated Efficiency as defined in Equation 2.</p>
        <p>Efficiency(u,t) =</p>
        <sec id="sec-4-2-1">
          <title>Effectiveness(u, t)</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>NumberofQueries(u,t)</title>
          <p>In other words, Efficiency is defined as the Effectiveness
obtained per query, or how effective a query is in terms of
achieving a certain level of useful coverage.</p>
          <p>The performance for each user u at each time t was classified in
to the two classes; high performance and low performance based
on the following criteria:
Class = {
high ;if
low
;else</p>
          <p>Efficiency(u, t) ≥ Efficiency(u, t)
(3)
Using various user studies data available to us, we constructed
feature matrices which consist of all aforementioned features for
each minute of time t for all the users in each dataset, and
converted in to a long vector of features which we fed as the
input to the classification models used.1 The class labels were
generated as high/low performance at minute t+n based on the
1 In the interest of space and scope of work here, details of these
experiments have been omitted, but will be available for
discussion at the workshop.
above mentioned criteria and threshold and used as the output
class labels to be used in the n-step ahead prediction model. If a
class label at n-step ahead was correctly predicted based on the
features extracted up to time t from the classification model it
was considered as correctly classified and if not as misclassified.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4 EXPERIMENTS</title>
      <p>In order to evaluate whether users who are predicted to perform
at low performance in the future based on the current search
process, could benefit from this analysis to improve their search
process, we conducted some simple simulation analysis.
We considered the individual user search processes as a
collection of search paths, where each search path is defined as
the search process from the time a user issued a query up to the
time user issued another quite different query. This was found
out using generalized Levenshtein (edit) distance, which is a
commonly used distance metric for measuring the distance
between two character sequences. If the Levenshtein (edit)
distance between two subsequent queries were greater than 2
(assuming less than 2 was when there were changes in the
queries due to simple spelling mistakes or refining of the query),
we considered the search process from the former query to the
next query as a single search path.</p>
      <p>Following this method, we found the first search path of each
user and based on the features extracted up to the end of the first
search path, and based on the classification model learnt from
that corresponding n-step ahead prediction we predicted whether
the user is going to have low/high performance at the end of the
session. If the user was going to have low performance, then out
of the users who predicted to have high performance, we looked
at which high performing user has the lowest Levenshtein (edit)
distance between the queries issued by low performing user
within the first search path and considered it as a pair of users,
whom we are going to use in the simulation. Then, for each low
performing user and high performing user that was matched, we
switched the search process of low performing user at the end of
the first search path with the high performing user’s search path
up to t=T minutes, where T is the total number of minutes for a
session. Then we evaluated by switching the search process
early during the overall process whether it would benefit each
low performing user to improve their performance. We found
that we were able to move most of the underperforming search
processes to higher performance by early detection and
switching, while keeping the higher performing processes
unharmed.</p>
      <p>These simulations provide verification that by realizing early
during the search process whether a user is going to perform
well or not, one could recommend better search
processes/strategies for that user which would lead to uplifting
the search performance of a previously destined to low
performing user.</p>
    </sec>
    <sec id="sec-6">
      <title>5 CONCLUSION</title>
      <p>When it comes to prediction, information retrieval and filtering
systems are primarily focused on objects while assessing what
and if something could help the users. These approaches are
often system-dependent even though the process of information
seeking is usually user-specific. Personalization and
recommendations are frequently exercised as methods to address
user-specific IR and filtering, but still limited to comparing and
recommending objects, not focusing on underlying IR processes
that are carried out by the searchers. We presented a new
approach to address these shortcomings. We began by asking
whether we could model a user’s search process based on the
actions he/she is performing during an exploratory search task
and forecast how well that process will do in the future. This
was based on a realization that an information seeker’s search
goal/task can be mapped out as a series of actions, and that a
sequence of actions or choices the searcher makes, and
especially the search path he/she takes, affects how well he/she
will do. Thus, in contrast to approaches that measure the
goodness of search products (e.g., documents, queries) as a way
to evaluate the overall search effectiveness, we measured the
likelihood of an existing search process to produce good results.
Here we presented simulations to demonstrate what could
happen if one can make process-based predictions, but one could
develop an actual recommender system using the proposed
method. Another potential application of such prediction-based
method would be to use such approach in IR systems to provide
the awareness to users how their future performance will be
based on the current/past search process. The system could
identify that a user will have low performance if, he continues
this manner at an early stage of the process, and what could be
done to provide suggestions to improve overall performance.
Given that the proposed technique is independent of any specific
kind of system, and solely focused on user-based processes, it
will presumably be easy to apply it to a variety of IR systems
and situations irrespective of retrieval, ranking, or
recommendation algorithms. Finally, while we have used
datasets borrowed from previous user studies, one could easily
apply the proposed method to Web logs, TREC data, and other
forms of datasets with various user actions recorded over time.</p>
    </sec>
    <sec id="sec-7">
      <title>6 ACKNOWLEDGEMENTS</title>
      <p>The work reported here is supported by The Institute of Museum
and Library Services (IMLS) Cyber Synergy project as well as
IMLS grant # RE-04-12-0105-12. The author is also grateful to
his PhD students Chathra Hendahewa and Roberto
GonzalezIbanez for their valuable contributions to this work.
7
in collaborative IR systems. In Proceedings of the 75th Annual
Meeting of the Association for Information Science and
Technology (ASIS&amp;T). Baltimore, MD, USA.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Adar</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bershad</surname>
            ,
            <given-names>B. N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Gribble</surname>
            ,
            <given-names>S. D.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Why we search: visualizing and predicting user behavior</article-title>
          .
          <source>In Proceedings of World Wide Web (WWW) Conference</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Adomavicius</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Tuzhilin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Toward the Next Generation of Recommender Systems: A Survey of the State-ofthe-Art and Possible Extensions</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>17</volume>
          (
          <issue>6</issue>
          ),
          <fpage>734</fpage>
          -
          <lpage>749</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Agichtein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <article-title>&amp;</article-title>
          <string-name>
            <surname>Bennett. P.N.</surname>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Search interrupted: Understanding and predicting search task continuation</article-title>
          .
          <source>In Proceedings of the Annual ACM Conference on Research and Development in Information Retrieval (SIGIR</source>
          )
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Cronen-Townsend</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Croft</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Predicting query performance</article-title>
          .
          <source>In Proceedings of the Annual ACM Conference on Research and Development in Information Retrieval (SIGIR</source>
          )
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Dignum</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kruschwitz</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fasli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yunhyong</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dawei</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beresi</surname>
            ,
            <given-names>U.C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>De Roeck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Incorporating Seasonality into Search Suggestions Derived from Intranet Query Logs</article-title>
          .
          <source>In Proceedings of IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT)</source>
          <year>2010</year>
          , vol.
          <volume>1</volume>
          , no., pp.
          <fpage>425</fpage>
          -
          <lpage>430</lpage>
          , Aug.
          <volume>31</volume>
          <fpage>2010</fpage>
          -
          <lpage>Sept</lpage>
          .
          <fpage>3</fpage>
          <lpage>2010</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Diriye</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buscher</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Leaving so soon? Understanding and predicting web search abandonment rationales</article-title>
          .
          <source>In Proceedings of CIKM</source>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>González-Ibáñez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R. W.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Pseudocollaboration as a method to perform selective algorithmic mediation</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Gwizdka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Cognitive load on web search tasks</article-title>
          .
          <source>Workshop on Cognition and the Web, Information Processing</source>
          , Comprehension, and Learning. Granada, Spain. Available from http://eprints.rclis.org/14162/1/GwizdkaJ_WCW2008_short_paper _finalp.pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karnawat</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mydland</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>White</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Evaluating implicit measures to improve web search</article-title>
          .
          <source>ACM TOIS</source>
          ,
          <volume>23</volume>
          (
          <issue>2</issue>
          ):
          <fpage>147</fpage>
          −
          <lpage>168</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Predicting query performance using query, result and interaction features</article-title>
          .
          <source>In Proceedings of RIAO</source>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Query performance prediction</article-title>
          ,
          <source>Information Systems</source>
          , Volume
          <volume>31</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>7</given-names>
          </string-name>
          ,
          <string-name>
            <surname>November</surname>
            <given-names>2006</given-names>
          </string-name>
          , Pages
          <fpage>585</fpage>
          -
          <lpage>594</lpage>
          , ISSN 0306-
          <issue>4379</issue>
          ,
          <fpage>10</fpage>
          .1016/j.is.
          <year>2005</year>
          .
          <volume>11</volume>
          .003.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Järvelin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>IR research: systems, interaction, evaluation and theories</article-title>
          .
          <source>ACM SIGIR Forum</source>
          ,
          <volume>45</volume>
          (
          <issue>2</issue>
          ), 17. doi:
          <volume>10</volume>
          .1145/2093346.2093348
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Kyoung-jae</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Financial time series forecasting using support vector machines</article-title>
          .
          <source>Neurocomputing</source>
          , Volume
          <volume>55</volume>
          ,
          <string-name>
            <surname>Issues</surname>
          </string-name>
          1-2,
          <year>September 2003</year>
          , Pages
          <fpage>307</fpage>
          -
          <lpage>319</lpage>
          , ISSN 0925-
          <issue>2312</issue>
          ,
          <fpage>10</fpage>
          .1016/S0925-
          <volume>2312</volume>
          (
          <issue>03</issue>
          )
          <fpage>00372</fpage>
          -
          <lpage>2</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gwizdka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            , &amp;
            <surname>Belkin</surname>
          </string-name>
          ,
          <string-name>
            <surname>N. J.</surname>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Analysis and evaluation of query reformulations in different task types</article-title>
          .
          <source>American Society for Information Science</source>
          ,
          <volume>47</volume>
          (
          <issue>17</issue>
          ). Available from http://dl.acm.org/citation.cfm?id=
          <volume>1920331</volume>
          .
          <fpage>1920356</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gwizdka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>C.</given-names>
            , &amp;
            <surname>Belkin</surname>
          </string-name>
          ,
          <string-name>
            <surname>N. J.</surname>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Predicting task difficulty for different task types</article-title>
          .
          <source>In Proceedings of the Association for Information Science</source>
          ,
          <volume>47</volume>
          (
          <issue>16</issue>
          ). Available from http://dl.acm.org/citation.cfm?id=
          <volume>1920331</volume>
          .
          <fpage>1920355</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Ming-Chi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Using support vector machine with a hybrid feature selection method to the stock trend prediction</article-title>
          .
          <source>Expert Systems with Applications</source>
          , Volume
          <volume>36</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>8</given-names>
          </string-name>
          ,
          <string-name>
            <surname>October</surname>
            <given-names>2009</given-names>
          </string-name>
          , Pages
          <fpage>10896</fpage>
          -
          <lpage>10904</lpage>
          , ISSN 0957-
          <issue>4174</issue>
          ,
          <fpage>10</fpage>
          .1016/j.eswa.
          <year>2009</year>
          .
          <volume>02</volume>
          .038
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Ord</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hyndman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koehler</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Snyder</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Forecasting with Exponential Smoothing (The State Space Approach</article-title>
          ). Springer,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Radinski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Svore</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teevan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horvitz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bocharov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Modeling and predicting behavioral dynamics on the Web</article-title>
          .
          <source>In Proceedings of WWW</source>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Resnick</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Varian</surname>
            ,
            <given-names>H. R.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <source>Recommender Systems. Communications of the ACM</source>
          ,
          <volume>40</volume>
          (
          <issue>3</issue>
          ),
          <fpage>56</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Saracevic</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>Evaluation of evaluation in information retrieval</article-title>
          .
          <source>In Proceedings of the Annual ACM Conference on Research and Development in Information Retrieval (SIGIR)</source>
          (pp.
          <fpage>138</fpage>
          -
          <lpage>146</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Croft</surname>
            ,
            <given-names>W. B.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Evaluating high accuracy retrieval techniques</article-title>
          .
          <source>Proceedings of the Annual ACM Conference on Research and Development in Information Retrieval (SIGIR)</source>
          (pp.
          <fpage>2</fpage>
          -
          <lpage>9</lpage>
          ). Sheffield, UK.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Gonzalez-Ibanez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Evaluating the Synergic Effect of Collaboration in Information Seeking</article-title>
          .
          <source>Proceedings of the Annual ACM Conference on Research and Development in Information Retrieval (SIGIR)</source>
          (pp.
          <fpage>913</fpage>
          -
          <lpage>922</lpage>
          ). Beijing, China.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Marchionini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Awareness in Collaborative Information Seeking</article-title>
          .
          <source>Journal of American Society of Information Science and Technology (JASIST)</source>
          ,
          <volume>61</volume>
          (
          <issue>10</issue>
          ),
          <fpage>1970</fpage>
          -
          <lpage>1986</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Shannon</surname>
            ,
            <given-names>C. E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Weaver</surname>
          </string-name>
          , W. Mathematical Theory of Communication. Urbana, IL: University of Illinois Press,
          <year>1963</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>White</surname>
            ,
            <given-names>R. W.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Assessing the scenic route: Measuring the value of search trails in web logs</article-title>
          .
          <source>In Proceedings of the Annual ACM Conference on Research and Development in Information Retrieval (SIGIR)</source>
          . Geneva, Switzerland.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>