<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>What-If Analysis: A Visual Analytics Approach to Information Retrieval Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Angelini</string-name>
          <email>angelini@dis.uniroma1.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Santucci</string-name>
          <email>santucci@dis.uniroma1.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianmaria Silvello</string-name>
          <email>silvello@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>La Sapienza” University of Rome</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper focuses on the innovative visual analytics approach realized by the Visual Analytics Tool for Experimental Evaluation (VATE2) system, which eases and makes more effective the experimental evaluation process by introducing the what-if analysis. The what-if analysis is aimed at estimating the possible effects of a modification to an Information Retrieval (IR) system, in order to select the most promising fixes before implementing them, thus saving a considerable amount of effort. VATE2 builds on an analytical framework which models the behavior of the systems in order to make estimations, and integrates this analytical framework into a visual part which, via proper interaction and animations, receives input and provides feedback to the user. We conducted an experimental evaluation to assess the numerical performances of the analytical model and a validation of the visual analytics prototype with domain experts. Both the numerical evaluation and the user validation have shown that VATE2 is effective, innovative, and useful.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        IR systems operate using a best match approach: in response to an often vague user
query, they return a ranked list of documents ordered by the estimation of their relevance
to that query. In this context effectiveness, meant as the ability of systems to retrieve
and better rank relevant documents while at the same time suppressing the retrieval of
not relevant ones, is the primary concern. Since there are no a-priori exact answers to
a user query, experimental evaluation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] based on effectiveness is the main driver of
research and innovation in the field. Indeed, the measurement of system performances
from the effectiveness point of view is basically the only means to determine the best
approaches and to understand how to improve IR systems.
      </p>
      <p>Nowadays, user tasks and needs are becoming increasingly demanding, the data
sources to be searched are rapidly evolving and greatly heterogeneous, the interaction
between users and IR systems is much more articulated, and the systems themselves
become increasingly complicated and constituted by many interrelated components. As an
example consider what web search is today: highly diversified results are returned from
web pages, news, social media, image and video search, products and more, and they
are all merged together by adaptive strategies driven by current and previous interaction
of the users with the system.</p>
      <p>Understanding and interpreting the results produced by experimental evaluation is
not a trivial task, due to the complex interactions among the components of an IR
system. Nevertheless, succeeding in this task is fundamental for detecting where systems
fail and hypothesizing possible fixes and improvements. As a consequences, this task is
mostly manual and requires huge amounts of time and effort.</p>
      <p>Moreover, after such activity, the researcher needs to come back to design and then
implement the modifications that the previous analysis suggested as possible solutions
to the identified problems. Afterwards, a new experimentation cycle needs to be started
to verify whether the introduced modifications actually give the expected improvement.
Therefore, the overall process of improving an IR system is extremely time and
resource demanding and proceeds through cycles where each new feature needs to be
implemented and experimented.</p>
      <p>The goal of this paper is to introduce a new phase in this cycle: we call it what-if
analysis and it falls between the experimental evaluation and the design and
implementation of the identified modifications. What-if analysis aims at estimating what the
effects of a modification to the IR system under examination could be before actually
being implemented. In this way researchers and developers can get a feeling of whether
a modification is worth being implemented and, if so, they can go ahead with its
implementation followed by a new evaluation and analysis cycle for understanding whether
it has produced the expected outcomes.</p>
      <p>What-if analysis exploits Visual Analytics (VA) techniques to make researchers and
developers: (i) interact with and explore the ranked result list produced by an IR system
and the achieved performances; (ii) hypothesize possible causes of failure and their
fixes; (iii) estimate the possible impact of such fixes through a powerful analytical
model of the system behavior.</p>
      <p>What-if analysis is a major step forward since it can save huge amounts of time and
effort in IR system development and, to the best of our knowledge, it has never been
attempted before.</p>
      <p>The paper is organized as follows: Section 2 discusses some related works; Section 4
explains in detail the proposed analytical framework which is then experimentally
evaluated in Section 5; Section 6 presents the visual analytics environment called Visual
Analytics Tool for Experimental Evaluation (VATE2); and, Section 7 draws some
conclusions and presents an outlook for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        Experimental evaluation is based on the Cranfield methodology [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] which makes use of
experimental collections C = (D; T; GT ) consisting of: a set of documents D
representing the domain of interest; a set of topics T , which simulates and abstracts actual user
information needs; and, the ground-truth GT , i.e. a kind of “correct” answer, where for
each topic t 2 T the documents d 2 D relevant to it are determined. System outputs
are then scored with respect to the ground-truth using whole breadth of performance
measures [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The ground-truth can consist of both binary relevance judgments or
multi-graded ones [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>Experimental evaluation is a demanding activity in terms of effort and required
resources that benefits from using shared datasets, which allow for repeatability of the
experiments and comparison among state-of-the-art approaches. Therefore, over the
last 20 years, experimental evaluation has been carried out in large-scale evaluation
campaigns at international level, such as the Text REtrieval Conference (TREC)3 in the
US and the Conference and Labs of the Evaluation Forum (CLEF)4 in Europe.</p>
      <p>
        The activity described in the previous section and aimed at understanding how and
why a system has failed is called failure analysis. To give the reader an idea of how
demanding it can be, let us consider the case of the the Reliable Information Access (RIA)
workshop [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which was aimed at systematically investigating the behavior of just
one component in an IR system, namely the relevance feedback module. Harman and
Buckley in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] reported that, for analyzing 8 systems, 28 people from 12 organizations
worked for 6 weeks requiring from 11 to 40 person-hours per topic for 150 overall
topics. These figures do not take into account the time then needed to implement the
identified modifications and perform another evaluation cycle to understand if they had
the desired effect.
      </p>
      <p>
        VA is typically exploited for the presentation and exploration of the documents
managed by an IR system [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. However, much less attention has been devoted to applying
VA techniques to the analysis and exploration of the performances of IR systems [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
To the best of our knowledge, our previous work is the most systematic attempt. We
explored several ways in which VA can help the interpretation and exploration of system
performances [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This preliminary work led to the development of Visual Information
Retrieval Tool for Upfront Evaluation (VIRTUE), a fully-fledged VA prototype which
specifically supports performance and failure analysis [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        We then started to explore the need for what-if analysis in the context of large-scale
evaluation campaigns, such as TREC or CLEF. In this context, evaluators do not have
access to the tested systems but they can only examine the final outputs, i.e. the ranked
result lists returned for each topic. In [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] we continued the study for the evaluation
campaigns and we set up an analytical framework for trying to learn the behavior of a
system just from its outputs, in order to obtain a rough estimation of the possible effects
of a modification to the system.
      </p>
      <p>Therefore, in this paper, we put VATE2 in a different context from the one of
evaluation campaigns. Indeed, the present version of VATE2 is thought for designers and
developers of IR systems who have access to all the internals of the system being tested
and they have the know-how to hypothesizing possible fixes and improvements.</p>
    </sec>
    <sec id="sec-3">
      <title>3 Intuitive Overview</title>
      <p>
        In order to understand how what-if analysis works we need to recall the basic ideas
underlying performance and failure analysis as designed and realized by VIRTUE [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        To quantify the performances of an IR system, VIRTUE adopts the Discounted
Cumulated Gain (DCG) family of measures [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] which have proved to be especially
well-suited for analyzing ranked results list. This is because they allow for graded
relevance judgments and embed a model of the user behavior while he scrolls down the
result list which also gives an account of its overall satisfaction.
      </p>
      <sec id="sec-3-1">
        <title>3 http://trec.nist.gov/ 4 http://www.clef-initiative.eu/</title>
        <p>We compare the result list produced by an experiment with respect to an ideal
ranking created starting from the relevant documents in the ground-truth, which represents
the best possible results that an experiment can return. In addition to what is typically
done, we compare the result list with respect to an optimal one created with the same
documents retrieved by the IR system but with an optimal ranking, i.e. a permutation
of the results retrieved by the experiment aimed at maximizing its performances by
sorting the retrieved documents in decreasing order of relevance. Therefore, the ideal
ranking compares the experiment at hand with respect to the best results possible, i.e.
also considering relevant documents not retrieved by the system, while the optimal
ranking compares an experiment with respect to what could have been done better with
the same retrieved documents.</p>
        <p>Looking at a performance curve, as the DCG curve is, it is not always easy to spot
what the critical regions in a ranking are. Indeed, DCG is a not-decreasing monotonic
function which increases only when a relevant document is found in the ranking.
However, when DCG does not increase, this could be due to two different reasons: either
you are in an area of the ranking where you are expected to put relevant documents but
you are putting a not relevant one and thus you do not gain anything; or, you are in an
area of the ranking where you are not expected to put relevant documents and, correctly,
you are putting a not relevant one, still gaining nothing.</p>
        <p>
          In order to overcome this and similar issues, we introduce the Relative Position (RP)
indicator [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]; RP allows us to quantify and explain what happens at each rank position
and its paired with a visual counterpart which eases the exploration of the performances
across the ranking. RP allows us to immediately grasp the most critical areas as we
can see in Figure 1 showing the VATE2 system. RP quantifies the effect of misplacing
relevant documents with respect to the ideal case, i.e. it accounts for how far a document
is from its ideal position. Overall, the greater the absolute value of RP is, the bigger the
distance of the document from its ideal interval is.
        </p>
        <p>We envision the following scenario. By exploiting VIRTUE a failure analysis is
conducted and the user hypothesizes the problem of the IR system at hand. At the same
time, the user hypothesizes that if he fixes such failure, a given relevant document would
be ranked higher than in the current system. As shown in Figure 1, what VATE2 offers
to the user is: (i) the possibility of dragging and dropping the target document in the
estimated position of the rank; (ii) the estimation of which other documents would be
affected by the movement of the target document and how the overall ranking would
be modified; (iii) the computation of the system performances according to the new
ranking.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Analytical Framework</title>
      <p>
        We introduce the basic notions regarding experimental evaluation in IR regarding the
functioning of VATE2; for a complete and formal definition of experimental evaluation
in IR refer to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>We consider relevance as an ordered set of naturals, say REL 2 N, where rel 2 REL
indicates the degree of relevance of a document d 2 D for a given topic t 2 T . The
ground truth associates a relevance degree to each document with respect to a topic. So,
let T be a set of topics and D a set of documents, then we can define a ground truth for
each topic, say tk 2 T , as a map GTtk of pairs (d; rel) with size jDj, where d 2 D is a
document and rel 2 REL is its relevance degree with respect to tk.</p>
      <p>We indicate with Ltk a list of triples (di; simi; reli) of length N representing the
ranked list of documents retrieved for a topic tk 2 T , where di 2 D is a document
(di 6= d j; 8i; j 2 [1; N] j i 6= j), simi 2 R is a degree indicating the similarity of di to
tk, and reli 2 REL indicates the relevance degree of di for tk; the triples in Ltk are in
decreasing order of similarity degree such that simi sim j if i &gt; j.</p>
      <p>Now, we can point out some methods to access the elements in a ranked list Lt that
we will use in the following. Lt (1; 1) = d1 returns the document in the first triple of
Lt , Lt (2; 1) = d2 the second document and so on; whereas, Lt (:; 1) returns the list of
documents in Lt . In the same vein, Lt (1; 2) = sim1 returns the similarity degree of the
first document in Lt and Lt (1; 3) = rel1 returns its relevance degree; moreover, Lt (1; :
) = (d1; sim1; rel1) returns the first triple in Lt .</p>
      <p>Relative Position (RP) is a measure that quantifies the misplacement of a document
in a ranked list with respect to the ideal one. In order to introduce RP we need to define
the concepts of minimum rank and maximum rank of a given relevance degree building
on the definition of ideal ranked list. Given an ideal ranked list It with length N 2 N+
and a relevance degree rel 2 REL, then the minimum rank minIt (rel) returns the first
position i 2 [1; N] at which we find a document with relevance degree equal to rel, while
the maximum rank maxIt (rel) returns the last position i at which we find a document
with relevance degree equal to rel in the ideal ranking.</p>
      <p>RPLt [i] =
80
&gt;
&lt;</p>
      <p>i
&gt;:i
minIt Lt (i; 3)
maxIt Lt (i; 3)
if minIt Lt (i; 3))
if i &lt; minIt Lt (i; 3)
if i &gt; maxIt Lt (i; 3)
i
maxIt Lt (i; 3)</p>
      <p>The RP measure points out the instantaneous and local effect of misplaced
documents and how much they are misplaced with respect to the ideal ranking It . In the
following definition, zero values denote documents which are within the ideal
interval; positive values denote documents which are ranked below their ideal interval, i.e.
documents of higher relevance degree that are in a position of the ranking where less
relevant ones are expected; and negative values denote documents which are above their
ideal interval, i.e. less relevant documents that are in a position of the ranking where
documents of higher relevance degree are expected.</p>
      <p>In VATE2, document clustering is adopted in the context of the failure hypothesis:
“closely associated documents tend to be affected by the same failures”, stating the
common intuition that a given failure will affect documents with common features,
and, consequently, that a fix for that failure will have an effect on the documents sharing
those common features.; once the user has selected a target document, say d j, within a
ranked list Lt , our goal is to get a cluster of documents, Cd j , similar to d j, where the
similarity is quantified by the IR system, say SX , which generated Lt .</p>
      <p>The creation of a cluster of documents similar to d j is very close to the operation
done by the SX IR system to get a ranked list of documents starting from a topic t.
Indeed, SX , takes the topic t and calculates the similarity between the topic and each
document in D; afterwards it returns a ranked list Lt of documents ordered by decreasing
similarity to the topic. The document cluster creation methodology we adopt in VATE2
follows this very procedure: given a target document d j we use it as topic and we submit
it to the IR system SX which returns a ranked list of documents, say Cd j , ordered by
decreasing similarity to d j.</p>
      <p>The first document in Cd j is always d j and then we encounter progressively less
similar documents. Cd j tell us which documents are seen in “the same way” by the IR
system being tested and that will probably be affected by the same issues found for d j.</p>
      <p>We limit to 10 the number of documents in the cluster to be moved in order to
consider only the most similar documents to the target document d j.</p>
      <p>A cluster Cd j is defined as a list of pairs (dk; simk) where dk 2 D is a document and
simk is the similarity of dk to d j. The first document in Cd j is d j and, by definition, it has
the maximum similarity value in the cluster. For every considered cluster we normalize
the similarities in the [0; 1] range by dividing their values by maximum similarity value
in the cluster.</p>
      <p>The clusters of documents so defined play a central role in the document movement
estimation of VATE2. Indeed, once a user spots a misplaced document, say d4, and
s/he decides to move it upward, the ten documents in the Cd4 cluster are also moved
accordingly.</p>
      <p>We developed two variations of movement: a simple one called constant movement
(cm) and a slightly more complex one called similarity-based movement (sbm). We
present the constant movement first and then we build on it to explain the
similaritybased one.</p>
      <p>Let us consider a general environment where a ranked list Lt is composed of N
triples such that the list of documents is d1; d2; : : : ; dN , where the subscript of the
documents indicates their position in the ranking. As a consequence of the failure hypothesis,
if we move a document d j 2 Lt from position j to position k – which means that we
move d j of D = j k positions – we also move the documents in the cluster C j of D
positions accordingly.</p>
      <p>The constant movement is based on three assumptions:
1. Linear movement: if d j is moved upward from position j to position k where D =
j k, all the documents in its cluster C j are moved by D in the same direction.
2. Cluster independence: the movement of the cluster C j does not imply the movement
of other clusters. This means that when we move d j in the position of dk, dk is
influenced by the movement, but the cluster Ck is not.
3. Unary shifting: if C j is moved by D positions, then the other documents in the
ranking have to make room for them and thus they are moved downward by one
position.</p>
      <p>The pseudo-code of the constant movement is reported by Algorithm 1
(implementing the actual movement) and Algorithm 2 (implementing the operations necessary to
reorder the ranked list after a movement)5. It starts to move by D the last document
5 For the sake of simplicity in these algorithms we employ four convenience methods:
POSITION(L; d) which returns the rank index of d in L, SIZE(L) which returns the number of
element of L, ADD(L; L0 ) which adds the elements of L0 at the end of L and GETCLUSTER(Lt; dj)
which returns the cluster of d j.</p>
      <p>Algorithm 1: MOVEMENT</p>
      <p>Input: The ranked list Lt, the document dj to be moved, the target rank position tPos.</p>
      <p>Output: The ranked list Lt after the movements.
1 Cdj GETCLUSTER(Lt; dj)
2 sPos POSITION(Lt; Cdj (1; 1))
3 for i SIZE(Cdj ) to 1 do
4 oldPos POSITION(Lt; Cdj (i; 1))
5 if oldPos == 0 then
6 oldPos SIZE(Lt) + 1
7 newPos oldPos (sPos tPos)
8 Lt REORDERLIST(Lt; oldPos; newPos; Cdj (i; 1))
9 end
10 return Lt
in the cluster Cd6 which is d4, so we can see that d4 is put in the place of d1 and d1 is
shifted downward by one position; afterwards, the algorithm repeats the same operation
for all the other documents in the cluster generating the reordered list Lt0 .</p>
      <p>We have seen that in a general setting, if we move d j upwards by D positions, all
the documents in C j move accordingly by D position. There are cases where this is
not possible, because the movement is capped on the top by one or more documents in
the cluster. As an example, consider a movement upward of d j, if there is a document
dw 2 C j such that w &lt; D , then dw cannot be moved upwards by D position, but at most
by w. In this case, dw is moved by w positions while the other documents, if possible,
are moved by D positions.</p>
      <p>We can see that this movement can be easily changed by altering the three starting
assumptions. For instance, one can decide that constant movement is no longer a valid
assumption, e.g. by saying that when d j in moved by D positions, the documents in C j
are moved by D s , where s is a variable calculated on the basis of the documents
rank or similarity score. The similarity-based movement does exactly so by changing
the way in which the new document positions are calculated starting from the D value
defined for the starting document d j; the new movement is obtained by substituting the
instruction at line 7 of Algorithm 1 with the following:
newPos
oldPos
1
sPos tPos
sPos</p>
      <p>Cdj (i; 2)</p>
      <p>We can see that the new position of a document di is weighted by two terms Cdj (i; 2)
which is the normalized similarity of di to d j in the cluster Cd j and sPos tPos which
sPos
determines the relative movement of the starting document d j. Every document di is
moved by the same increment of d j (e.g. 1) weighted by the normalized similarity of di
in the cluster. Basically, d j, which has similarity 1 by definition, is always moved by the
number of positions indicated by the user, whereas the other documents in the cluster
are moved by a number of position depending on the similarity to d j: the higher it is the
bigger the movement upward.</p>
    </sec>
    <sec id="sec-5">
      <title>Experimental Evaluation of the Analytical Framework</title>
      <p>VATE2 is expected to work as follows: the user examines a bugged system SB, identifies
the cause of a possible failure and makes a hypothesis about how the fixed version of the
system SF would rank the documents by dragging the spotted document in the expected
position.</p>
      <p>To conduct an experiment in a controlled environment which accounts for this
behaviour, we start from a properly working IR system SF and we produce a “bugged”
version of it SB by changing one component or one specific feature at a time. Then,
we consider all the possible movements that move a relevant document from a wrong
position in SB to the correct one in SF and we count how many times the prediction
of VATE2, i.e. improvement or deterioration of the performances, corresponds to the
actual improvement or deterioration passing from SB to SF .</p>
      <p>
        The system we used is Terrier ver. 4.06[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], an open source and widely used system
in the IR field which is developed and maintained by the University of Glasgow. To run
the experiment, we used a standard and openly available experimental collection C , the
TREC 8, 1999, Ad-Hoc collection [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        We experimented the use case about the stemmer and we setup the Terrier
system with four different stemmers by keeping all the other components fixed, namely:
Porter [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] stemmer, Weak Porter stemmer, Snowball stemmer and no stemmer.
      </p>
      <p>We considered those pairs of systems SB and SF , which correspond to sensible and
useful cases in practice, i.e. when you pass from a lesser performing system SB to a
better one SF , as for example when you pass from no stemming SB to stemming SF .</p>
      <sec id="sec-5-1">
        <title>6 http://terrier.org/</title>
        <p>For each topic, we identified the set of all the possible predictions T P, where each
misplaced relevant document d j in LBt is moved upwards along with its cluster Cd j to
the correct position determined by its rank in the fixed ranked list LFt , thus generating
a predicted ranked list LPt .</p>
        <p>We computed the DCG for each of the above ranked lists: DCGLBt and DCGLFt
indicate the DCG of the bugged system SB and fixed system SF while DCGLPt indicates
the DCG of the predicted system SP for the i-th possible movement in T P.</p>
        <p>We consider a prediction by VATE2 to be correct if a performance improvement (or
deterioration) between the actual bugged system SB and the fixed one SF corresponds
to a performance improvement (or deterioration) between the actual bugged systems SB
and the predicted one SP. Let sgn(x) = 1 if x 0 and sgn(x) = 1 if x &lt; 0; then for
each possible prediction p 2 T P we define the Correct Prediction (CP) measure as:
CPt [p] =
sgn DCGLFt</p>
        <sec id="sec-5-1-1">
          <title>DCGLBt</title>
          <p>+ sgn DCGLPt</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>DCGLBt</title>
          <p>2
Lastly, we define the Prediction Precision (PP) for topic t as the number of correct
predictions over the total number of possible predictions: PPt = jT1Pj åp2T P CPt [p]</p>
          <p>PPt ranges between 0 and 1, where 0 indicates that no correct prediction has been
made and 1 indicates that all the predictions were correct.</p>
          <p>In Table 1 we report the mean DCG value (DCG is calculated topic by topic and
then it is averaged over all 50 topics) for the four different stemmers. We can see that
there are substantial differences between the systems, in particular the “Weak Porter”
and the “No Stemmer” systems have much lower performances with respect to the best
one which is the “Porter” system. In Table 2 we report the results of the tests in terms
of Prediction Precision (PP) averaged over all the topics for the considered pairs of
systems. We can see that the PP is in general satisfactory and it is higher for those pairs
where the difference in DCG is higher – e.g. SB = No Stemmer and SF = Snowball.</p>
          <p>
            Even if the constant movement behaves better than the similarity-based one in 3
out of 5 considered cases, there is no clear evidence that one of the two movements
performs better since they are not significantly different from the statistical point of
view according to Student’s t test [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] which returns a p-value p = 0:9459. Therefore,
in the running implementation of VATE2, we decided to use the constant movement in
VATE2 because its behavior is more intuitive to the users.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6 Visual Analytics Environment</title>
      <p>
        In Figure 1(a) we can see an overview of VATE2 system which is available at the URL:
http://ims-ws.dei.unipd.it/vate_ui/ and a video is available here: http://
ims.dei.unipd.it/video/VATE/sigir-vate.mp4. VATE2 functioning has been
described in details in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>We can see that the system is structured in three main components. The
Experimental collection information (A) which allows the user to inspect and interact with
the information regarding the experimental collection. More in detail, it is divided into
three sub-components. The first is the “Experiment Selection” where the user can select
the experimental collection, the experiment to analyze and the evaluation measure and
its parameters. The second sub-component is the “Topic Information” composed of the
structured description of the topic and the topic selection grid. The third sub-component
is the “Document Information” reporting the content of the document under analysis.</p>
      <p>The Ranked list exploration (B) which is placed on the center and shows a visual
representation of the ranked list. More in detail, the documents are represented as
rectangles ordered by rank from top to bottom where the color indicates the RP value. The
intensity of the color encodes the severity of the misplacement, the more intense the
worse the misplacement. This visualization provides the user with an intuitive insight
into the quality of the ranking.</p>
      <p>The Performance view (C) which is placed on the right side and shows the
performance curves of the selected experiment. The yellow curve is the ideal one, the magenta
curve is the optimal one and the cyan curve is the experiment one. The user can analyze
the trend of the experiment by comparing the behavior of its curve with the ideal and
optimal ranking by spotting the possible areas of improvement.</p>
      <p>The user can interactively select the topic to be analyzed in the topic selection grid
and the ranked list and the performance curves are updated accordingly to the selected
topic for the given experiment. The ranked list can be dynamically inspected by
hovering the mouse over the documents. Moreover, the user can interact with the
“Performance view” by hovering the mouse over the curves which, by means of a tooltip,
reports information about the document and the performance score.</p>
      <p>As shown in Figure 1, concerning what-if analysis, once the user selects a document,
the system displaces on the right the rectangles corresponding to the documents in its
similarity cluster and reports their identifiers also on the right. Once the user selects a
document, s/he can drag it to a new position in the ranked list; afterwards, the movement
algorithm is triggered and moves the document along with its similarity cluster in the
new positions. This action is visually shown to the user and it is represented with an
animated movement of the corresponding rectangles to the new positions. After the
movement the ranked list is split in two parts: the old ranked list on the left and the new
ranked list produced after the movement on the right. In this way the user can visually
compare the effects of the movement and see what other documents have been affected
by it.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>In this paper we explored the application of VA techniques to the problem of IR
experimental evaluation and we have seen how joining a powerful analytical framework with a
proper visual environment can foster the introduction of a new, yet highly needed phase,
which is the what-if analysis. Indeed, improving or fixing an IR system is an extremely
resource demanding activity and what-if analysis can help in getting an estimate of what
is worth doing, thus saving time and effort.</p>
      <p>We designed and developed the VATE2 system which has proven to be robust and
well-suited for its purposes from a two-fold point of view. The experimental evaluation
has numerically shown that the analytical engine, the failure hypothesis and the
corresponding way of clustering documents together with the document movement
estimation algorithms are satisfactory. The validation with domain experts has confirmed that
VATE2 is innovative, addresses an open and relevant problem and provides an effective
and intuitive solution to it.</p>
      <p>As future work, we plan to explore what happens when multiple movements are
considered all together. This will require an extension of the analytical engine in order to
account for the possible inter-dependencies among the different movements. Moreover,
also the visual analytics environment will require a substantial modification in order
to support users in intuitively dealing with multiple movements, interacting with the
history of the performed movements and moving back and forth within it.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Angelini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , G. Santucci, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          .
          <article-title>Visual Interactive Failure Analysis: Supporting Users in Information Retrieval Evaluation</article-title>
          .
          <source>In Proc. 4th Symposium on Information Interaction in Context (IIiX</source>
          <year>2012</year>
          ), pages
          <fpage>195</fpage>
          -
          <lpage>203</lpage>
          . ACM Press,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Angelini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , G. Santucci, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          .
          <article-title>Improving Ranking Evaluation Employing Visual Analytics. In Information Access Evaluation meets Multilinguality, Multimodality, and Visualization</article-title>
          .
          <source>Proceedings of the Fourth International Conference of the CLEF Initiative (CLEF</source>
          <year>2013</year>
          ), pages
          <fpage>29</fpage>
          -
          <lpage>40</lpage>
          . LNCS 8138, Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Angelini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , G. Santucci, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello. VIRTUE</surname>
          </string-name>
          :
          <article-title>A visual tool for information retrieval performance evaluation and failure analysis</article-title>
          .
          <source>Journal of Visual Languages &amp; Computing (JVLC)</source>
          ,
          <volume>25</volume>
          (
          <issue>4</issue>
          ):
          <fpage>394</fpage>
          -
          <lpage>413</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Angelini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , G. Santucci, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          .
          <article-title>Visual Analytics for Information Retrieval Evaluation (VAIR E¨</article-title>
          <year>2015</year>
          ).
          <source>In Advances in Information Retrieval. Proc. 37th European Conference on IR Research (ECIR</source>
          <year>2015</year>
          ), pages
          <fpage>709</fpage>
          -
          <lpage>812</lpage>
          . LNCS 9022, Springer,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>M.</given-names>
            <surname>Angelini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , G. Santucci, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          .
          <article-title>A visual analytics approach for whatif analysis of information retrieval systems</article-title>
          .
          <source>In Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , Pisa, Italy,
          <source>July 17-21</source>
          ,
          <year>2016</year>
          . Accepted for publication,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>C. W.</given-names>
            <surname>Cleverdon</surname>
          </string-name>
          .
          <source>The Cranfield Tests on Index Languages Devices. In Readings in Information Retrieval</source>
          , pages
          <fpage>47</fpage>
          -
          <lpage>60</lpage>
          . Morgan Kaufmann Publisher, Inc.,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sabetta</surname>
          </string-name>
          , G. Santucci, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Tino</surname>
          </string-name>
          .
          <article-title>Visual Comparison of Ranked Result Cumulated Gains</article-title>
          .
          <source>In Proc. 2nd International Workshop on Visual Analytics (EuroVA</source>
          <year>2011</year>
          ), pages
          <fpage>21</fpage>
          -
          <lpage>24</lpage>
          . Eurographics Association,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , G. Silvello,
          <string-name>
            <given-names>H.</given-names>
            <surname>Keskustalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pirkola</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Ja</surname>
          </string-name>
          <article-title>¨rvelin. The Twist Measure for IR Evaluation: Taking User's Effort Into Account</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology (JASIST)</source>
          ,
          <volume>67</volume>
          :
          <fpage>620</fpage>
          -
          <lpage>648</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>D.</given-names>
            <surname>Harman</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Buckley</surname>
          </string-name>
          .
          <source>Overview of the Reliable Information Access Workshop. Information Retrieval</source>
          ,
          <volume>12</volume>
          (
          <issue>6</issue>
          ):
          <fpage>615</fpage>
          -
          <lpage>641</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>D. K. Harman</surname>
          </string-name>
          . Information Retrieval Evaluation. Morgan &amp; Claypool Publishers, USA,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>K. Ja</surname>
          </string-name>
          <article-title>¨rvelin and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Keka</surname>
          </string-name>
          <article-title>¨la¨inen. Cumulated Gain-Based Evaluation of IR Techniques</article-title>
          .
          <source>ACM Transactions on Information Systems (TOIS)</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ):
          <fpage>422</fpage>
          -
          <lpage>446</lpage>
          ,
          <year>October 2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>J. Keka</surname>
          </string-name>
          <article-title>¨la¨inen and K. Ja¨rvelin. Using Graded Relevance Assessments in IR Evaluation</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology (JASIST)</source>
          ,
          <volume>53</volume>
          (
          <issue>13</issue>
          ):
          <fpage>1120</fpage>
          -
          <lpage>1129</lpage>
          ,
          <year>November 2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. I. Ounis,
          <string-name>
            <given-names>G.</given-names>
            <surname>Amati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Plachouras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Lioma</surname>
          </string-name>
          .
          <article-title>Terrier: A High Performance and Scalable Information Retrieval Platform</article-title>
          .
          <source>In Proceedings of ACM SIGIR'06 Workshop on Open Source Information Retrieval (OSIR</source>
          <year>2006</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>An algorithm for suffix stripping</article-title>
          .
          <source>Program</source>
          ,
          <volume>14</volume>
          (
          <issue>3</issue>
          ):
          <fpage>130</fpage>
          -
          <lpage>137</lpage>
          ,
          <year>July 1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>T.</given-names>
            <surname>Sakai</surname>
          </string-name>
          . Metrics, Statistics, Tests.
          <source>In Bridging Between Information Retrieval and Databases - PROMISE Winter School</source>
          <year>2013</year>
          , Revised Tutorial Lectures, pages
          <fpage>116</fpage>
          -
          <lpage>163</lpage>
          . Lecture Notes in Computer Science (LNCS) 8173, Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Student</surname>
          </string-name>
          .
          <source>The Probable Error of a Mean. Biometrika</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          ,
          <year>March 1908</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. K.</given-names>
            <surname>Harman</surname>
          </string-name>
          .
          <article-title>Overview of the Eigth Text REtrieval Conference (TREC8)</article-title>
          .
          <source>In The Eighth Text REtrieval Conference (TREC-8)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          . National Institute of Standards and Technology (NIST),
          <source>Special Publication 500-246</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          . Visualization for Information Retrieval. Springer-Verlag,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>