<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparing ANOVA Approaches to Detect Significantly Diferent IR Systems*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Discussion Paper</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guglielmo Faggioli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Padova</institution>
          ,
          <addr-line>Padova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <abstract>
        <p>The ultimate goal of the evaluation is to understand when two IR systems are (significantly) diferent. To this end, many comparison procedures have been developed over time. However, to date, most reproducibility eforts focused just on reproducing systems and algorithms, almost fully neglecting to investigate the reproducibility of the methods we use to compare our systems. In this paper, we focus on methods based on ANalysis Of VAriance (ANOVA), which explicitly model the data in terms of diferent contributing efects, allowing us to obtain a more accurate estimate of significant diferences. In this context, we compare statistical analysis methods based on “traditional” ANOVA (tANOVA) to those based on a bootstrapped version of ANOVA (bANOVA) and those performing multiple comparisons relying on a more conservative Family-wise Error Rate (FWER) controlling approach to those relying on a more lenient False Discovery Rate (FDR) controlling approach. Our findings highlight that, compared to the tANOVA approaches, bANOVA presents greater statistical power, at the cost of lower stability.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Comparing IR systems and identifying when they are significantly diferent is a critical task
for both industry and academia [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2, 3, 4, 5</xref>
        ]. The literature still lacks reproducibility studies on
the statistical tools used to compare the performance of such systems and algorithms. Using
reproducible statistical tools is crucial to drawing robust inferences and conclusions. In this
context, ANalysis Of VAriance (ANOVA) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a widely used technique, where we model
performance as a linear combination of factors, such as topic and system efects, and, by
developing more and more sophisticated models, we accrue higher sensitivity in determining
significant diferences among systems. We focus on two recently developed ANOVA models,
bANOVA, developed by Ferro and Sanderson [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and tANOVA, developed by Voorhees et al.
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Voorhees et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] used sharding of the document corpus to obtain the replicates of the
performance score for every (topic, system) pairs needed to develop a model accounting not
only for the main efects, but also for the interaction between topics and systems; Voorhees
et al. also used an ANOVA version based on residuals bootstrapping [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which we call bANOVA.
Similarly, Ferro and Sanderson [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] used document sharding as well but they developed a more
comprehensive model, based on traditional ANOVA, which also accounts for the shard factor,
the shard*system interaction, and the topic*shard interaction; we call this approach tANOVA.
Another fundamental aspect to consider when comparing several IR systems is the need to
adjust for multiple comparisons [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. Indeed, when comparing just two systems, significance
tests control the Type-I error at the significance level  . However, when  simultaneous tests
are carried out, the probability of committing at least one Type-I error increases up to 1 −
(1 −  ). To correct for the multiple comparisons problem, Voorhees et al. adopted a lenient
False Discovery Rate (FDR) correction by Benjamini and Hochberg [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]; Ferro and Sanderson
used a conservative Family-wise Error Rate (FWER) correction, using the Honestly Significant
Diference ( HSD) method by Tukey [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In conclusion, we identified three aspects that can
impact the reproducibility of the above-mentioned ANOVA approaches: i) the strategy used to
obtain replicates, ii) the kind of ANOVA used, and iii) the control procedure for the pairwise
comparisons problem. Our work investigates behaviour of tANOVA and bANOVA (Voorhees et al.
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]) under diferent experimental settings – with respect to the above-mentioned focal points –
and the generalizability of their results.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Experimental Analysis</title>
      <sec id="sec-2-1">
        <title>2.1. Experimental Setup &amp; ANOVA Models</title>
        <p>
          Akin to Voorhees et al., we used two collections: the TREC-3 Adhoc track [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and TREC-8
Adhoc track [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. In this work we report results only on TREC-8, we refer to [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] for all the
experiments. We use Average Precision (AP) as performance measure. We consider three
ANOVA models: (MD1) : a traditional two-way ANOVA that accounts only for the topic and
the system factors; (MD2) : A second model, similar to the previous one, that considers also
the interaction between topics and systems; (MD3) : A third model that includes also the shard
factor and all the interactions between diferent factors.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Impact of the multiple comparison strategies and bootstrapping</title>
        <p>
          To investigate the diferences between ANOVA approaches, our first analysis compares the
number of statistically significantly diferent ( s.s.d.) system pairs found by them. We consider the
following multiple comparison procedures: HSD for tANOVA, as originally proposed by Ferro and
Sanderson, indicated with tANOVA(HSD); Benjamini-Hochberg (BH) for bANOVA, as originally
proposed by Voorhees et al., indicated with bANOVA(BH); and, BH for tANOVA, indicated with
tANOVA(BH). tANOVA with Benjamini-Hochberg correction is here employed and analyzed for
the first time. It takes the p-values on the diference between levels of the factors produced
by the traditional ANOVA, but corrects them using the BH correction. The rationale behind
it is that it enjoys the statistical properties provided by the ANOVA while granting a higher
discriminative power, due to the BH correction procedure. zero has been used as interpolation
strategy; in Section 2.3 we empirically show that the interpolation strategy has a negligible
efect on the results. Finally, we experiment all the models from (MD1) to (MD3) with all the
ANOVA approaches; note that (MD3) has not been studied before for bANOVA and this represents
another generalizability aspect. Table 1 reports the results averaged over the five samples of
shards together with their condfience interval. Numbers on the diagonal of Table 1 describe how
many pairs of systems are considered s.s.d. by a given approach; numbers above the diagonal
are the additional s.s.d. pairs found by one method with respect to the other. Table 1 shows
that, as the complexity of the model increases from (MD1) to (MD3), the pairs of systems
deemed significantly diferent increase as well, confirming previous findings in the literature.
tANOVA(HSD) controls tANOVA(BH) since all the s.s.d. pairs for tANOVA(HSD) are significant also
for tANOVA(BH); this was expected since FWER controls FDR [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. It is possible see this by
considering the diferences between approaches (above diagonal): by summing the diference
between tANOVA(HSD) and tANOVA(BH) to the tANOVA(HSD) you obtain back the number of
s.s.d. pairs identified by t ANOVA(BH). However, this pattern holds also for bANOVA(BH) and
tANOVA(BH), i.e. all the s.s.d. pairs of tANOVA(BH) are s.s.d. pairs for bANOVA(BH) too. While the
relation between BH and HSD was expected, this finding sheds some light on the diference
between using a traditional or a bootstrapped version of ANOVA. In summary, most of the
increase in the s.s.d. pairs is due to the correction procedure rather than the use of bootstrap.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Stability of ANOVA Models with respect to Diferent Interpolation Values</title>
        <p>Both tANOVA and bANOVA are based on the concept of “corpus sharding”: divide the corpus in
non-overlapping subcopora, and use those to compute the systems performance. Problems arise
when a shard do not contain any relevant documents, since several Information Retrieval (IR)
measures are not defined. Thus, we study the impact of the interpolation strategy, i.e. how
to substitute missing values for topics without any relevant document on a given shard, for
the diferent approaches. If a shard does not contain any relevant document for a topic, we
interpolate the missing value using 4 possible strategies: zero; lq, the value of the lower
quartile of the measure scores; mean, the average value of the measure scores; and, one. To
assess the stability with respect to interpolation strategies, we resample shards 5 times ans
we consider the number of Passive Disagreements (PD), i.e. the number of pairs of systems
A and B for which an approach considers A to be significantly better than B on a sample but
A is not significantly better than B on the other sample. Here, for space reasons, we report
only the results for tANOVA(HSD) and bANOVA(BH), being the tANOVA(BH) midway between
these two. Table 2 reports the average PD counts together with their confidence interval for
models (MD2) and (MD3). Values on the diagonal are the average PD observed using the same
interpolation strategy, but over the pairs of shards samples. The upper triangle of the Table
contains the average PD when using two diferent interpolation values. Table 2 shows what
happens if, using model (MD2) by Voorhees et al., instead of re-sampling shards we use an
interpolation value. We can note that, as the interpolation value increases, the PD count on
the diagonal tends to increase too. When it comes to the upper triangles, we interestingly find
that bANOVA(BH) is much less sensitive to the interpolation values than tANOVA(HSD), being the
PD counts substantially lower. The bootstrapped version of ANOVA (bANOVA) appears to be
less stable with respect to the resharding (higher diagonal values). This phenomenon is likely
due to its greater discriminative power: since a small evidence for bANOVA is enough to assess
when two systems are diferent, the random resharding might produce spurious evidence and
thus large variation among diferent samples. In the part of Table 2 concerning (MD3), both
tANOVA(HSD) and bANOVA(BH) have upper triangle equal to zero, and thus are independent from
the interpolation values. Indeed, the bANOVA approach samples the residuals and Ferro and
Sanderson proved that they are independent of the interpolation value for (MD3). Therefore,
using (MD3) also the bootstrap approach by Voorhees et al. does not need to re-sample shards.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusions and Future Work</title>
      <p>
        In this work, we compared bANOVA [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and tANOVA approaches under diferent conditions. We
found out that tANOVA tends to be more robust than bANOVA with respect to the actual random
shards used, suggesting more reliability in drawing the same conclusions. On the other hand,
when using partial ANOVA models like (MD2) which are not able to deal with shards without
relevant documents, bANOVA is more robust than tANOVA to the chosen interpolation value.
Regarding the multiple comparison strategy, we have found that tANOVA with HSD is more
restrictive than bANOVA but tANOVA with BH correction behaves similarly to bANOVA. Overall,
we can conclude that, the decision of the model and the correction technique depends on the
ifnal aim of the researcher. If stability is more important, t ANOVA(HSD) is preferable, since it is
more stable with respect to random shards and less computationally expensive. Conversely,
if the focus is on the number of pairs, bANOVA(BH) gives the maximum boost, at the price of
lower stability for random shards. Future work will investigate the use of uneven-size random
shards, instead of the even-size ones used in the literature so far.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Faggioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <article-title>System efect estimation by sharding: A comparison between anova approaches to detect significant diferences</article-title>
          ,
          <source>in: European Conference on Information Retrieval</source>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Hull</surname>
          </string-name>
          ,
          <article-title>Using Statistical Testing in the Evaluation of Retrieval Experiments</article-title>
          ,
          <source>in: Proc. SIGIR</source>
          ,
          <year>1993</year>
          , pp.
          <fpage>329</fpage>
          -
          <lpage>338</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Savoy</surname>
          </string-name>
          , Statistical Inference in
          <source>Retrieval Efectiveness Evaluation, Information Processing &amp; Management</source>
          <volume>33</volume>
          (
          <year>1997</year>
          )
          <fpage>495</fpage>
          -
          <lpage>512</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Carterette</surname>
          </string-name>
          ,
          <source>Multiple Testing in Statistical Analysis of Systems-Based Information Retrieval Experiments</source>
          ,
          <source>ACM Trans. Inf. Syst</source>
          <volume>30</volume>
          (
          <year>2012</year>
          ) 4:
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          :
          <fpage>34</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Culpepper</surname>
          </string-name>
          , G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kurland</surname>
          </string-name>
          ,
          <article-title>Topic dificulty: Collection and query formulation efects</article-title>
          ,
          <source>ACM Transactions on Information Systems</source>
          <volume>40</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rutherford</surname>
          </string-name>
          ,
          <article-title>ANOVA and ANCOVA. A GLM Approach</article-title>
          , 2nd ed., John Wiley &amp; Sons, New York, USA,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Sanderson, Improving the Accuracy of System Performance Estimation by Using Shards</article-title>
          ,
          <source>in: Proc. SIGIR</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>805</fpage>
          -
          <lpage>814</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Samarov</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Soborof</surname>
          </string-name>
          ,
          <article-title>Using Replicates in Information Retrieval Evaluation</article-title>
          ,
          <source>ACM Trans. Inf. Syst</source>
          <volume>36</volume>
          (
          <year>2017</year>
          )
          <volume>12</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          :
          <fpage>21</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Efron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Tibshirani</surname>
          </string-name>
          ,
          <article-title>An Introduction to the Bootstrap, Chapman</article-title>
          and Hall/CRC, USA,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Fuhr</surname>
          </string-name>
          ,
          <article-title>Some Common Mistakes In IR Evaluation, And How They Can Be Avoided</article-title>
          ,
          <source>SIGIR Forum 51</source>
          (
          <year>2017</year>
          )
          <fpage>32</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sakai</surname>
          </string-name>
          ,
          <article-title>On Fuhr's Guideline for IR Evaluation, SIGIR Forum 54 (</article-title>
          <year>2020</year>
          ) p14:
          <fpage>1</fpage>
          -
          <lpage>p14</lpage>
          :
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Benjamini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hochberg</surname>
          </string-name>
          ,
          <article-title>Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing</article-title>
          ,
          <source>J. Royal Stat. Soc</source>
          .
          <volume>57</volume>
          (
          <year>1995</year>
          )
          <fpage>289</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Tukey</surname>
          </string-name>
          ,
          <source>Comparing Individual Means in the Analysis of Variance, Biometrics</source>
          <volume>5</volume>
          (
          <year>1949</year>
          )
          <fpage>99</fpage>
          -
          <lpage>114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D. K.</given-names>
            <surname>Harman</surname>
          </string-name>
          ,
          <article-title>Overview of the Third Text REtrieval Conference (TREC-3)</article-title>
          ,
          <source>in: Proc. TREC</source>
          ,
          <year>1994</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. K.</given-names>
            <surname>Harman</surname>
          </string-name>
          ,
          <article-title>Overview of the Eigth Text REtrieval Conference (TREC-8)</article-title>
          ,
          <source>in: Proc. TREC</source>
          ,
          <year>1999</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Hsu</surname>
          </string-name>
          , Multiple Comparisons.
          <article-title>Theory and methods</article-title>
          , Chapman and Hall/CRC, USA,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>