<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On the Reproduciblity and Robustness of Query Performance Prediction Experiments - An Extended Abstract</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Suchana Datta</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Debasis Ganguly</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mandar Mitra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Derek Greene</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indian Statistical Institute</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University College Dublin</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Glasgow</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Query performance prediction (QPP), i.e. the process of estimating the retrieval quality of an IR system, has attracted the attention of the IR research community for several years. A diverse range of pre-retrieval (e.g. AvgIDF) and post-retrieval approaches (e.g. WIG, NQC, UEF) have been proposed for QPP. Specifically, given a query and an IR system, QPP methods compute a score that is indicative of the efectiveness of the system for the given query. While this score is typically not interpreted as a statistical estimate of a specific evaluation metric (e.g. AP or nDCG), it is indeed expected to be correlated with a standard evaluation measure computed over a ground-truth set of assessed relevant documents. Indeed, the efectiveness of a QPP method is determined by measuring the correlation between its predicted efectiveness scores and the values of some standard evaluation metric over a set of queries. This abstract summarizes our ECIR 2022 research article 'An Analysis of Variations in the Efectiveness of Query Performance Prediction' [ 1], where we analysed the relative stability of QPP outcomes (rank correlations) with respect to changes in the IR models used to derive the top-retrieved documents, or the IR evaluation metrics used to order a given set of queries from easy to dificult. As per our findings, we emphasize in this abstract that such variations in QPP results (both in terms of the absolute values themselves and also in terms of the relative efectiveness of diferent QPP systems) can lead to dificulties in reproducing QPP experiment results on standard datasets. We now summarize the research questions and the findings of our study [ 1]. The context of a QPP experiment depends on 3 factors, which are i) a list of top- documents retrieved, ii) the IR model or scoring function that is used to derive this list, and iii) an IR evaluation function (e.g., average precision or AP) that is used to induce an ordering over the set of queries (e.g., low AP to high AP indicating a spectrum of easy to dificult queries). The first research question RQ1, that we investigated in our paper, [1] is - 'Do variations in the QPP contexts lead to significant diferences in measured QPP outcomes ?'. Next, as the second</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>research question RQ2, we investigated the following: ‘Do variations in the QPP contexts lead
to significant diferences in the relative efectiveness of diferent QPP methods ?’.</p>
      <p>
        To investigate the above research questions in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], we conducted QPP experiments1 on
the widely-used TREC-Robust dataset, which consists of 249 queries. To obtain diverse QPP
contexts for our experiments, we tried out a number of diferent combinations of IR models
and IR evaluation metrics. Specifically, as IR models we employed a) language modeling with
Jelinek-Mercer smoothing (LMJM), b) language modeling with Dirichlet smoothing (LMDir),
and c) BM25. As choices for the IR evaluation metric, we considered a) AP, b) nDCG, c) P@10,
and d) recall. We compared seven diferent QPP methods in our experiments, namely a) AvgIDF,
b) Clarity, c) WIG, d) NQC, and three variants of UEF derived from three diferent base QPP
models - e) UEF(Clarity), f) UEF(WIG) and g) UEF(NQC). We now summarize the main findings
of our experimental study (for more details see [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]).
• Observations related to RQ1 (diferences in QPP outcomes):
– With NQC as the QPP method, we observed that LMJM yielded the highest deviations in
observed QPP outcomes (specifically, rank correlation computed by  ) across the 4 diferent
IR metrics used to obtain the QPP ground-truth (a reference order of query dificulty). The
diference between the highest and the lowest rank correlation values were significant
(0.3657 with recall and 0.2061 with P@10). This indicates that QPP methods do not
generalize well enough across diferent IR metrics. In other words, it is hard to consistently
predict which queries are easy and which ones are dificult across the diferent notions of
how this query dificulty itself is defined (e.g., via a precision or a recall oriented measure).
– With NQC as the QPP method, we observed significant diferences in the highest and
lowest QPP outcomes, across diferent IR models. The highest diference in rank correlation
was recorded between BM25 ( = 0.3563) and LMDir ( = 0.4354), the QPP ground-truth
being defined with respect to AP. This shows that the efectiveness of a QPP method is
also not consistent for diferent IR models. In other words, it is not easy to predict that
which queries are easy and which ones are dificult consistently well enough for diferent
IR systems.
• Observations related to RQ2 (diferences in relative efectiveness of QPP methods):
– We observed that the relative ranks of QPP systems (ordered by the efectiveness measure
 ) changed significantly across diferent IR metrics. The highest disagreement in the ranks
were observed between AP@10 and recall@1000. This indicates that what may be the
best QPP method for predicting the query dificulty induced by AP@10 may not be so for
predicting query performance with respect to recall@1000.
– Similar trends were also observed for diferences in the choice of IR models in the QPP
context. The highest disagreements in the relative performance of QPP methods were
observed for the P@10 metric across LMJM and LMDir. This indicates that what may be
the best QPP method for predicting the retrieval efectiveness of LMJM may not be the best
one when it comes to predicting the system performance of LMDir.
      </p>
      <p>
        1Implementation available at: https://github.com/suchanadatta/qpp-eval.git
The main takeaway from this extended abstract is that since our extensive investigation in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
has shown that QPP outcomes are indeed sensitive to the experimental setup used, any future
experiment on QPP should emphasize clear specification of the experimental setup to warrant
better reproducibility.
      </p>
      <p>Acknowledgement
The first and the fourth authors were supported by the Science Foundation Ireland (SFI) grant
number SFI/12/RC/2289_P2.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ganguly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Datta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Greene</surname>
          </string-name>
          ,
          <article-title>An analysis of variations in the efectiveness of query performance prediction</article-title>
          ,
          <source>in: Proc. of ECIR'22</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>215</fpage>
          -
          <lpage>229</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>