<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Do Hard Topics Exist? A Statistical Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Discussion Paper</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Shane Culpepper</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guglielmo Faggioli</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oren Kurland</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>RMIT University</institution>
          ,
          <addr-line>Melbourne</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Technion, Israel Institute of Technology</institution>
          ,
          <addr-line>Haifa</addr-line>
          ,
          <country country="IL">Israel</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Padova</institution>
          ,
          <addr-line>Padova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Several recent studies have explored the interaction efects between topics, systems, corpora, and components when measuring retrieval efectiveness. However, all of these previous studies assume that a topic or information need is represented by a single query. In reality, users routinely reformulate queries to satisfy an information need. Recently there has been renewed interest in the notion of “query variations” which are essentially multiple user formulations for an information need. Like many retrieval models, some queries are highly efective while others are not. In this work 1, we explore the fundamental problem of studying the interaction components of an IR experimental collection. Our findings show that query formulations have a comparable efect size to the topic factor itself, which is known to be the factor with the greatest efect size in prior ANOVA studies. This suggests that topic dificulty is an artifact of the collection considered and highlights the importance of further research in understanding link between the complexity of a topic and the query rewriting in IR related tasks.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The interplay between simple keyword queries and large document collections has challenged
researchers in Information Retrieval (IR) for more than half a century. Some queries are highly
efective, while others perform poorly, and changing the ranking models to compensate for
dificult queries can have negative efects on the performance of queries that were performing
well previously. This notion of query dificulty has received a great deal of attention over the
years. For example, NIST ran the Robust Track in 2004 and 2005 to reexamine sets of queries
which had performed poorly across all systems evaluated in the Ad hoc track [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It is clear that
certain queries challenge even the best performing systems. The distinction between a topic
(information need) and a query can have a profound impact on the efectiveness of retrieval, as
well as how IR researchers typically categorize and compare system performance [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. In this
paper, we reexamine the idea of query dificulty from the topic perspective, where a topic can
have many diferent query formulations and the retrieval system and the underlying document
collection can change. We explore this issue by addressing the following research questions:
RQ1: How does the formulation of a topic impact system performance within corpora?
RQ2: How does the formulation of a topic impact system performance across corpora?
RQ3: How does topic dificulty vary across corpora based on the formulation of a topic?
RQ1 allows us to investigate the efect size of topics and query formulations with respect to
systems and their components in order to better understand what contributes to topic dificulty
and the magnitude of the efects. RQ2 extends RQ1 by looking at what happens across corpora
and allows us to also explore corpora-specific topic / query formulations. Finally, RQ3 examines
the topic dificulty across multiple corpora. To address our research questions we develop
a set of ANalysis Of VAriance (ANOVA) models which allow us to break down the overall
system performance into topic, query formulation, system, and corpora efects.Additionally, to
investigate RQ3 more deeply, we measure variance in arbitrarily ranked topics across corpora.
The key idea is that the likelihood of observing arbitrary rank orderings of topics by efectiveness
is analogous to topic dificulty being an intrinsic property. That is, high volatility in topic
ordering suggests that topic hardness is not absolute. Rather it is an artifact of system / corpora
interaction. Our experiments highlights that the idea of a single topic being dificult is an artifact
of collection design, and the topic dificulty can reliably be circumvented through careful query
reformulation. This is a promising step in a fundamentally important problem in IR — that
of robust system efectiveness. The paper is organized as follows: Section 2 discusses the
experimental setup and the experimental findings; finally, Section 3 draws some conclusions
and outlooks for future work.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Experiments</title>
      <p>
        Data and Methods. We used the following collections: TREC Robust 2004 Ad Hoc, TREC
Common CORE 2017, and TREC Common CORE 2018 for our experiments. The Robust Ad
Hoc track contains approximately 528K documents; the TREC 2017 Common CORE contains
over 1.8 million articles; finally, the TREC Common CORE 2018 track roughly containing 600K
news articles. A large seed set of 3,402 human curated query formulations originally developed
using the TREC Robust 2004 Ad Hoc search collection were used in our experiments [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
RQ1: Efect Size of Query Formulation within Corpora .
      </p>
      <p>
        To determine the impact of the topics, their formulations and the systems on the overall
performance, we run an ANOVA [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] on a Grid of Points (GoP) of 144 diferent systems (9
ranking functions, 2 stemmers, 4 query expansion approaches and 2 stoplists). We also include
in the ANOVA model the interaction between the systems and the topics. Table 1 provides a
summary of the ANOVA efect size for each factor using each corpus separately. We observe
similar performance trends across all of them. All factors are statistically significant. The topic
factor has a large-size efect size, and it is indeed the largest efect for this configuration. We
can also clearly see that query formulations also have a large efect size in our experiments
– approaching the topic efect size – suggesting that query formulations strongly influence
topic dificulty. Overall, query formulation has the second largest efect, with nearly 1.5 times
the size of the topic*system interaction which has historically been a point of emphasis in
similar performance comparisons. This provides important evidence that query formulation is
crucial in retrieval efectiveness, and has deeper implications in rethinking the way many IR
experiments currently formalize query / topic dificulty.
      </p>
      <p>RQ2 and RQ3: Efect Size of Query Formulation across Corpora . To study the efect of
the topics and query formulations across corpora, we slightly change the ANOVA model, by
adding the additional corpus factor. Table 2 shows the results for the ANOVA across-corpora.
The introduction of the new factor allows to compute the efect of some interactions between
factors that could not be computed previously. All the factors are statistically significant. We can
observe that the topic factor has a large-size efect even across corpora and that the system factor
becomes a moderately large-size efect, being bigger than in the single corpus case (see Table 1);
the corpus factor has a medium-size efect. We also note that the query formulation factor has a
remarkably large-size efect, even across corpora, observed here for the first time, suggesting
it is a key contributor to topic dificulty. Both the query formulation*corpus interaction and
the topic*system*corpus interaction, observed here for the first time, are clearly important
large-size efects. Overall, these findings provide further evidence supporting the possibility
that dificult topics do not actually exist in any absolute sense.</p>
      <p>Topic Dificulty . To determine whether the dificulty is an intrinsic property of the topic
we execute the following experiment. We randomly sample 20,000 permutations of the topics.
For each of such permutations and for each pair system-corpus, using a greedy approach we
select formulation to represent each of the topics to maximize the correlation between the
ranking of the topics based on the Average Precision (AP) and the random permutation. We then
select, among all the rankings of topics, the one that maximizes the correlation with the random
permutation of topics. We finally compute the Kendall’s  correlation between the sampled
and the constructed rankings of topics. We have a mean Kendall’s  of 0.85, indicating that the
queries selected to induce the desired topic rankings were consistently close to the arbitrary
target ordering. This is a strong evidence that topics can be “arbitrarily” easy or dificult across
many diferent formulations, corpora and system combinations. This is empirical evidence
that topic dificulty is not an intrinsic property of an information need – meaning that query
formulation based on a corpus and retrieval system, can be combined to sort topics arbitrarily
based on a performance goal. How can such a finding help researchers develop more efective
retrieval systems? Firstly, it is worth noting that current evaluation paradigms usually consider
a single formulation for each topic. Such an arrangement prevents us to observe system behavior
with small changes to each query. We believe that multiple formulations are a key omission
in our current evaluation campaigns, and we are hopeful future campaigns will incorporate
them into their methodology. If our goal is to model real performance of systems, collections
creators should explore how to best include multiple formulations of each topic. Our isolation of
multiple formulations of topics has allowed us to study in detail the concept of “topic dificulty”,
which is construct of a specific retrieval configuration – the collection, the system and the query
which represents the topic – and not a property intrinsic to the topic alone.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusion</title>
      <p>In this work, we have presented a comprehensive ANOVA analysis that compares the efect sizes
across multiple corpora and retrieval system configurations. We have also generalized previous
model configurations in order to incorporate a new nesting factor which maps an information
need (topic) to multiple query formulations. The removal of the constraint of a 1:1 mapping
between a query and a topic has led to several interesting observations which have important
implications on the notion of topic dificulty. We also propose an analysis methodology, based
on a permutation algorithm, to further explore topic dificulty. Based on this new knowledge,
we were able to show conclusive evidence that topic dificulty is not an intrinsic property and
therefore query formulations should be included in future evaluation campaigns.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Culpepper</surname>
          </string-name>
          , G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Oren</surname>
          </string-name>
          ,
          <article-title>Topic dificulty: Collection and query formulation efects</article-title>
          ,
          <source>Transactions on Information Systems</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          ,
          <article-title>Overview of the trec 2004 robust retrieval track</article-title>
          .,
          <source>in: Proc. TREC</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bailey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mofat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Scholer</surname>
          </string-name>
          , P. Thomas,
          <article-title>User Variability and IR System Evaluation</article-title>
          ,
          <source>in: Proc. of SIGIR</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>625</fpage>
          -
          <lpage>634</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Faggioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Zendel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Culpepper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Scholer</surname>
          </string-name>
          ,
          <article-title>An enhanced evaluation framework for query performance prediction</article-title>
          ,
          <source>in: Proc. ECIR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>115</fpage>
          -
          <lpage>129</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Benham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Culpepper</surname>
          </string-name>
          ,
          <article-title>Risk-reward trade-ofs in rank fusion</article-title>
          ,
          <source>in: Proc. ADCS</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Banks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Over</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.-F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Blind Men and
          <article-title>Elephants: Six Approaches to TREC data</article-title>
          ,
          <source>Information Retrieval 1</source>
          (
          <year>1999</year>
          )
          <fpage>7</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Faggioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <article-title>System efect estimation by sharding: A comparison between anova approaches to detect significant diferences</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>