<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scales of Evaluation Measures: From Theory to Experimentation∗</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvia Pontarollo</string-name>
          <email>spontaro@math.unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Ferrante</string-name>
          <email>ferrante@math.unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eleonora Losiouk</string-name>
          <email>elosiouk@math.unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Information Engineering, University of Padua</institution>
          ,
          <addr-line>Padua</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Mathematics, University of Padua</institution>
          ,
          <addr-line>Padua</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Evaluation measures are the basis for quantifying the performance
of IR systems and measurement scales play a central role since
they determine the operations that can be performed with the
measured values and, as a consequence, the statistical analyses that
can be applied. Stevens [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] identifies four major types of scales
with increasing properties: (i) the nominal scale consists of discrete
unordered values, i.e. categories; (ii) the ordinal scale introduces a
natural order among the values; (iii) the interval scale preserves the
equality of intervals or diferences; and (iv) the ratio scale preserves
the equality of ratios. For example, mean and variance should be
computed only when relying on interval scales.
      </p>
      <p>
        We present our formal theory of IR evaluation measures [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
based on the representational theory of measurement [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ], to
determine whether and when IR measures are interval scales.
      </p>
      <p>We found that common set- based retrieval measures – namely
Precision, Recall, and F-measure – always are interval scales in the
case of binary relevance while this does not happen in the
multigraded relevance case. In the case of rank-based retrieval measures
– namely AP, gRBP, DCG, and ERR – only gRBP is an interval scale
when we choose a specific value of the parameter p and define a
specific total order among systems while all the other IR measures
are not interval scales. We also introduce some brand new set-based
and rank-based IR evaluation measures which ensure to be interval
scales.</p>
      <p>
        Finally, we discuss the outcomes of an extensive evaluation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
based on standard TREC collections, to study how our theoretical
ifndings impact on the experimental ones. In particular, we report
here a correlation analysis to study the relationship among the
above-mentioned state-of-the-art evaluation measures and their
scales.
      </p>
    </sec>
    <sec id="sec-2">
      <title>SET-BASED MEASURES</title>
      <p>
        Let us start by introducing an order relation ⪯ on the set of judged
runs R(N). Let rˆ, sˆ ∈ R(N ) such that rˆ , sˆ, and let k be the biggest
relevance degree at which the two runs difer for the first time, i.e.
∗Extended abstract of [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]
k = max{j ≤ c : {i : rˆi = aj } , {i : sˆi = aj } }. We strictly order
any pair of distinct system runs as follows
rˆ ≺ sˆ ⇔
{i : rˆi = ak } &lt; {i : sˆi = ak } .
(1)
      </p>
      <p>R(N ) is a totally ordered set with respect to the ordering ⪯
defined by (1). As for any totally order set, R(N ) is a poset consisting
of only one maximal chain (the whole set); therefore it is graded of
rank |R(N )| − 1, where R(N ) = NN+c since it consists of all the N
combinations of c + 1 = |REL| objects with repetition. Since R(N )
is graded of rank |R(N )| − 1, there exists a unique rank function
ρ(rˆ) : R(N ) −→ N such that ρ(0ˆ) = 0 and ρ(sˆ) = ρ(rˆ) + 1 if sˆ covers
rˆ:
ρ(rˆ) =
ÕN δa(rˆj ) + N − j ,
j=1 N − j + 1
where rˆ = {rˆ1, . . . , rˆN } ∈ R(N ) with rˆi ⪯ rˆi+1 for any i &lt; N .</p>
      <p>The natural distance is then given by ℓ(rˆ, sˆ) = ρ(sˆ) − ρ(rˆ), for
rˆ, sˆ ∈ R(N ) such that rˆ ⪯ sˆ, and we can define the diference as
∆ rˆsˆ = ℓ(rˆ, sˆ) if rˆ ⪯ sˆ, otherwise ∆ rˆsˆ = −ℓ(sˆ, rˆ). (R(N ), ⪯d ) is a
diference structure. Thus the rank function is an interval scale and
we are able to define a new interval-scale measure that follows:</p>
      <sec id="sec-2-1">
        <title>The Set-Based Total Order (SBTO) measure on</title>
      </sec>
      <sec id="sec-2-2">
        <title>Definition 2.1.</title>
        <p>(R(N ), ⪯d ) is:
SBTO(rˆ) = ρ(rˆ) =
ÕN δa(rˆj ) + N − j .
j=1 N − j + 1
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>RANK-BASED MEASURES</title>
      <p>Top-heaviness is a central property in Information Retrieval (IR),
stating that the higher a system ranks relevant documents the
better it is. If we apply this property at each rank position and we
take to extremes the importance of having a relevant document
ranked higher, we can define a strong top-heaviness property which,
in turn, will induce a total ordering among runs.</p>
      <p>Let rˆ, sˆ ∈ R(N ) such that rˆ , sˆ, then there exists k = min{j ≤
N : rˆ[j] , sˆ[j]} &lt; ∞, and we order system runs as follows
rˆ ≺ sˆ ⇔ rˆ[k] ≺ sˆ[k] .</p>
      <p>
        This order prefers a single relevant document ranked higher to
any number of relevant documents, with same relevance degree or
higher, ranked just below it
(uˆ[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], . . . , uˆ[m], a0, ac , . . . , ac ), ≺ (uˆ[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], . . . , uˆ[m], aj , a0, . . . , a0) ,
(2)
(3)
(4)
      </p>
      <sec id="sec-3-1">
        <title>Binary Relevance – T08</title>
        <p>Measure Pair Topic-by-Topic
Precision vs SBTO 1.0000
Recall vs SBTO 1.0000
F-measure vs SBTO 1.0000
Precision vs Recall 1.0000
SBTO vs RBTO 0.4358</p>
      </sec>
      <sec id="sec-3-2">
        <title>Multi-graded Relevance – T26</title>
        <p>Measure Pair Topic-by-Topic
Generalized Precision vs SBTO 0.7325
Generalized Recall vs SBTO 0.7325
Generalized Precision vs Generalized Recall 1.0000
SBTO vs RBTO 0.3895
for any 1 ≤ j ≤ c, for any length N ∈ N and any m ∈ {0, 1, . . . , N −
1}. This is why we call it strong top-heaviness.</p>
        <p>R(N ) is totally ordered with respect to ⪯ and is graded of rank
(c + 1)N − 1. Therefore, there is a unique rank function ρ : R(N ) −→
{0, 1, . . . , (c + 1)N − 1} which is given by:
where δa is the indicator function.</p>
        <p>
          Let us set δa (rˆ) = (δa(rˆ[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]), . . . , δa(rˆ[N ])). If we look at δa (rˆ)
as a string, the rank function is exactly the conversion in base 10
of the number in base c + 1 identified by δa (rˆ) and the ordering
among runs ⪯ corresponds to the ordering ≤ among numbers in
base c + 1.
        </p>
        <p>The natural distance is then given by ℓ(rˆ, sˆ) = ρ(sˆ) − ρ(rˆ), for
rˆ, sˆ ∈ R(N ) such that rˆ ⪯ sˆ, and we can define the diference
as ∆ rˆsˆ = ℓ(rˆ, sˆ) if rˆ ⪯ sˆ, otherwise ∆ rˆsˆ = −ℓ(sˆ, rˆ). (R(N ), ⪯d ) is
a diference structure. As done before in the set-based case, an
interval scale measure on (R(N ), ⪯d ) is given by the rank function
itself.</p>
        <p>Definition 3.1. The Rank-Based Total Order (RBTO) interval-scale
measure on (R(N ), ⪯d ) is:</p>
        <p>RBTO(rˆ) = ρ(rˆ) =</p>
        <p>N
Õ δa(rˆ[i])(c + 1)N −i
i=1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 EXPERIMENTS</title>
      <p>We explore the following research question: “How to scales
determine the relationship among evaluation measures?”, i.e. what is the
relationship between measures which are interval scales, ordinal
scale or which are not on any scale? To this end, we will perform
Kendall’s τ correlation analysis on TREC 08 Ad-hoc (binary
judgements) and TREC 26 Core (multi-graded judgments) collections.</p>
      <p>Tables 1 and 2 report the correlation analysis in the case of
set-based and rank-based evaluation measures for both binary and
multi-graded relevance.</p>
      <p>Note that two interval scale measures order systems in the same
way on the same topic and their correlation must be 1.0. However,
this may be no more true, if you first average performance across
(5)
(6)</p>
      <sec id="sec-4-1">
        <title>Binary Relevance – T08</title>
        <p>Measure Pair Topic-by-Topic
RBP p = 1/2 vs RBTO 1.0000
RBP p = 0.2 vs RBTO 0.9985
RBP p = 0.8 vs RBTO 0.8553
AP vs RBTO 0.6099</p>
      </sec>
      <sec id="sec-4-2">
        <title>Multi-graded Relevance – T26</title>
        <p>
          Measure Pair Topic-by-Topic
gRBP p = 1/3 vs RBTO 1.0000
gRBP p = 1/3, W3 = [
          <xref ref-type="bibr" rid="ref1 ref3">0, 1, 3</xref>
          ] vs RBTO 0.9867
gRBP p = 0.2 vs RBTO 0.9996
gRBP p = 0.8 vs RBTO 0.7420
DCG vs RBTO 0.3774
ERR vs RBTO 0.9468
RBTO W1 = [
          <xref ref-type="bibr" rid="ref1 ref2">0, 1, 2</xref>
          ] vs RBTO W2 = [
          <xref ref-type="bibr" rid="ref2 ref4">0, 2, 4</xref>
          ] 1.0000
RBTO W1 = [
          <xref ref-type="bibr" rid="ref1 ref2">0, 1, 2</xref>
          ] vs RBTO W3 = [
          <xref ref-type="bibr" rid="ref1 ref3">0, 1, 3</xref>
          ] 0.9866
all the topics and then compute the correlation, which is the typical
way of computing Kendall’s τ correlation [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>This can be, for example, observed in Table 1 where Precision,
Recall, F-measure, and SBTO are all transformation of the same
interval scale and thus their topic-by-topic correlation is 1; on
the other hand, their overall correlation, i.e. the traditional one,
is diferent from 1.0 because of the efect of the recall base when
averaging across topics. This suggest that the diference between
Precision and Recall (τ = 0.85) is not due to them ranking systems
diferently but just to the fact that the recall base alters the scale
properties from topic to topic.</p>
        <p>Another interesting case is RBP. For p = 1/2 (p = 1/3 in the
multigraded case) it is an interval-scale; for p &lt; 1/2 it is an ordinal
but not interval scale and its correlation starts departing from 1.0;
the efect is much more pronounced for RBP with p &gt; 1/2 which
is neither an ordinal nor an interval scale, suggesting that simply
acting on a parameter of a measure can completely alter its scale
properties.</p>
        <p>
          A final interesting case is RBTO with diferent weights for the
relevance degrees: W1 = [
          <xref ref-type="bibr" rid="ref1 ref2">0, 1, 2</xref>
          ] vs W2 = [
          <xref ref-type="bibr" rid="ref2 ref4">0, 2, 4</xref>
          ] keep the RBTO
on an interval-scale while W1 = [
          <xref ref-type="bibr" rid="ref1 ref2">0, 1, 2</xref>
          ] vs W3 = [
          <xref ref-type="bibr" rid="ref1 ref3">0, 1, 3</xref>
          ] show that
it stops to be an interval-scale. Indeed, our theoretical findings [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
demonstrate that, in the multi-graded case, the interval-scale
property is complied with only if the weights of the relevance degrees
are on a ratio scale, which is not the case for W3 = [
          <xref ref-type="bibr" rid="ref1 ref3">0, 1, 3</xref>
          ].
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ferrante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Losiouk</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>How do interval scales help us with better understanding IR evaluation measures? Information Retrieval Journal (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ferrante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Pontarollo</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A General Theory of IR Evaluation Measures</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering (TKDE) 31, 3 (March</source>
          <year>2019</year>
          ),
          <fpage>409</fpage>
          -
          <lpage>422</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. H.</given-names>
            <surname>Krantz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. D.</given-names>
            <surname>Luce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Suppes</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Tversky</surname>
          </string-name>
          .
          <year>1971</year>
          .
          <article-title>Foundations of Measurement. Additive and Polynomial Representations</article-title>
          . Vol.
          <volume>1</volume>
          . Academic Press, USA.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Stevens</surname>
          </string-name>
          .
          <year>1946</year>
          .
          <article-title>On the Theory of Scales of Measurement</article-title>
          . Science, New Series 103,
          <issue>2684</issue>
          (
          <year>June 1946</year>
          ),
          <fpage>677</fpage>
          -
          <lpage>680</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Variations in relevance judgments and the measurement of retrieval efectiveness</article-title>
          .
          <source>In SIGIR</source>
          <year>1998</year>
          ,
          <volume>315</volume>
          -
          <fpage>323</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>