<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Eect of Inter-Assessor Disagreement on IR System Evaluation: A Case Study with Lancers and Students</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tetsuya Sakai</string-name>
          <email>tetsuyasakai@acm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Waseda University</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>31</fpage>
      <lpage>38</lpage>
      <abstract>
        <p>is paper reports on a case study on the inter-assessor disagreements in the English NTCIR-13 We Want Web (WWW) collection. For each of our 50 topics, pooled documents were independently judged by three assessors: two “lancers” and one Waseda University student. A lancer is a worker hired through a Japanese part time job matching website, where the hirer is required to rate the quality of the lancer's work upon task completion and therefore the lancer has a reputation to maintain. Nine lancers and ve students were hired in total; the hourly pay was the same for all assessors. On the whole, the inter-assessor agreement between two lancers is statistically signicantly higher than that between a lancer and a student. We then compared the system rankings and statistical signicance test results according to dierent qrels versions created by changing which asessors to rely on: overall, the outcomes do dier according to the qrels versions, and those that rely on multiple assessors have a higher discriminative power than those that rely on a single assessor. Furthermore, we consider removing topics with relatively low inter-assessor agreements from the original topic set: we thus rank systems using 27 high-agreement topics, aer removing 23 low-agreement topics. While the system ranking with the full topic set and that with the high-agreement set are statistically equivalent, the ranking with the high-agreement set and that with the low-agreement set are not. Moreover, the low-agreement set substantially underperforms the full and the high-agreement sets in terms of discriminative power. Hence, from a statistical point of view, our results suggest that a high-agreement topic set is more useful for nding concrete research conclusions than a low-agreement one.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>•Information systems ! Retrieval eectiveness;
inter-assessor agreement; p-values; relevance assessments;
statistical signicance</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>While IR researchers oen view laboratory IR evaluation results as
something objective, at the core of any laboratory IR experiments
lie the relevance assessments, which are the result of subjective
judgements of documents by a person, or multiple persons, based
on a particular (intepretation of an) information need. Hence it is
of utmost importance for IR researchers to understand the eects
Copying permied for private and academic purposes.</p>
      <p>EVIA 2017, co-located with NTCIR-13, Tokyo, Japan.
© 2017 Copyright held by the author.
of the subjective nature of the relevance assessment process on the
nal IR evaluation results.</p>
      <p>
        is paper reports on a case study on the inter-assessor
disagreements in a recently-constructed ad hoc web search test
collection, namely, the English NTCIR-13 We Want Web (WWW)
collection [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. For each of our 50 topics, pooled documents were
independently judged by three assessors: two “lancers” and one
Waseda University student. A lancer is a worker hired through
a Japanese part time job matching website1, where the hirer is
required to rate the quality of the lancer’s work upon task
completion and therefore the lancer has a reputation to maintain2. Nine
lancers and ve students were hired in total; the hourly pay was the
same for all assessors. On the whole, the inter-assessor agreement
between two lancers is statistically signicantly higher than that
between a lancer and a student (Section 3). We then compared
the system rankings and statistical signicance test results
according to dierent qrels versions created by changing which asessors
to rely on: overall, the outcomes do dier according to the qrels
versions, and those that rely on multiple assessors have a higher
discriminative power (i.e., the ability to obtain many statistically
signicant system pairs [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]) than those that rely on a single
assessor (Section 4.1). Furthermore, we consider removing topics with
relatively low inter-assessor agreements from the original topic
set: we thus rank systems using 27 high-agreement topics, aer
removing 23 low-agreement topics. While the system ranking with
the full topic set and that with the high-agreement set are
statistically equivalent, the ranking with the high-agreement set and that
with the low-agreement set are not. Moreover, the low-agreement
set substantially underperforms the full and the high-agreement
sets in terms of discriminative power (Section 4.2). Hence, from a
statistical point of view, our results suggest that a high-agreement
topic set is more useful for nding concrete research conclusions
than a low-agreement one.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK/NOVELTY OF OUR WORK</title>
      <p>
        Studies on the eect of inter-assessor (dis)agreement on IR system
evaluation have a long history; Bailey et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] provides a concise
survey on this topic covering the period 1969-2008. More recent
work in the literature includes Carteree and Soboro [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Webber,
Chandar, and Carteree [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], Demeester et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] Megorskaya,
Kukushkin, and Serdyukov [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], Wang et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], Ferrante, Ferro,
and Maistro [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and Maddalena et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Among these studies, the
work of Voorhees [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] from 2000 (or the earlier version reported at
SIGIR 1998) is probably one of the most well-known; below, we rst
1 hp://www.lancers.jp/ (in Japanese). See also hps://www.techinasia.com/
lancers-produces-200-million-freelancing-gigs-growing (in English)
2 e lancer then rates the hirer; therefore the hirer also has a reputation to maintain
on the website.
highlight the dierences between her work and the present study,
since the primary research question of the present study is whether
her well-known ndings generalise to our new test collection with
experimental seings that are quite dierent from hers in several
ways. Aer that, we also briey compare the present study with
the recent, closely-related work of Maddalena et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] from ICTIR
2017.
      </p>
      <p>
        Voorhees [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] examined the eect of using dierent qrels
versions on ad hoc IR system evaluation. Her experiments used the
TREC-4 and TREC-6 data3. In particular, in her experiments with
the 50 TREC-4 topics, she hired two additional assessors in addition
to the primary assessor who created the topic, and discussed the
pairwise inter-assessor agreement in terms of overlap as well as
recall and precision: overlap is dened as the size of the
intersection of two relevant sets divided by the size of the union; recall
and precison are dened by one of the relevant set as the gold
data. However, it was not quite the case that the three assessors
judged the same document pool independently: the document sets
provided to the additional assessors were created aer the primary
assessment, by mixing both relevant and nonrelevant documents
from the primary assessor’s judgements. Moreover, all documents
judged relevant by the primary assessor but not included in the
document set for the additional assessors were counted towards
the set intersection when computing the inter-assessor agreement.
Her experiments with the TREC-6 experiments relied on a dierent
seing, where University of Waterloo created their own pools and
relevance assessments independent of the original pools and
assessments. She considered binary relevance only4, and therefore she
considered Average Precision and Recall at 1000 as eectiveness
evaluation measures. Her main conclusion was: “e actual value
of the eectiveness measure was aected by the dierent conditions,
but in each case the relative performance of the retrieved runs was
almost always the same. ese results validate the use of the TREC
test collections for comparative retrieval experiments.”
      </p>
      <p>e present study diers from that of Voorhees in the following
aspects at least:</p>
      <p>We use a new English web search test collection contructed
for the NTCIR-13 WWW task, with depth-30 pools.</p>
      <p>For each of our 50 topics, the same pool was completely
independently judged by three assessors. Nine assessors
were hired through the lancers website, and an additional
ve assessors were hired at Waseda University, so that
each topic was judged by two lancers and one student.
We collected graded relevance assessments from each
assessor: highly relevant (2 points), relevant (1 point),
nonrelevant (0) and error (0) for cases where the web pages
to judge could not be displayed. When consolidating the
multiple assessments, the raw scores were added to form
more ne-grained graded relevance data.</p>
      <p>
        We use graded relevance measures at cuto 10
(representing the quality of the rst search engine result page),
namely nDCG@10, Q@10, and nERR@10 [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], which are
the ocial measures of the WWW task.
3 e document collections are: disks 2 and 3 for TREC-4; disks 4 and 5 for TREC-6 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
4 e original Waterloo assessments on a tertiary scale, but were collapsed into binary
for her analysis.
      </p>
      <p>
        As our topics were sampled from a query log, none of our
assessors are the topic originators (or “primary” [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] or
“gold” assessors [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]); e assessors were not provided with
any information other than the query (e.g., a narrative
eld [
        <xref ref-type="bibr" rid="ref1 ref8">1, 8</xref>
        ]): the denition for a highly relevant document
was: “it is likely that the user who entered this search query
will nd this page relevant”; that for a relevant document
was: “it is possible that the user who entered this search
query will nd this page relevant ” [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        We discuss inter-assessor agreeement and system ranking
agreement using stastical tools, namely, linear weighted κ
with 95%CIs (which, unlike raw overlap measures, takes
chance agreement into account [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) and Kendall’s τ with
95%CIs. Moreover, we employ the randomised Tukey HSD
test [
        <xref ref-type="bibr" rid="ref16 ref3">3, 16</xref>
        ] to discuss the discrepancies in statistical
signicance test results. Furthermore, we consider removing
topics that appear to be unreliable in terms of inter-assessor
agreement.
      </p>
      <p>
        While the recent work of Maddalena et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] addressed
several research questions related to inter-assessor agreement, one
aspect of their study is closely related to our analysis with
highagreement and low-agreement topic sets. Maddalena et al. utilised
the TREC 2010 Relevance Feedback track data and exactly ve
dierent relevance assessments for each ClueWeb document, and
used Krippendorph’s α [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] to quantify the inter-assessor agreement.
ey dened high-agreement and lowagreeement topics based on
Krippendorph’s α , and reported that high-agreement topics can
predict the system ranking with the full topic set more accurately
than low-agreement topics. e analysis in the present study diers
from the above as discussed below:
      </p>
      <p>Krippendorph’s α disregards which assessments came from
which assessors, as it is a measure of the overall reliability
of the data. In the present study, where we only have three
assessors, we are more interested in the agreement between
every pair of assessors and hence utilise Cohen’s linear
weighted κ. Hence our denition of a high/low-agreement
topic diers from that of Maddalena et al.: according to
our denition, a topic is in high agreement if the κ is
statistically signicantly positive for every pair of assessors.
While Maddalena et al. discussed the absolute
eectiveness scores and system rankings only, the present study
discusses statistical signicance testing aer replacing the
full topic set with the high/low-agreement set.</p>
      <p>
        Maddalena et al. focussed on nDCG; we discuss the three
aforementioned ocial measures of the NTCIR-13 WWW
task [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
3
      </p>
      <p>DATA
e NTCIR-13 WWW English subtask created 100 test topics; 13
runs were submied from three teams. We acknowledge that this
is a clear limitation of the present study: we would have liked a
larger number of runs from a larger number of teams. However, we
claim that this limitation does not invalidate neither our approach
to analysing inter-assessor disagreement nor the actual results on
the system ranking and statistical signicance. We hope to repeat
the same analysis on a larger set of runs in the next round of the
WWW task.</p>
      <p>For evaluating the 13 runs submied to the WWW task, we
created a depth-30 pool for each of the 100 topics and this resulted
in a total of 22,912 documents to judge. We hired nine lancers who
speak English through the lancers website: the job call and the
relevance assessment instructions were published on the website
in English. None of them had any prior experience in relevance
assessments. Topics were assigned at random to the nine assessors
so that each topic had two independent judgments from two lancers.
e ocial relevance assessments of the WWW task were formed
by consolidating the two lancer scores: since each lancer gave 0, 1,
or 2, the nal relevance levels were L0-L4.</p>
      <p>For the present study, we focus on a subset of the above test set,
which contains 50 topics whose topic IDs are odd numbers. e
number of pooled documents for this topic set is 11,214. We then
hired ve students from the Department of Computer Science and
Engineering, Waseda University, to provide a third set of judgments
for each of the 50 topics. e instructions given to them were
identical to those given to the lancers. e students also did not
have any prior experience in relevance assessments. Moreover,
lancers and students all received an hourly pay of 1,200 Japanese
Yen. However, hiring lancers is more expensive, because we have
to pay about 20% to Lancers the company on top of what we pay
to the individual lancers. e purpose of collecting the third set
of assessments was to compare the lancer-lancer inter-assessor
agreement with the lancer-student agreement, which should shed
some light on the reliability of the dierent assessor types. All of
the assessors completed the work in about one month.</p>
      <p>
        It should be noted that all of our assessors are “bronze”
according to the denition by Bailey et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]: they are neither topic
originators nor topic experts.
      </p>
      <p>To quantify inter-assessor agreement, we compute Cohen’s
linear weighted κ for every assessor pair, where, for example, the
weight for a ¹0; 2º disagreement is 2 and that for a ¹0; 1º or ¹1; 2º
disagreement is 15. It should be noted that κ represents how much
agreement there is beyond chance.</p>
      <p>Table 1 summarises the inter-assessor agreement results. e
“raw scores” section shows the 3 3 confusion matrices for each
pair of assessors; the counts were summed across topics, although
lancer1, lancer2, and student are actually not single persons. e
linear weighted κ’s were computed based on these matrices. It can
be observed that the lancer-lancer κ is statistically signicantly
higher than the lancer-student κ’s, which means that the lancers
agree with each other more than they do with students. While the
lack of gold data prevents us from concluding that lancers are more
reliable than students, it does suggest that lancers are worth hiring
if we are looking for high inter-assessor agreement. As we shall see
in Section 4.1, the discriminative power results discussed in also
support this observation.</p>
      <p>We also computed per-topic linear weighted κ’s so that the
assessments of exactly three individuals are compared against one
another: the mean, minimum and maximum values are also
provided in Table 1. It can be observed that the lowest per-topic κ
observed is 0:128, for “lancer1 vs. student”; the 95%CI for this
instance was » 0:262; 0:0006¼, suggesting the lack of agreement
beyond chance6.</p>
      <p>
        Table 1 also shows the number of topics for which the per-topic
κ’s were not statistically signicantly positive, that is, the 95%CI
lower limits were not positive, as exemplied by the above instance.
ese numbers indicate that the lancer-lancer agreements were
statistically signicantly positive for 50 8 = 42 topics while the
lancer-student agreements were statistically signicantly positive
for only 35 (31) topics. Again, the lancers agree with each other
more than they do with students.
5 Fleiss’ κ [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], designed for more than two assessors, is applicable to nominal categories
only; the same goes for Randoph’s κfree [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]; see also our discussion on Krippendor’s
α [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] in Section 2.
6 It should be noted that negative κ’s are not unsual in the context of inter-assessor
agreement: for example, according to a gure from Bailey et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], when a “gold”
assessor (i.e., top originator) was compared with a bronze asessor, a version of κ was
in the » 0:6; 0:4¼ range for one topic, despite the fact that the assessors must have
read the narrative elds of the TREC Enterprise 2007 test collection [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>ere were 27 topics where all three per-topic κ’s were
statistically signicantly positive; for the remaining 23 topics, at least one
per-topic κ was not. We shall refer to the set of the former 27 topics
as the high-agreement set and the laer as the low-agreement set.
We shall utilise these subsets in Section 4.2.</p>
      <p>Also in Table 1, the binary Cohen’s κ row shows the κ values
aer collapsing the 3 3 matrices into 2 2 matrices by treating
highly relevant and relevant as just relevant. Again, the
lancerlancer κ is statistically signicantly higher than the lancer-student
κ’s. Finally, the table shows the raw agreement based on the 2 2
confusion matrices: the counts of ¹0; 0º and ¹1; 1º are divided by
those of ¹0; 0º, ¹0; 1º, ¹1; 0º, and ¹1; 1º. It can be observed that only
the lancer-lancer agreement exceeds 70%.</p>
      <p>We summed the raw scores of lancer1, lancer2, and student to
form a qrels set which we call all3; we also summed the raw scores
of lancer1 and lancer2 to form a qrels set which we call 2lancers.
Table 2 shows the distribution of documents across the relevance
levels. Note that all3 and 2lancers are on 7-point and 5-point scales,
respectively, while the others are on a 3-point scale. In this way,
we preserve the views of individual assessors instead of collapsing
the assessments into binary or to force them to reach a consensus.
Note that nDCG@10, Q@10, and nERR@10 can fully utilise the
rich relevance assessments. As we shall see in the next section, this
approach to combining multiple relevance assessments is benecial.</p>
      <p>
        For alternatives to simply summing up the raw assessor scores,
we refer the reader to Maddalena et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and Sakai [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]: these
approaches are beyond the scope of the present study.
4
4.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>RESULTS AND DISCUSSIONS</title>
      <p>Dierent Qrels Versions
e previous section showed that assessors do disagree, and that
the lancers agree with each other more than they do with students.
is section investigates the eect of inter-assessor disagreement
on system ranking and statistical signicance through comparisons
across the ve qrels versions: all3, 2lancers, lancer1, lancer2, and
student.</p>
      <p>
        4.1.1 System Ranking. Figure 1 visualises the system rankings
and the actual mean eectiveness scores according to the ve
different qrels, for nDCG@10, Q@10, and nERR@10. In each graph,
the runs have been sorted by the all3 scores, and therefore if every
curve is monotonically decreasing, that means all the qrels versions
produce system rankings that are identical to the one based on all3.
First, it can be observed that the absolute eectiveness scores dier
depending on the qrels version used, just as Voorhees [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] observed
with Average Precision and Recall@1000. Second, and more
importantly, the ve system rankings are not identical: for example, in
Figure 1(c), the top performing run according to nERR@10 with
all3 is only the h best run according to the same measure with
lancer1. e nDCG@10 and Q@10 curves are relatively consistent
across the qrels versions. Table 3 quanties the above observation
in terms of Kendall’s τ , with 95%CIs: while the CI upper limits
show that all the rankings are statisticallly equivalent, the widths
of the CIs due to the small sample size (13 runs) suggest that the
results should be viewed with caution. For example, the τ for the
aforementioned case of all3 vs. lancer1 with nERR@10 is 0.821
0.8
0.7
0.6
0.5
0.4
all3 2lancers lancer1 lancer2 student
(95%CI»0:616; 1:077¼). e actual system swaps that Figure 1 shows
are probably more important than these summary statistics.
      </p>
      <p>
        4.1.2 Statistically Significant Dierences across Systems. e
next and perhaps more important question is: how do the dierent
qrels versions aect pairwise statistical signicance test results? If
the researcher is interested in the dierence between every system
pair, a proper multiple comparison procedure should be employed
to ensure that the familywise error rate is bounded above by the
signicance criterion α [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this study we use the
distributionfree randomised Tukey HSD test using the Discpower tool7, with
B = 10; 000 trials [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. e input to the tool is a topic-by-run score
matrix: in our case, for every combination of evaluation measure
and qrels version, we have a 50 13 matrix.
7 hp://research.nii.ac.jp/ntcir/tools/discpower-en.html
      </p>
      <p>Table 4 shows the results of comparing the outcomes of
signicance test results at α = 0:05: for example, Table 4(a) shows that,
in terms of nDCG@10, all3 obtained 2 + 29 = 31 statistically
signicantly dierent run pairs, while 2lancers obtained 29 + 2 = 31,
and that the two qrels versions had 29 pairs in common. us the
Statistical Signicance Overlap (SSO) is 29¹2 + 29 + 2º = 87:9%. e
table shows that the student results disagree relatively oen with
the others: for example, in Table 4(b), Q@10 with lancer1 have 11
statistically signicantly dierent run pairs that are not statistically
signicantly dierent according to the same measure with student,
while the opposite is true for ve pairs. e two qrels versions have
15 pairs in common and the SSO is only 48.4%. us, dierent qrels
versions can lead to dierent research conclusions.</p>
      <p>
        Table 5 shows the number of statistically signicantly dierent
run pairs (i.e., discriminative power [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]) deduced from Table 4.
It can be observed that combining multiple assessors’ labels and
thereby having ne-grained relevance levels can result in high
discriminative power, and also that student underperforms the
others in terms of discriminative power. us, it appears that
student is not only dierent from the two lancers: they fail to
provide many signicantly dierent pairs.
4.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Using Reliable Topics Only</title>
      <p>In Section 3, we dened a high-agreement set containing 27 topics
and a low-agreement set containg 23 topics. A high-agreement
topic is one for which every assessor pair “statistically agreed,”
in the sense that the 95%CI lower limit of the per-topic κ was
positive. A low-agreement topic is one for which at least one
assessor pair did not show any agreement beyond chance, and
therefore deemed unreliable. While to the best of our knowledge
this kind of close analysis of inter-assessor agreement is rarely done
prior to evaluating the submied runs, removing such topics at an
early stage may be a useful practice for ensuring test collection
reliability. Hence, in this section, we focus on the all3 qrels, and
compare the evaluation outcomes when the full topic set (50 topics)
is replaced with just the high-agreement set or even just the
lowagreement set. e fact that the high-agreement and low-agreement
sets are similar in sample size is highly convenient for comparing
them in terms of discriminative power.</p>
      <p>4.2.1 System Ranking. Figure 2 visualises the system rankings
and the actual mean eectiveness scores according to the three
topic sets, for nDCG@10, Q@10, and nERR@10. Again, in each
graph, the runs have been sorted by the all3 scores (mean over 50
topics). Table 6 compares the system rankings in terms of Kendall’s
τ with 95%CIs. e values in bold indicate the cases where the two
rankings are statistically not equivalent. It can be observed that
while the system rankings by the full set and the high-agreement
set are statistically equivalent, those by the full set and the
lowagreement set are not. us, the properties of the high-agreement
topics appear to be dominant in the full topic set.
4.2.2 Statistically Significanct Dierences across Systems.
Table 7 compares the outcomes of statistical signicance test results
(Randomised Tukey HSD with B = 10; 000 trials) across the three
topic sets in a way similar to Table 4. Note that the two subsets
are inherently less discriminative than the full set as the sample
sizes are about half that of the full set. It can be observed that
the set of statistically signifantly dierent pairs according to the
high-agreement (low-agreement) set is always a subset of the set
of statistically signifantly dierent pairs according to the full topic
set. More interestingly, the set of statistically signifantly dierent
pairs according to the low-agreement set is almost a subet of the
set of statistically signifantly dierent pairs according to the
highagreement set: for example, Table 7(b) shows that there is only one
system pair for which the low-agreement set obtained a statistically
signicant dierence while the high-agreement set did not in terms
of Q@10.
all3 (27 good topics)
all3 (23 bad topics)</p>
      <p>CONCLUSIONS AND FUTURE WORK
is paper reported on a case study involving only 13 runs
contributed from only three teams. Hence we do not claim that our
nding will generalise; we merely hope to apply the same
methodology to test collections that will be created for the future rounds of
the WWW task and possibly even other tasks. Our main ndings
using the English NTCIR-13 WWW test collection are as follows:
Lancer-lancer inter-assessor agreements are statistically
signicantly higher than lancer-student agreements. e
student qrels is less discriminative than the lancers qrels.
While the lack of gold data prevents us from concluding
which type of assessors is more reliable, these results
suggest that hiring lancers has some merit despite the extra
cost.</p>
      <p>Dierent qrels versions based on dierent (combinations
of) assessors can lead to somewhat dierent system
rankings and statistical signicance test results. Combining
multiple assessors’ labels to form ne-grained relevance
levels is benecial in terms of discriminative power.</p>
      <p>Removing 23 low-agreement topics (in terms of inter-assessor
agreement) from the full set of 50 topics prior to evaluating
runs did not have a major impact on the evaluation results,
as the properties of the 27 high-agreement topics are
dominant in the full set. However, replacing the high-agreement
set with the low-agreement set resulted in a statistically
signicantly dierent system ranking, and substantially
lower discriminative power. Hence, from a statistical point
of view, our results suggest that a high-agreement topic set
is more useful for nding concrete research conclusions
than a low-agreement one.</p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGEMENTS</title>
      <p>I thank the PLY team (Peng Xiao, Lingtao Li, Yimeng Fan) of my
laboratory for developing the PLY relevance assessment tool and
collecting the assessments. I also thank the NTCIR-13 WWW task
organisers and participants for making this study possible.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Bailey</surname>
          </string-name>
          , Nick Craswell adn Arjen P. de Vries, and Ian Soboro.
          <year>2008</year>
          .
          <article-title>Overview of the TREC 2007 Enterprise Track</article-title>
          .
          <source>In Proceedings of TREC</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Bailey</surname>
          </string-name>
          , Nick Craswell, Ian Soboro, Paul omas, Arjen P. de Vries, and
          <string-name>
            <given-names>Emine</given-names>
            <surname>Yilmaz</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Relevance Assessment: Are Judges Exchangeable</article-title>
          and Does It Maer?.
          <source>In Proceedings of ACM SIGIR</source>
          <year>2008</year>
          .
          <volume>667</volume>
          -
          <fpage>674</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Ben</given-names>
            <surname>Cartere</surname>
          </string-name>
          e.
          <year>2012</year>
          .
          <article-title>Multiple testing in statistical analysis of systems-based information retrieval experiments</article-title>
          .
          <source>ACM TOIS 30</source>
          ,
          <issue>1</issue>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Ben</given-names>
            <surname>Cartere</surname>
          </string-name>
          e and Ian Soboro.
          <year>2010</year>
          . e E
          <article-title>ect of Assessor Errors on IR System Evaluation</article-title>
          .
          <source>In Procceedings of ACM SIGIR</source>
          <year>2010</year>
          .
          <volume>539</volume>
          -
          <fpage>546</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] omas Demeester, Robin Aly, Djoerd Hiemstra, Dong Nguyen, Dolf Trieschnigg, and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Develder</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Exploiting User Disagreement for Web Search Evaluation: an Experimental Approach</article-title>
          .
          <source>In Proceedings of ACM WSDM</source>
          <year>2014</year>
          .
          <volume>33</volume>
          -
          <fpage>42</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Marco</given-names>
            <surname>Ferrante</surname>
          </string-name>
          , Nicola Ferro, and
          <string-name>
            <given-names>Maria</given-names>
            <surname>Maistro</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>AWARE: Exploiting Evaluation Measures to Combine Multiple Assessors</article-title>
          .
          <source>ACM TOIS 36</source>
          ,
          <issue>2</issue>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Joseph</surname>
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Fleiss</surname>
          </string-name>
          .
          <year>1971</year>
          .
          <article-title>Measuring Nominal Scale Agreement among Many Raters</article-title>
          .
          <source>Psychological Bulletin</source>
          <volume>76</volume>
          ,
          <issue>5</issue>
          (
          <year>1971</year>
          ),
          <fpage>378</fpage>
          -
          <lpage>382</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Donna</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Harman</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>e TREC Ad Hoc Experiments</article-title>
          .
          <source>In TREC: Experiment and Evaluation in Information Retrieval</source>
          , Ellen M.
          <article-title>Voorhees and Donna</article-title>
          K. Harman (Eds.). e MIT Press, Chapter 4.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Klaus</surname>
            <given-names>Krippendor.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Content Analysis: An Introduction to Its Methodology (ird Edition)</article-title>
          .
          <source>Sage Publications.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Cheng</surname>
            <given-names>Luo</given-names>
          </string-name>
          , Tetsuya Sakai, Yiqun Liu, Zhicheng Dou, Chenyan Xiong, and
          <string-name>
            <given-names>Jingfang</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Overview of the NTCIR-13 WWW Task</article-title>
          .
          <source>In Proceedings of NTCIR-13.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Eddy</surname>
            <given-names>Maddalena</given-names>
          </string-name>
          , Kevin Roitero, Gianluca Demartini, and
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Mizzaro</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Considering Assessor Agreement in IR Evaluation</article-title>
          .
          <source>In Proceedings of ACM ICTIR</source>
          <year>2017</year>
          .
          <volume>75</volume>
          -
          <fpage>82</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Olga</surname>
            <given-names>Megorskaya</given-names>
          </string-name>
          , Vladimir Kukushkin, and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Serdyukov</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>On the Relation between Asessor's Agreement and Accuracy in Gamied Relevance Assessment</article-title>
          .
          <source>In Proceedings of ACM SIGIR</source>
          <year>2015</year>
          .
          <volume>605</volume>
          -
          <fpage>614</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Justus</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Randolph</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Free-Marginal Multirater Kappa (Multirater κfree): An Alternative to Fleiss' Fixed Marginal Multirater Kappa</article-title>
          .
          <source>In Joensuu Learning and Instruction Symposium</source>
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Tetsuya</given-names>
            <surname>Sakai</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Evaluating Evaluation Metrics based on the Bootstrap</article-title>
          .
          <source>In Proceedings of ACM SIGIR</source>
          <year>2006</year>
          .
          <volume>525</volume>
          -
          <fpage>532</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Tetsuya</given-names>
            <surname>Sakai</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Alternatives to Bpref</article-title>
          .
          <source>In Proceedings of ACM SIGIR</source>
          <year>2007</year>
          .
          <volume>71</volume>
          -
          <fpage>78</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Tetsuya</given-names>
            <surname>Sakai</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Evaluation with Informational and Navigational Intents</article-title>
          .
          <source>In Proceedings of WWW</source>
          <year>2012</year>
          .
          <volume>499</volume>
          -
          <fpage>508</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Tetsuya</given-names>
            <surname>Sakai</surname>
          </string-name>
          .
          <year>2014</year>
          . Metrics, Statistics, Tests.
          <source>In PROMISE Winter School</source>
          <year>2013</year>
          :
          <article-title>Bridging between Information Retrieval</article-title>
          and
          <source>Databases (LNCS 8173)</source>
          .
          <fpage>116</fpage>
          -
          <lpage>163</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Tetsuya</given-names>
            <surname>Sakai</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Unanimity-Aware Gain for Highly Subjective Assessments</article-title>
          .
          <source>In Proceedings of EVIA</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Ellen</given-names>
            <surname>Voorhees</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Variations in Relevance Judgments and the Measurement of Retrieval Eectiveness</article-title>
          .
          <source>Information Processing and Management</source>
          (
          <year>2000</year>
          ),
          <fpage>697</fpage>
          -
          <lpage>716</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Yulu</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Garrick Sherman,
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Lin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Miles</given-names>
            <surname>Efron</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Assessor Dierences and User Preferences in Tweet Timeline Generation</article-title>
          .
          <source>In Proceedings of ACM SIGIR</source>
          <year>2015</year>
          .
          <volume>615</volume>
          -
          <fpage>624</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>William</given-names>
            <surname>Webber</surname>
          </string-name>
          , Praveen Chandar, and Ben Carteree.
          <year>2012</year>
          .
          <article-title>Alternative Assessor Disagreement and Retrieval Depth</article-title>
          .
          <source>In Proceedings of ACM CIKM</source>
          <year>2012</year>
          .
          <volume>125</volume>
          -
          <fpage>134</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>