<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CENTRE@CLEF2019: Overview of the Replicability and Reproducibility Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ERR@10</email>
          <email>ERR@100</email>
          <email>ERR@1000</email>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Norbert Fuhr</string-name>
          <email>norbert.fuhr@uni-due.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Maistro</string-name>
          <email>maistro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetsuya Sakai</string-name>
          <email>tetsuyasakai@acm.org</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ian Soboro</string-name>
          <email>ian.soboroff@nist.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of Standards and Technology (NIST)</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Copenhagen</institution>
          ,
          <country country="DK">Denmark</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Duisburg-Essen</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Waseda University</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Reproducibility has become increasingly important for many research areas, among those IR is not an exception and has started to be concerned with reproducibility and its impact on research results. This paper describes our second attempt to propose a lab on reproducibility named CENTRE, held during CLEF 2019. The aim of CENTRE is to run both a replicability and reproducibility challenge across all the major IR evaluation campaigns and to provide the IR community with a venue where previous research results can be explored and discussed. This paper reports the participant results and preliminary considerations on the second edition of CENTRE@CLEF 2019.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Reproducibility is becoming a primary concern in many areas of science [
        <xref ref-type="bibr" rid="ref16">16, 24</xref>
        ]
as well as in computer science, as also witnessed by the recent ACM policy on
result and artefact review and badging.
      </p>
      <p>
        Also in Information Retrieval (IR) replicability and reproducibility of the
experimental results are becoming a more and more central discussion items in the
research community [
        <xref ref-type="bibr" rid="ref12 ref4">4, 12, 17, 23, 28</xref>
        ]. We now commonly nd questions about
the extent of reproducibility of the reported experiments in the review forms of
all the major IR conferences, such as SIGIR, CHIIR, ICTIR and ECIR, as well
as journals, such as ACM TOIS. We also witness to the raise of new activities
aimed at verifying the reproducibility of the results: for example, the
\Reproducibility Track" at ECIR since 2015 hosts papers which replicate, reproduce
and/or generalize previous research results.
      </p>
      <p>Nevertheless, it has been repeatedly shown that best TREC systems still
outperform o -the-shelf open source systems [4{6, 22, 23]. This is due to many
di erent factors, among which lack of tuning on a speci c collection when using
default con guration, but it is also caused by the lack of the speci c and advanced
components and resources adopted by the best systems.</p>
      <p>
        It has been also shown that additivity is an issue, since adding a component
on top of a weak or strong base does not produce the same level of gain [
        <xref ref-type="bibr" rid="ref6">6, 22</xref>
        ].
This poses a serious challenge when o -the-shelf open source systems are used
as stepping stone to test a new component on top of them, because the gain
might appear bigger starting from a weak baseline.
      </p>
      <p>
        Moreover, besides the problems encountered in replicating/reproducing
research, we lack any well established measure to assess and quantify the extent
to which something has been replicated/reproduced. In other terms, even if a
later researcher can manage to replicate or reproduce an experiment, to which
extent can we claim that the experiment is successfully replicated or reproduced?
For the replicability task we can compare the original measure score with the
score of the replicated run, as done in [
        <xref ref-type="bibr" rid="ref14 ref15">15, 14</xref>
        ]. However, this can not be done for
reproducibility, since the reproduced system is obtained on a di erent data set
and it is not directly comparable with the original system in terms of measure
scores.
      </p>
      <p>
        Finally, both a Dagstuhl Perspectives Workshop [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and the recent SWIRL
III strategic workshop [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] have put on the IR research agenda the need to develop
both better explanatory models of IR system performance and new predictive
models, able to anticipate the performance of IR systems in new operational
conditions.
      </p>
      <p>Overall, the above considerations stress the need and urgency for a systematic
approach to reproducibility and generalizability in IR. Therefore, the goal of
CLEF, NTCIR, TREC REproducibility (CENTRE) at CLEF 2019 is to run a
joint CLEF/NTCIR/TREC task on challenging participants:
{ to replicate and reproduce best results of best/most interesting systems in
previous editions of CLEF/NTCIR/TREC by using standard open source
IR systems;
{ to contribute back to the community the additional components and
resources developed to reproduce the results in order to improve existing open
source systems;
{ to start exploring the generalizability of our ndings and the possibility of
predicting IR system performance;
{ to investigate possible measures for replicability and reproducibility in IR.</p>
      <p>The paper is organized as follows: Section 2 introduces the setup of the
lab; Section 3 discusses the participation and the experimental outcomes; and,
Section 4 draws some conclusions and outlooks possible future works.</p>
    </sec>
    <sec id="sec-2">
      <title>Evaluation Lab Setup</title>
      <sec id="sec-2-1">
        <title>Tasks</title>
        <p>Similarly to its previous edition, CENTRE@CLEF 2019 o ered the following
two tasks:
{ Task 1 - Replicability : the task focuses on the replicability of selected
methods on the same experimental collections;
{ Task 2 - Reproducibility : the task focuses on the reproducibility of selected
methods on di erent experimental collections;</p>
        <p>For Replicability and Reproducibility we refer to the ACM Artifact Review
and Badging de nitions6:
{ Replicability (di erent team, same experimental setup): the measurement
can be obtained with stated precision by a di erent team using the same
measurement procedure, the same measuring system, under the same
operating conditions, in the same or a di erent location on multiple trials. For
computational experiments, this means that an independent group can
obtain the same result using the author's own artifacts. In CENTRE@CLEF
2019 this meant to use the same collections, topics and ground-truth on
which the methods and solutions have been developed and evaluated.
{ Reproducibility (di erent team, di erent experimental setup): The
measurement can be obtained with stated precision by a di erent team, a di erent
measuring system, in a di erent location on multiple trials. For
computational experiments, this means that an independent group can obtain the
same result using artifacts which they develop completely independently. In
CENTRE@CLEF 2019 this meant to use a di erent experimental collection,
but in the same domain, from those used to originally develop and evaluate
a solution.</p>
        <p>
          For Task 1 and Task 2, CENTRE@CLEF 2019 teams up with the
OpenSource IR Replicability Challenge (OSIRRC) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] at SIGIR 2019. Therefore,
participating groups could consider to submit their runs both to CENTRE@CLEF
2019 and OSIRRC 2019, where the second venue requires to submit the runs as
Docker images.
        </p>
        <p>Besides Task 1 and Task 2, CENTRE@CLEF 2019 o ered also a new pilot
task:
{ Task 3 - Generalizability : the task focuses on collection performance
prediction and the goal is to rank (sub-)collections on the basis of the expected
performance over them.</p>
        <p>In details, Task 3 was instatiated as follows:
6 https://www.acm.org/publications/policies/
artifact-review-badging
{ Training : participants need to run plain BM25 and, if they wish, also their
own system on the test collection used for TREC 2004 Robust Track (they
are allowed to use the corpus, topics and qrels). Participants need to identify
features of the corpus and topics that allow them to predict the system score
with respect to Average Precision (AP).
{ Validation: participants can use the test collection used for TREC 2017
Common Core Track (corpus, topics and qrels) to validate their method and
determine which set of features represent the best choice for predicting AP
score for each system. Note that the TREC 2017 Common Core Track topics
are an updated version of the TREC 2004 Robust track topics.
{ Test (submission): participants need to use the test collection used for TREC
2018 Common Core Track (only corpus and topics). Note that the TREC
2018 Common Core Track topics are a mix of \old" and \new" topics, where
old topics were used in TREC 2017 Common Core track. Participants will
submit a run for each system (BM25 and their own system) and an additional
le (one for each system) including the AP score predicted for each topic.
The score predicted can be a single value or a value with the corresponding
con dence interval.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Replicability and Reproducibility Targets</title>
        <p>
          For the previous edition of CENTRE@CLEF 2018 [
          <xref ref-type="bibr" rid="ref14 ref15">15, 14</xref>
          ] we selected the target
runs for replicability and reproducibility among the Ad Hoc tasks in previous
editions of CLEF, TREC, and NTCIR. However, even though CENTRE@CLEF
2018 had 17 enrolled teams, eventually only one team managed to submit a run.
One of the main issues reported by the participating team is the lack of the
external resources exploited in the original paper, which are no longer available [19].
Therefore, for CENTRE@CLEF 2019 we decided to focus on more recent papers
submitted at TREC Common Core Track in 2017 and 2018.
        </p>
        <p>
          To select the target runs from the TREC 2017 and 2018 Common Core Tracks
we did not consider the impact of the proposed approaches in terms of number
of citations, since both the tracks are recent and the citations received by the
submitted papers are not signi cant. Therefore, we looked at the nal ranking
of runs reported in the tracks overviews [
          <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
          ] and we chose the best performing
runs which exploit open source search systems and do not make use of additional
relevance assessments, which are not available to di erent teams.
        </p>
        <p>Below we list the runs selected as targets of replicability and reproducibility
among which the participants can choose. For each run, we specify the
corresponding collection for replicability and for reproducibility. For more
information, the list also provides references to the papers describing those runs as well
as the overviews describing the overall task and collections.</p>
        <p>{ Runs: WCrobust04 and WCrobust0405 [18]</p>
        <p>
          Task Type: TREC 2017 Common Core Track [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
Replicability: New York Times Annotated Corpus, with TREC 2017
Common Core Topics
{ Runs: RMITFDA4 and RMITEXTGIGADA5 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
        </p>
        <p>
          Reproducibility: TREC Washington Post Corpus, with TREC 2018
Common Core Topics
Task Type: TREC 2018 Common Core Track [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]
Replicability: TREC Washington Post Corpus, with TREC 2018
Common Core Topics
Reproducibility: New York Times Annotated Corpus, with TREC
2017 Common Core Topics
        </p>
        <p>Since these runs were not originally thought for being used as targets of a
replicability/reproducibility exercise, we contacted the authors of the papers to
inform them and ask their consent to use the runs.</p>
        <p>The participants in CENTRE@CLEF 2019 were not provided with the
corpora necessary to perform the tasks. The following collections were needed to
perform the task:
{ The New York Times Annotated Corpus7 contains over 1:8 million articles
written and published by the New York Times between January 1, 1987 and
June 19, 2007. The text in this corpus is formatted in News Industry Text
Format (NITF), which is an XML speci cation that provides a standardized
representation for the content and structure of discrete news articles. The
dataset is available upon payment of a fee.
{ The TREC Washington Post Corpus8 contains 608 180 news articles and
blog posts from January 2012 through August 2017. The articles are stored
in JSON format, and include title, byline, date of publication, kicker (a
section header), article text broken into paragraphs, and links to embedded
images and multimedia. The dataset is publicly available and free of charge.
{ The TREC 2004 Robust Corpus9 corresponds to the set of documents on
TREC disks 4 and 5, minus the Congressional Record. This document set
contains approximately 528 000 documents. The dataset is available upon
payment of a fee.</p>
        <p>Finally, Table 1 reports the topics used for the three tasks, with the
corresponding number of documents and pool sizes. An example of topic is reported
in the Figure 1 for TREC 2018 Common Core Track.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Evaluation Measures</title>
        <p>
          Task 1 - Replicability: As done in the previous edition of CENTRE [
          <xref ref-type="bibr" rid="ref14 ref15">15, 14</xref>
          ], the
quality of the replicability runs has been evaluated from two points of view:
7 https://catalog.ldc.upenn.edu/LDC2008T19
8 https://trec.nist.gov/data/wapost/
9 https://trec.nist.gov/data/qa/T8_QAdata/disks4_5.html
where T is the total number of topics, P is the total number of concordant
pairs (document pairs that are ranked in the same order in both vectors)
Q the total number of discordant pairs (document pairs that are ranked
(2)
in opposite order in the two vectors), U and V are the number of ties,
respectively, in the rst and in the second ranking.
        </p>
        <p>Note that the de nition of Kendall's in Equation (2) is originally proposed
for permutations of the same set of items, therefore it is not applicable whenever
two rankings do not contain the same set of documents. However, for real
rankings of systems it is highly likely that two lists do not contain the same set of
items, thus we performed some pre-processing with the runs before computing
Kendall's in Equation (2).</p>
        <p>In details, consider a xed topic t, the original ranking rt;orig and the
replicated ranking rt;replica. If one of the rankings contains a document that is not
retrieved by the other ranking, we de ne the rank position of that document
as zero. For example, if for a document d, d 2 rt;orig, but d 62 rt;replica, then
the rank position of d in rt;replica is zero. Whenever the two rankings contains
the same set of documents, Equation (2) is not a ected by this pre-processing
step and the computation of Kendall's tau is performed as usual. Furthermore,
if two rankings retrieves di erent documents and place them in the same rank
positions, Kendall's tau will still be equal to 1, and the comparison is performed
just with respect to the relative order of the documents retrieved by both the
rankings.</p>
        <p>Task 2 - Reproducibility: Since for the reproducibility runs we do not have an
already existing run to compare against, we compare the reproduced run score
with respect to a baseline run, to see whether the improvement over the baseline
is comparable between the original collection C and the new collection D. In
particular we compute the E ect Ratio (ER), which is also exploited in
CENTRE@NTCIR 14 [25].</p>
        <p>In details, given two runs, we refer to the A-run, as the advanced run, and
B-run, as the baseline run, where the A-run has been reported to outperform the
B-run on the original test collection C. The intuition behind ER is to evaluate
to which extent the improvement on the original collection C is reproduced on a
new collection D. For any evaluation measure M , let MiC (A) and MiC (B) denote
the score of the A-run and that of the B-run for the i-th topic of collection C
(1 i TC ). Similarly, let MiD(A0) and MiD(B0) denote the scores for the
reproduced A-run and B-run respectively, on the new collection D. Then, ER is
computed as follows:</p>
        <p>ER(</p>
        <sec id="sec-2-3-1">
          <title>MrDeproduced;</title>
          <p>1 PTD
MoCrig) = TD 1 iP=1TC
TC i=1</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>MiD;reproduced</title>
        </sec>
        <sec id="sec-2-3-3">
          <title>MiC;orig</title>
          <p>(3)
where MiC;orig = MiC (A) MiC (B) is the per-topic improvement of the original
advanced and baseline runs for the i-th topic on C. Similarly MiD;reproduced =
MiC (A0) MiC (B0) is the per-topic improvement of the reproduced advanced
and baseline runs for the i-th topic on D. Note that the per-topic improvement
can be negative, for those topics where the advanced run fails to outperform the
baseline run.</p>
          <p>If ER 0, that means that the replicated A-run failed to outperform the
replicated B-run: the replication is a complete failure. If 0 &lt; ER &lt; 1, the
replication is somewhat successful, but the e ect is smaller compared to the
original experiment. If ER = 1, the replication is perfect in the sense that the
original e ect has been recovered as is. If ER &gt; 1, the replication is successful,
and the e ect is actually larger compared to the original experiment.</p>
          <p>Finally, ER in Equation (3) is instantiated with respect to AP, nDCG and
ERR. Furthermore, as suggested in [25], ER is computed even for the replicability
task, by replacing MiD;reproduced with MiC;replica in Equation (3).
Task 3 - Generalizability: For the generalizability task we planned to compare
the predicted run score with the original run score. This is measured with Mean
Absolute Error and RMSE between the predicted and original measures scores,
with respect to AP, nDCG and ERR. However, we did not receive any run for
the generalizability task, so we did not put in practice this part of the evaluation
task.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Participation and Outcomes</title>
      <p>19 groups registered for participating in CENTRE@CLEF2019, but
unfortunately only one group succeeded in submitting two replicability runs and two
reproducibility runs. No runs were submitted for the generalizability task.</p>
      <p>
        The team from the University of Applied Science TH Koln [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] replicated
and reproduced the runs by Grossman and Cormack [18], i.e. WCrobust04 and
WCrobust0405. They could not replicate the runs by Benham et Al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] since
they do not have access to the Gigaworld dataset10, which is publicly available
upon payment of a fee. The dataset is necessary to perform the external query
expansion exploited by the selcted runs from [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>Eventually, the participating team submitted four o cial runs and four
uno cial runs described in Table 2. The runs and all the code is publicly available
online11.
10 https://catalog.ldc.upenn.edu/LDC2012T21
11 https://bitbucket.org/centre_eval/c2019_irc/src/master/</p>
      <p>The paper by Grossman and Cormack [18] exploits the principle of automatic
routing runs: rst, a logistic regression model is trained with the relevance
judgments from one or more collections for each topic, then the model is used to
predict relevance assessments of documents from a di erent collection. Both the
training and the prediction phases are done on a topic-wise basis.</p>
      <p>The routing process represented a challenge for the participating team, which
initially submitted a set of four o cial runs, where some of the topics were
missing. For example, the o cial run irc task1 WCrobust0405 001 contains
only 33 topics, while the corresponding original run WCrobust0405 contains
all the 50 topics. The participating team could not understand how to derive
document rankings for those topics such that no training topics were available
for the logistic regression model. For example, when they were attempting to
replicate WCrobust0405, they exploited as training set the intersection
between the topics from TREC 2004 Robust and TREC 2005 Robust. Then, for
the prediction phase, only 33 topics from TREC 2017 Common Core were
contained in the training set, and no prediction could be performed for the
remaining topics. Due to similar issues, the o cial irc task2 WCrobust04 001 and
irc task2 WCrobust0405 001 contain 25 and 15 topics respectively.</p>
      <p>Afterwards, the participating team contacted the authors of the original
paper, Grossman and Cormack [18], to understand how to derive rankings even
when there are no training topics available. The authors clari ed that for
WCrobust0405 the training set contains both the topics from TREC 2004
Robust and TREC 2005 Robust, and when a topic is not contained in TREC
2005 Robust, they used just the TREC 2004 Robust collection as training set.
Therefore, the authors submitted four additional uno cial runs, where both
irc task1 WCrobust04 001 and irc task1 WCrobust0405 001 contain
all the 50 topics, while the reproduced runs irc task2 WCrobust04 001 and
irc task2 WCrobust04 001 contain 25 topics. Note that some of the topics
are missing for the reproduced runs, since no training data is available for 25
out of the 50 topics of TREC 2018 Common Core.</p>
      <p>In the following we report the evaluation results for the replicability and
reproducibility tasks, both for the o cial and uno cial submissions.</p>
      <p>Table 3 and Table 4 report AP, nDCG and ERR scores for the o cial
replicated runs. As shown by RMSE, the replication task was fairly successful with
respect to AP and nDCG, while when ERR is considered, RMSE is greater than
0:2, showing that it is harder to replicate ERR than the other evaluation
measures. Indeed, it is well known that ERR is highly sensitive to the position of
relevant documents at the very beginning of the ranking, thus even the
misplacement of a single relevant documents may cause a signi cant drop in ERR
score.</p>
      <p>Furthermore, as the cut-o increases, even RMSE for AP and nDCG
increases, showing that the replication is less accurate at lower cut-o levels. On
the other side, RMSE for ERR is almost constant when the cut-o increases,
showing once more that ERR focuses on the top rank positions rather than
considering the whole ranking.
0:0506
0:2252
0:3821
0:1442
0:3883
0:6299
0:0473
0:2541
0:4428
0:1490
0:4268
0:6883
nDCG@1000
ERR@100</p>
      <p>Similarly, Table 5 reports AP, nDCG and ERR scores for the uno cial
replicated run irc task1 WCrobust0405 001. Note that the o cial and uno cial
replicated run irc task1 WCrobust04 001 are identical, therefore the
evaluation scores for this uno cal run are the same reported in Table 3 and are
omitted in the following.</p>
      <p>Again, we can observe that the replication task is more successful for RMSE
with respect to AP and nDCG than ERR. Furthermore, RMSE increases as the
cut-o increases, meaning that the accuracy of the replicated run decreases as
the cut-o level increases.</p>
      <p>By comparing the o cial and uno cial evaluation results for irc
task1WCrobust0405 001, in Table 4 and Table 5 respectively, we can note that
RMSE score are quite similar, showing that the uno cial run is fairly accurate
even on the additional topics.</p>
      <p>Table 6 reports ER for the replication task with the o cial runs. We
considered WCrobust0405 as advanced run and WCrobust04 as baseline run,
therefore the per-topic improvement is computed as WCrobust0405 scores
minus WCrobust04 scores for each topic. For the replicated o cial runs, we
needed to select from irc task1 WCrobust04 001 the 33 topics contained in
irc task1 WCrobust0405 001, otherwise we could not compute the mean
per-topic improvement.</p>
      <p>ER shows that the replication task is fairly successful for AP, while it is
less successful for nDCG and ERR. Furthermore, ER &gt; 1 highlights that the
di erence between the advanced and the baseline run is more pronounced in the
replicated runs than in the original runs. Again, it can be noted that as the
cut-o increases, the accuracy of the replicability exercise decreases for AP and
nDCG, while it is almost constant for ERR.
0:0078
0:0446
0:0556
0:0233
0:0597
0:0578
0:0078
0:0446
0:0556
0:0233
0:0597
0:0578</p>
      <p>between the original and replicated runs.</p>
      <p>Replicated Run
irc task1 WCrobust04 001 o cial
irc task1 WCrobust0405 001 o cial</p>
      <p>Original Run
WCrobust04</p>
      <p>WCrobust0405
irc task1 WCrobust0405 001 uno cial</p>
      <p>WCrobust0405
0:0222
0:0034
0:0107
0:0073
0:0316
0:0199
0:0021
0:0046</p>
      <p>Fig. 3. First 10 rank positions for
irc task1 WCrobust04 001 for
topic 307 form TREC 2017 Common</p>
      <p>Core Track.</p>
      <p>Analogously, Table 7 reports ER for the replication task with the uno
cial runs. We considered WCrobust0405 as advanced run and WCrobust04
as baseline run, therefore the per-topic improvement is computed as in
Table 6 and the rst column is equal. Both the replicated uno cial runs
contain the same 50 topics, therefore the per-topic improvement is computed as
irc task1 WCrobust0405 001 scores minus irc task1 WCrobust04 001
scores for each topic.</p>
      <p>When the whole set of 50 topics is considered, the replication is fairly
successful with respect to all the measure, with ER ranging between 0:83 and 1:12.
The only exception is represented by AP@10, where the replicated runs fails to
replicate the per-topic improvements. Again, the accuracy of the replicated runs
decreases as the cut-o increases.</p>
      <p>Table 8 reports the Kendall's correlation between the original and
replicated runs, both for the o cial and uno cial runs. We computed Kendall's at
di erent cut-o levels, where we rst trimmed the runs at the speci ed cut-o
and subsequently computed Kendall's between the trimmed runs.</p>
      <p>Table 8 shows that the replication was not successful for any of the runs in
terms of Kendall's . This means that even if the considered replicated runs were
similar to the original runs in terms of placement of relevant and non relevant
documents, they actually retrieves di erent documents.</p>
      <p>Figure 2 and Figure 3 shows the rst 10 rank positions for WCrobust04 and
its replicated version irc task1 WCrobust04 001, for topic 307 from TREC
2017 Common Core Track. We can observe that even if the runs retrieves a
similar set of documents, the relative position of each document is di erent. For
example, document 309412 is at rank position 1 for the original run, but at
nDCG@1000
0:0078
0:0446
0:0556
0:0233
0:0597
0:0578
0:1042
0:1019
0:1019
0:0122
0:0431
0:0579
0:0298
0:0767
0:0898
0:0124
0:0142
0:0135</p>
      <p>ER
1:5641
0:9664
1:0414
1:2790
1:2848
1:5536
0:1190
0:1394
0:1325
rank position 2 for the replicated run, similary document 733642 is at rank
position 1 for the replicated run and at rank position 5 for the original run.
Moreover, document 241240 is at rank position 3 for the replicated run, but it
does not apper on the rst 10 positions for the original run.</p>
      <p>Table 8, Figure 2 and Figure 3 highlights how hard is to replicate the exact
ranking of documents. Therefore, whenever a replicability task is considered,
comparing the evaluation scores with RMSE or ER might not be enough, since
these approaches consider just the position of relevant and not relevant
documents, and overlook the actual ranking of documents.</p>
      <p>Finally, Table 9 reports the mean per-topic improvement and ER for the
o cial runs from the reproducibility task. As done for the replicability task, we
considered WCrobust0405 as advanced run and WCrobust04 as baseline run
on the test collection from TREC 2017 Common Core Track. For the reproduced
o cial runs, we needed to select from irc task2 WCrobust04 001 the 15
topics contained in irc task2 WCrobust0405 001 from TREC 2018 Common
Core Track, otherwise we could not compute the per topic improvement.</p>
      <p>Furthermore, when the cut-o increases, the accuracy of the
reproducibility exercise increases for AP, while it decreases for nDCG and remains almost
constant for ERR.
nDCG@1000
0:0078
0:0446
0:0556
0:0233
0:0597
0:0578
0:1042
0:1019
0:1019
0:0065
0:0241
0:0336
0:0155
0:0426
0:0509
0:0004
0:0033</p>
      <p>Similarly, Table 10 reports the mean per-topic improvement and ER for the
uno cial runs from the reproducibility task. We considered WCrobust0405 as
advanced run and WCrobust04 as baseline run on TREC 2017 Common Core,
therefore the per-topic improvement is computed as in Table 9 and the rst
column is equal. Both the reproduced uno cial runs contain the same 25 topics,
therefore the per-topic improvement is computed as irc task2
WCrobust0405 001 scores minus irc task2 WCrobust04 001 scores for each topic.</p>
      <p>The best reproducibility results are obtained with respect to AP@10 and
nDCG@1000, thus the e ect of the advanced run over the baseline run is
better reproduced at the beginning of the ranking for AP, and when the whole
ranked list is considered, for nDCG. Again, ERR is the hardest measure to be
reproduced, indeed it has the lowest ER score for each cut-o level.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>
        This paper reports the results on the second edition of CENTRE@CLEF2019.
A total of 19 participants enrolled in the lab, however just one group managed
to submit two replicability runs and two reproducibility runs. As reported in
Section 3, the participating team could not reproduce the runs from Benham et
Al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], due to the lack of the Gigaworld dataset, but they managed to replicate
and reproduce the runs from Grossman and Cormack [18]. More details regarding
the implementation are described in their paper [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The experimental results show that the replicated runs are fairly successful
with respect to AP and nDCG, while the lowest replicability results are obtained
with respect to ERR. As ERR mainly focuses on the beginning of the ranking,
misplacing even a single relevant document can deteriorate ERR score and have
a great impact on the replicability evaluation scores.</p>
      <p>Moreover, whenever replicability is considered, RMSE and ER are not enough
to evaluate the replicated runs. Indeed, they only account for the position of
relevant and not relevant documents by considering the similarity between the
original scores and the replicated scores, and they overlook the actual ranking
of documents. When the runs are evaluated with Kendall's to account for
the actual position of the documents in the ranking, the experiments show that
the replicability is not successful at all, with Kendall's values close to 0. This
con rms that, even if it is possible to achieve similar scores in terms of IR
evaluation measures, it is challenging to replicate the same documents ranking.</p>
      <p>When it comes to reproducibility, there are no well-established evaluation
measures to determine to which extent a system can be reproduced. Therefore,
we compute ER, rstly exploited in [25], which focuses on the reproduction of
the improvement of an advanced run over a baseline run. The experiments show
that reproducibility was fairly successful in terms of AP@10 and nDCG@1000,
while, similarly to the replicability task, ERR is the hardest measure in terms
of reproducibility success.</p>
      <p>
        Finally, as reported in [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ], the lack of participation is a signal that the
IR community is somehow overlooking replicability and reproducibility issues.
As it also emerged from a recent survey within the SIGIR community [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], while
there is a very positive attitude towards reproducibility and it is considered very
important from a scienti c point of view, there are many obstacles to it such as
the e ort required to put it into practice, the lack of rewards for achieving it, the
possible barriers for new and inexperienced groups, and, last but not least, the
(somehow optimistic) researcher's perception that their own research is already
reproducible.
      </p>
      <p>For the next edition of the lab we are planning to propose some changes in
the lab organization to increase the interest and participation of the research
community. First, we will target for more popular systems to be replicated and
reproduced, moreover we will consider other tasks than the AdHoc, as for
example the medical or other popular domains.
17. Fuhr, N.: Some Common Mistakes In IR Evaluation, And How They Can Be</p>
      <p>
        Avoided. SIGIR Forum 51(3), 32{41 (December 2017)
18. Grossman, M.R., Cormack, G.V.: MRG UWaterloo and WaterlooCormack
Participation in the TREC 2017 Common Core Track. In: Voorhees and Ellis [26]
19. Jungwirth, M., Hanbury, A.: Replicating an Experiment in Cross-lingual
Information Retrieval with Explicit Semantic Analysis. In: Cappellato et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
20. Kendall, M.G.: Rank correlation methods. Gri n, Oxford, England (1948)
21. Kenney, J.F., Keeping, E.S.: Mathematics of Statistics { Part One. D. Van
Nostrand Company, Princeton, USA, 3rd edn. (1954)
22. Kharazmi, S., Scholer, F., Vallet, D., Sanderson, M.: Examining Additivity and
Weak Baselines. ACM Transactions on Information Systems (TOIS) 34(4), 23:1{
23:18 (June 2016)
23. Lin, J., Crane, M., Trotman, A., Callan, J., Chattopadhyaya, I., Foley, J., Ingersoll,
G., Macdonald, C., Vigna, S.: Toward Reproducible Baselines: The Open-Source
IR Reproducibility Challenge. In: Ferro, N., Crestani, F., Moens, M.F., Mothe, J.,
Silvestri, F., Di Nunzio, G.M., Hau , C., Silvello, G. (eds.) Advances in
Information Retrieval. Proc. 38th European Conference on IR Research (ECIR 2016). pp.
357{368. Lecture Notes in Computer Science (LNCS) 9626, Springer, Heidelberg,
Germany (2016)
24. Munafo, M.R., Nosek, B.A., Bishop, D.V.M., Button, K.S., Chambers, C.D.,
Percie du Sert, N., Simonsohn, U., Wagenmakers, E.J., Ware, J.J., Ioannidis, J.P.A.:
A manifesto for reproducible science. Nature Human Behaviour 1, 0021:1{0021:9
(January 2017)
25. Sakai, T., Ferro, N., Soboro , I., Zeng, Z., Xiao, P., Maistro, M.: Overview of the
NTCIR-14 CENTRE Task. In: Ishita, E., Kando, N., Kato, M.P., Liu, Y. (eds.)
Proc. 14th NTCIR Conference on Evaluation of Information Access Technologies.
pp. 494{509. National Institute of Informatics, Tokyo, Japan (2019)
26. Voorhees, E.M., Ellis, A. (eds.): The Twenty-Sixth Text REtrieval Conference
Proceedings (TREC 2017). National Institute of Standards and Technology (NIST),
Special Publication 500-324, Washington, USA (2018)
27. Voorhees, E.M., Ellis, A. (eds.): The Twenty-Seventh Text REtrieval
Conference Proceedings (TREC 2018). National Institute of Standards and Technology
(NIST), Washington, USA (2019)
28. Zobel, J., Webber, W., Sanderson, M., Mo at, A.: Principles for Robust Evaluation
Infrastructure. In: Agosti, M., Ferro, N., Thanos, C. (eds.) Proc. Workshop on Data
infrastructurEs for Supporting Information Retrieval Evaluation (DESIRE 2011).
pp. 3{6. ACM Press, New York, USA (2011)
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Allan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arguello</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azzopardi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bailey</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balog</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bast</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belkin</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berberich</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>von Billerbeck</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Capra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carterette</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clarke</surname>
            ,
            <given-names>C.L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Collins-Thompson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Craswell</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croft</surname>
            ,
            <given-names>W.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Culpepper</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalton</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demartini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diaz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dietz</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eickho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuhr</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Geva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hau</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hawking</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamps</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiseleva</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Mizzaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            , Mo at, A.,
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.Y.</given-names>
            ,
            <surname>Olteanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Ounis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Radlinski</surname>
          </string-name>
          , F.,
          <string-name>
            <surname>de Rijke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scholer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sitbon</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smucker</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soboro</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spina</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suel</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thom</surname>
            , J., Thomas,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trotman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Vries</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yilmaz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <source>Research Frontiers in Information Retrieval { Report from the Third Strategic Workshop on Information Retrieval in Lorne (SWIRL</source>
          <year>2018</year>
          ).
          <source>SIGIR Forum</source>
          <volume>52</volume>
          (
          <issue>1</issue>
          ) (
          <year>June 2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Allan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>D.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Gysel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.:</given-names>
          </string-name>
          <article-title>TREC 2017 Common Core Track Overview</article-title>
          . In: Voorhees and Ellis [
          <volume>26</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Allan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>D.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Gysel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.:</given-names>
          </string-name>
          <article-title>TREC 2018 Common Core Track Overview</article-title>
          . In: Voorhees and Ellis [
          <volume>27</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Arguello</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crane</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diaz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trotman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <source>Report on the SIGIR 2015 Workshop on Reproducibility</source>
          , Inexplicability, and
          <article-title>Generalizability of Results (RIGOR)</article-title>
          .
          <source>SIGIR Forum</source>
          <volume>49</volume>
          (
          <issue>2</issue>
          ),
          <volume>107</volume>
          {
          <issue>116</issue>
          (
          <year>December 2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Armstrong</surname>
            ,
            <given-names>T.G.</given-names>
          </string-name>
          , Mo at, A.,
          <string-name>
            <surname>Webber</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zobel</surname>
          </string-name>
          , J.:
          <source>Has Adhoc Retrieval Improved Since</source>
          <year>1994</year>
          ? In: Allan,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Aslam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.A.</given-names>
            ,
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Zobel</surname>
          </string-name>
          ,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (eds.)
          <source>Proc. 32nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR</source>
          <year>2009</year>
          ). pp.
          <volume>692</volume>
          {
          <fpage>693</fpage>
          . ACM Press, New York, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Armstrong</surname>
            ,
            <given-names>T.G.</given-names>
          </string-name>
          , Mo at, A.,
          <string-name>
            <surname>Webber</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zobel</surname>
            ,
            <given-names>J.: Improvements</given-names>
          </string-name>
          <string-name>
            <surname>That Don't Add</surname>
            <given-names>Up</given-names>
          </string-name>
          :
          <article-title>Ad-Hoc Retrieval Results Since 1998</article-title>
          . In: Cheung,
          <string-name>
            <given-names>D.W.L.</given-names>
            ,
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.Y.</given-names>
            ,
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.W.</given-names>
            ,
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.J</surname>
          </string-name>
          . (eds.)
          <source>Proc. 18th International Conference on Information and Knowledge Management (CIKM</source>
          <year>2009</year>
          ). pp.
          <volume>601</volume>
          {
          <fpage>610</fpage>
          . ACM Press, New York, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Benham</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gallagher</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mackenzie</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scholer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mo</surname>
            <given-names>at</given-names>
          </string-name>
          , A.,
          <string-name>
            <surname>Culpepper</surname>
            ,
            <given-names>J.S.:</given-names>
          </string-name>
          <article-title>RMIT at the 2018 TREC CORE Track</article-title>
          . In: Voorhees and Ellis [
          <volume>27</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Breuer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schaer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Replicability and Reproducibility of Automatic Routing Runs</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          , Muller, H. (eds.)
          <source>CLEF 2019 Working Notes. CEUR Workshop Proceedings (CEUR-WS.org)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>J.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soulier</surname>
            ,
            <given-names>L</given-names>
          </string-name>
          . (eds.):
          <source>CLEF 2018 Working Notes. CEUR Workshop Proceedings (CEUR-WS.org)</source>
          ,
          <source>ISSN 1613-0073</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Clancy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hau</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakai</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Z.Z.: The</given-names>
          </string-name>
          <string-name>
            <surname>SIGIR 2019 OpenSource IR Replicability</surname>
          </string-name>
          <article-title>Challenge (OSIRRC 2019)</article-title>
          . In: Chevalier,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gaussier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Piwowarski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Maarek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.Y.</given-names>
            ,
            <surname>Scholer</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.)
          <source>Proc. 42nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR</source>
          <year>2019</year>
          )
          <article-title>(</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuhr</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grefenstette</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konstan</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castells</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daly</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ekstrand</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Geyer</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Ku ik, T.,
          <string-name>
            <surname>Linden</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>J.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perego</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shapira</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soboro</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tintarev</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verspoor</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willemsen</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zobel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <source>The Dagstuhl Perspectives Workshop on Performance Modeling and Prediction. SIGIR Forum</source>
          <volume>52</volume>
          (
          <issue>1</issue>
          ) (
          <year>June 2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuhr</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , Jarvelin,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Kando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Lippold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Zobel</surname>
          </string-name>
          , J.:
          <article-title>Increasing Reproducibility in IR: Findings from the Dagstuhl Seminar on \Reproducibility of Data-Oriented Experiments in e-Science"</article-title>
          .
          <source>SIGIR Forum</source>
          <volume>50</volume>
          (
          <issue>1</issue>
          ),
          <volume>68</volume>
          {82 (
          <year>June 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>SIGIR Initiative to Implement ACM Artifact Review and Badging</article-title>
          .
          <source>SIGIR Forum</source>
          <volume>52</volume>
          (
          <issue>1</issue>
          ) (
          <year>June 2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maistro</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakai</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soboro</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>CENTRE@CLEF2018: Overview of the Replicability Task</article-title>
          . In: Cappellato et al. [
          <volume>9</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maistro</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakai</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soboro</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Overview of CENTRE@CLEF 2018: a First Tale in the Systematic Reproducibility Realm</article-title>
          . In: Bellot,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Trabelsi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Mothe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Murtagh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.Y.</given-names>
            ,
            <surname>Soulier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>SanJuan</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , N. (eds.)
          <article-title>Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Nineth International Conference of the CLEF Association (CLEF</source>
          <year>2018</year>
          ). pp.
          <volume>239</volume>
          {
          <fpage>246</fpage>
          . Lecture Notes in Computer Science (LNCS), Springer, Heidelberg, Germany (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Freire</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuhr</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . (eds.)
          <source>: Report from Dagstuhl Seminar</source>
          <volume>16041</volume>
          :
          <article-title>Reproducibility of Data-Oriented Experiments in e-Science</article-title>
          .
          <source>Dagstuhl Reports</source>
          , Volume
          <volume>6</volume>
          ,
          <string-name>
            <surname>Number</surname>
            <given-names>1</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schloss</surname>
            <given-names>Dagstuhl</given-names>
          </string-name>
          {
          <article-title>Leibniz-Zentrum fu</article-title>
          r Informatik,
          <source>Germany</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>