<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the 2019 Open-Source IR Replicability Challenge (OSIRRC 2019)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ryan Clancy</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudia Hauf</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jimmy Lin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetsuya Sakai</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ze Zhong Wu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Padua</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Waterloo</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Waseda University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Open-Source IR Replicability Challenge (OSIRRC 2019), organized as a workshop at SIGIR 2019, aims to improve the replicability of ad hoc retrieval experiments in information retrieval by gathering a community of researchers to jointly develop a common Docker specification and build Docker images that encapsulate a diversity of systems and retrieval models. We articulate the goals of this workshop and describe the “jig” that encodes the Docker specification. In total, 13 teams from around the world submitted 17 images, most of which were designed to produce retrieval runs for the TREC 2004 Robust Track test collection. This exercise demonstrates the feasibility of orchestrating large, community-based replication experiments with Docker technology. We envision OSIRRC becoming an ongoing community-wide efort to ensure experimental replicability and sustained progress on standard test collections.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>The importance of repeatability, replicability, and reproducibility is
broadly recognized in the computational sciences, both in
supporting desirable scientific methodology as well as sustaining empirical
progress. The Open-Source IR Replicability Challenge (OSIRRC
2019), organized as a workshop at SIGIR 2019, aims to improve the
replicability of ad hoc retrieval experiments in information retrieval
by building community consensus around a common technical
specification, with reference implementations. This overview paper
is an extended version of an abstract that appears in the SIGIR
proceedings.</p>
      <p>In order to precisely articulate the goals of this workshop, it is
ifrst necessary to establish common terminology. We use the above
terms in the same manner as recent ACM guidelines pertaining to
artifact review and badging:1
• Repeatability (same team, same experimental setup): a researcher
can reliably repeat her own computation.
• Replicability (diferent team, same experimental setup): an
independent group can obtain the same result using the authors’ own
artifacts.
• Reproducibility (diferent team, diferent experimental setup): an
independent group can obtain the same result using artifacts
which they develop completely independently.</p>
      <p>This workshop tackles the replicability challenge for ad hoc
document retrieval, with three explicit goals:
(1) Develop a common Docker specification to support images that
capture systems performing ad hoc retrieval experiments on
standard test collections. The solution that we have developed
is known as “the jig”.
(2) Build a curated library of Docker images that work with the jig
to capture a diversity of systems and retrieval models.
(3) Explore the possibility of broadening our eforts to include
additional tasks, diverse evaluation methodologies, and other
benchmarking initiatives.</p>
      <p>
        Trivially, by supporting replicability, our proposed solution enables
repeatability as well (which, as a recent case study has shown [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ],
is not as easy as one might imagine). It is not our goal to directly
address reproducibility, although we do see our eforts as an important
stepping stone.
      </p>
      <p>
        We hope that the fruits of this workshop can fuel empirical
progress in ad hoc retrieval by providing competitive baselines
that are easily replicable. The “prototypical” research paper of
this mold proposes an innovation and demonstrates its value by
comparing against one or more baselines. The often-cited
metaanalysis of Armstrong et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] from a decade ago showed that
researchers compare against weak baselines, and a recent study
by Yang et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] revealed that, a decade later, the situation has
not improved much—researchers are still comparing against weak
baselines. Lin [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] discussed social aspects of why this persists, but
there are genuine technical barriers as well. The growing
complexity of modern retrieval techniques, especially neural models that
are sensitive to hyperparameters and other details of the training
regime, poses challenges for researchers who wish to demonstrate
that their proposed innovation improves upon a particular method.
Solutions that address replicability facilitate in-depth comparisons
between existing and proposed approaches, potentially leading to
more insightful analyses and accelerating advances.
      </p>
      <p>Overall, we are pleased with progress towards the first two goals
of the workshop. A total of 17 Docker images, involving 13
diferent teams from around the world, were submitted for evaluation,
comprising the OSIRRC 2019 “image library”. These images
collectively generated 49 replicable runs for the TREC 2004 Robust
Track test collection, 12 replicable runs for the TREC 2017 Common
Core Track test collection, and 19 replicable runs for the TREC 2018
Common Core Track test collection. With respect to the third goal,
this paper ofers our future vision—but its broader adoption by the
community at large remains to be seen.
2</p>
    </sec>
    <sec id="sec-2">
      <title>BACKGROUND</title>
      <p>
        There has been much discussion about reproducibility in the
sciences, with most scientists agreeing that the situation can be
characterized as a crisis [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We lack the space to provide a comprehensive
review of relevant literature in the medical, natural, and
behavioral sciences. Within the computational sciences, to which at least
a large portion of information retrieval research belongs, there
have been many studies and proposed solutions, for example, a
recent Dagstuhl seminar [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Here, we focus on summarizing the
immediate predecessor of this workshop.
      </p>
      <p>
        Our workshop was conceived as the next iteration of the
OpenSource IR Reproducibility Challenge (OSIRRC), organized as part of
the SIGIR 2015 Workshop on Reproducibility, Inexplicability, and
Generalizability of Results (RIGOR) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This event in turn traces
its roots back to a series of workshops focused on open-source IR
systems, which is widely understood as an important component
of reproducibility. The Open-Source IR Reproducibility Challenge2
brought together developers of open-source search engines to
provide replicable baselines of their systems in a common environment
on Amazon EC2. The product is a repository that contains all code
necessary to generate ad hoc retrieval baselines, such that with a
single script, anyone with a copy of the collection can replicate
the submitted runs. Developers from seven diferent systems
contributed to the evaluation, which was conducted on the GOV2
collection. The details of their experience are captured in an ECIR
2016 paper [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>In OSIRRC 2019, we aim to address two shortcomings with the
previous exercise as a concrete step in moving the field forward.
From the technical perspective, the RIGOR 2015 participants
developed scripts in a shared VM environment, and while this was
suficient to support cross-system comparisons at the time, the
scripts were not suficiently constrained, and the entire setup
suffered from portability and isolation issues. Thus, it would have
been dificult for others to reuse the infrastructure to replicate the
results—in other words, the replicability experiments themselves
were dificult to replicate. We believe that Docker, which is a
popular standard for containerization, ofers a potential solution to these
technical challenges.</p>
      <p>Another limitation of the previous exercise was its focus on
“bag of words” baselines, and while some participants did submit
systems that exploited richer models (e.g., term dependence models
and pseudo-relevance feedback), there was insuficient diversity in
the retrieval models examined. Primarily due to these issues, the
exercise has received less follow-up and uptake than the organizers
had originally hoped.
3</p>
    </sec>
    <sec id="sec-3">
      <title>DOCKER AND “THE JIG”</title>
      <p>From a technical perspective, our eforts are built around Docker, a
widely-adopted Linux-centric technology for delivering software in
lightweight packages called containers. The Docker Engine hosts
one or more of these containers on physical machines and manages
their lifecycle. One key feature of Docker is that all containers run
on a single operating system kernel; isolation is handled by Linux
kernel features such as cgroups and kernel namespaces. This makes
containers far more lightweight than virtual machines, and hence
easier to manipulate. Containers are created from images, which
are typically built by importing base images (for example, capturing
a specific software distribution) and then overlaying custom code.
The images themselves can be manipulated, combined, and modified
as first-class citizens in a broad ecosystem. For example, a group
can overlay several existing images from public sources, add in its
own code, and in turn publish the resulting image to be further
used by others.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>General Design</title>
      <p>As defined by the Merriam-Webster dictionary, a jig is “a device
used to maintain mechanically the correct positional relationship
between a piece of work and the tool or between parts of work
during assembly”. The central activity of this workshop revolved
around the co-design and co-implementation of a jig and Docker
images that work with the jig for ad hoc retrieval. Of course, in our
context, the relationship is computational instead of mechanical.</p>
      <p>Shortly after the acceptance of the workshop proposal at SIGIR
2019, we issued a call for participants who were interested in
contributing Docker images to our efort; the jig was designed with the
input of these participants. In other words, the jig and the images
co-evolved with feedback from members of the community. The
code of the jig is open source and available on GitHub.3</p>
      <p>Our central idea is that each image would expose a number of
“hooks” that correspond to a point in the prototypical lifecycle of
an ad hoc retrieval experiment: for example, indexing a collection,
running a batch of queries, etc. These hooks then tie into code that
captures whatever retrieval model a particular researcher wishes to
package in the image—for example, a search engine implemented in
Java or C++. The jig is responsible for triggering the hooks in each
image in a particular sequence according to a predefined lifecycle
model, e.g., first index the collection, then run a batch of queries,
ifnally evaluate the results. We have further built tooling that
applies the jig to multiple images to facilitate large-scale experiments.
More details about the jig are provided in the next section, but first
we overview a few design decisions.</p>
      <p>Quite deliberately, the current jig does not make any demands
about the transparency of a particular image. For example, the
search hook can run an executable whose source code is not publicly
available. Such an image, while demonstrating replicability, would
not allow other researchers to inspect the inner workings of a
particular retrieval method. While such images are not forbidden
in our design, they are obviously less desirable than images based
on open code. In practice, however, we anticipate that most images
will be based on open-source code.</p>
      <p>One technical design choice that we have grappled with is how
to get data “into” and “out of” a container. To be more concrete,
for ad hoc retrieval the container needs access to the document
collection and the topics. The jig also needs to be able to obtain
the run files generated by the image for evaluation. Generically,
there are three options for feeding data to an image: first, the data
can be part of the image itself; second, the data can be fetched
from a remote location by the image (e.g., via curl, wget, or some
other network transfer mechanism); third, the jig could mount an
external data directory that the container has access to. The first
two approaches are problematic for our use case: images need to
be shareable, or resources need to be placed at a publicly-accessible
location online. This is not permissible for document collections
where researchers are required to sign license agreements before
using. Furthermore, both approaches do not allow the possibility
of testing on blind held-out data. We ultimately opted for the third
2Note that the exercise is more accurately characterized as replicability and not
reproducibility; the event predated ACM’s standardization of terminology.</p>
      <sec id="sec-4-1">
        <title>3https://github.com/osirrc/jig</title>
        <p>approach: the jig mounts a (read-only) data directory that makes
the document collection available at a known location, as part of
the contract between the jig and the image (and similarly for topics).
A separate directory that is writable serves as the mechanism for
the jig to gather output runs from the image for evaluation. This
method makes it possible for images to be tested on blind held-out
documents and topics, as long as the formats have been agreed to
in advance.</p>
        <p>
          Finally, any evaluation exercise needs to define the test collection.
We decided to focus on newswire test collections because their
smaller sizes support a shorter iteration and debug cycle (compared
to, for example, larger web collections). In particular, we asked
participants to focus on the TREC 2004 Robust Track test collection,
in part because of its long history: a recent large-scale literature
meta-analysis comprising over one hundred papers [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] provides a
rich context to support historical comparisons.
        </p>
        <p>Participants were also asked to accommodate the following two
(more recent) test collections if time allowed:
• TREC 2017 Common Core Track, on the New York Times
Annotated Corpus.
• TREC 2018 Common Core Track, on the Washington Post Corpus.
Finally, a “reach” goal was to support existing web test collections
(e.g., GOV2 and ClueWeb). Although a few submitted images do
support one or more of these collections, no formal evaluation was
conducted on them.
3.2</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Implementation Details</title>
      <p>In this section we provide a more detailed technical description
of the jig. Note, however, that the jig is continuously evolving as
we gather more image contributions and learn about our design
shortcomings. We invite interested readers to consult our code
repository for the latest details and design revisions. To be clear, we
describe v0.1.1 of the jig, which was deployed for the evaluation.</p>
      <p>The jig is implemented in Python and communicates with the
Docker Engine via the Docker SDK for Python.4 In the current
specification, each hook corresponds to a script in the image that
has a specific name, resides at a fixed location, and obeys a
speciifed contract dictating its behavior. Each script can invoke its own
interpreter: common implementations include bash and Python.
Thus, via these scripts, the image has freedom to invoke arbitrary
code. In the common case, the hooks invoke features of an existing
open-source search engine packaged in the image.</p>
      <p>From the perspective of a user who is attempting to replicate
results using an image, two commands are available: one for
preparation (the prepare phase) and another for actually performing
the ad hoc retrieval run (the search phase). The jig handles the
execution lifecycle, from downloading the image to evaluating run
ifles using trec_eval. This is shown in Figure 1 as a timeline in the
canonical lifecycle, with the jig on the left and the Docker image
on the right. The two phases are described in detail below.</p>
      <p>During the prepare phase, the user issues a command specifying
an image’s repository (i.e., name) and tag (i.e., version) along with
a list of collections to index. As part of the contract between the jig
and an image, the jig mounts the document collections and makes</p>
      <sec id="sec-5-1">
        <title>4https://docker-py.readthedocs.io/en/stable/</title>
        <p>User specifies
&lt;image&gt;:&lt;tag&gt;
prepare</p>
        <p>phase
&lt;image&gt;:&lt;tag&gt;
search
phase
jig</p>
        <p>Starts image
Triggers hook
Triggers hook
Creates snapshot</p>
        <p>&lt;snapshot&gt;
Triggers hook with snapshot</p>
        <p>run files
trec_eval</p>
        <p>Docker
image
init hook
index hook
search hook
them readable by the image for indexing (see discussion in the
previous section). The jig triggers two hooks in the image:
• First, the init hook is executed. This hook is meant for actions
such as downloading artifacts, cloning repositories and compiling
source code, or downloading external resources (e.g., a knowledge
graph). Alternatively, these steps can be encoded directly in the
image itself, thus making init a no-op. These two mechanisms
largely lead to the same end result, and so it is mostly a matter
of preference for the image developer.
• Next, the index hook is executed. The jig passes in a JSON string
containing information such as the collection name, path, format,
etc. required for indexing. The image manages its own index,
which is not directly visible to the jig.</p>
        <p>After indexing has completed, the jig takes a snapshot of the image
via a Docker commit. This is useful as indexing generally takes
longer than a retrieval run, and this design allows multiple runs (at
diferent times) to be performed using the same index.</p>
        <p>During the search phase, the user issues a command specifying
an image’s repository (i.e., name) and tag (i.e., version), the
collection to search, and a number of auxiliary parameters such as the
topics file, qrels file, and output directory. This hook is meant to
perform the actual ad hoc retrieval runs, after which the jig
evaluates the output with trec_eval. Just as in the index hook, relevant
parameters are encoded in JSON. The image places run files in the
/output directory, which is mapped back to the host; this allows
the jig to retrieve the run files for evaluation.</p>
        <p>In addition to the two main hooks for ad hoc retrieval
experiments, the jig also supports additional hooks for added functionality
(not shown in Figure 1). The first of these is the interact hook
that allows a user to interactively explore an image in the state that
has been captured via a snapshot, after the execution of the index
hook. This allows, for example, the user to “enter” an interactive
shell in the image (via standard Docker commands), and allows
the user to explore the inner workings of an image. The hook also
allows users to interact with services that a container may choose
to expose, such as an interactive search interface, or even Jupyter
notebooks. With the interact hook, the container is kept alive
in the foreground, unlike the other hooks which exit immediately
once execution has finished.</p>
        <p>Finally, images may also implement a train hook, enabling an
image to train a retrieval model, tune hyper-parameters, etc. after
the index hook has been executed. The train hook allows the user
to specify training and test splits for a set of topics, along with
a model directory for storing the model. This output directory is
mapped back to the host and can be passed to the search hook
for use during retrieval. Currently, training is limited to the CPU,
although progress has been made to support GPU-based training.</p>
        <p>In the current design, the jig runs one image at a time, but
additional tooling around the jig includes a script that further automates
all interactions with an image so that experiments can be run end to
end with minimal human supervision. This script creates a virtual
machine in the cloud (currently, Microsoft Azure), installs Docker
engine and associated dependencies, and then runs the image using
the jig. All output is then captured for archival purposes.
4</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>SUBMITTED IMAGES AND RESULTS</title>
      <p>Although we envision OSIRRC to be an ongoing efort, the reality of
a physical SIGIR workshop meant that it was necessary to impose
an arbitrary deadline at which to “freeze” image development. This
occurred at the end of June, 2019. At that point in time, we received
17 images by 13 diferent teams, listed alphabetically as follows:
• Anserini (University of Waterloo)
• Anserini-bm25prf (Waseda University)
• ATIRE (University of Otago)
• Birch (University of Waterloo)
• Elastirini (University of Waterloo)
• EntityRetrieval (Ryerson University)
• Galago (University of Massachusetts)
• ielab (University of Queensland)
• Indri (TU Delft)
• IRC-CENTRE2019 (Technische Hochschule Köln)
• JASS (University of Otago)
• JASSv2 (University of Otago)
• NVSM (University of Padua)
• OldDog (Radboud University)
• PISA (New York University and RMIT University)
• Solrini (University of Waterloo)
• Terrier (TU Delft and University of Glasgow)
All except for two images were designed to replicate runs for the
TREC 2004 Robust Track test collection, which was the primary
target for the exercise. The EntityRetrieval image was designed
to perform entity retrieval (as opposed to ad hoc retrieval). The
IRC-CENTRE2019 image packages a submission to the CENTRE
reproducibility efort, 5 which targets a specific set of runs from
the TREC 2017 Common Core Track. A number of images also
support the Common Core Track test collections from TREC 2017
and 2018. Finally, a few images also provide support for the GOV2
and ClueWeb test collections, although these were not evaluated.</p>
      <p>Following the deadline for submitting images, the organizers ran
all images “from scratch” with v0.1.1 of the jig and the latest release
of each participant’s image. Using our script (see Section 3.2), each
image was executed sequentially on a virtual machine instance in
the Microsoft Azure cloud. Note that it would have been possible
to speed up the experiments by running the images in parallel,
each on its own virtual machine instance, but this was not done.
We used the instance type Standard_D64s_v3, which according
to Azure documentation is either based on the 2.4 GHz Intel Xeon
E5-2673 v3 (Haswell) processor or 2.3 GHz Intel Xeon E5-2673 v4
(Broadwell) processor. Since we have no direct control over the
physical hardware, it is only meaningful to compare eficiency
(i.e., performance metrics such as query latency) across diferent
images running on the same virtual machine instance. Nevertheless,
our evaluations focused solely on retrieval efectiveness. This is a
shortcoming, since a number of images packaged search engines
that emphasize query evaluation eficiency.</p>
      <p>The results of running the jig on the submitted images
comprise the “oficial” OSIRRC 2019 image library, and is available
on GitHub.6 We have captured all log output, run files, as well as
trec_eval output. These results are summarized below.</p>
      <p>For the TREC 2004 Robust Track test collection, 13 images
generated a total of 49 runs, the results of which are shown in
Table 1; the specific version of the image is noted. Efectiveness is
measured using standard rank retrieval metrics: average precision
(AP), precision at rank cutof 30 (P30), and NDCG at rank cutof 20
(NDCG@20). The table does not include runs from the following
images: Solrini and Elastirini (which are identical to Anserini runs),
EntityRetrieval (where relevance judgments are not available since
it was designed for a diferent task), and IRC-CENTRE2019 (which
was not designed to produce results for this test collection).</p>
      <p>As the primary goal of this workshop is to build community,
infrastructure, and consensus, we deliberately attempt to minimize
direct comparisons of run efectiveness in the presentation: runs
are grouped by image, and the image themselves are sorted
alphabetically. Nevertheless, a few important caveats are necessary for
proper interpretation of the results: Most runs perform no
parameter tuning, although at least one implicitly encodes cross-validation
results (e.g., Birch). Also, runs might use diferent parts of the
complete topic: the “title”, “description”, and “narrative” (as well as
various combinations). For details, we invite the reader to consult
the overview paper by each participating team.</p>
      <p>We see that the submitted images generate runs that use a diverse
set of retrieval models, including query expansion and
pseudorelevance feedback (Anserini, Anserini-bm25prf, Indri, Terrier),
term proximity (Indri and Terrier), conjunctive query processing
(OldDog), and neural ranking models (Birch and NVSM). Several
images package open-source search engines that are primarily focused
on eficiency (ATIRE, JASS, JASSv2, PISA). Although we concede
that there is an under-representation of neural approaches,
relative to the amount of interest in the community at present, there
are undoubtedly replication challenges with neural ranking
models, particularly with their training regimes. Nevertheless, we are
pleased with the range of systems and retrieval models that are
represented in these images.</p>
      <sec id="sec-6-1">
        <title>6https://github.com/osirrc/osirrc2019-library</title>
        <p>Results from the TREC 2017 Common Core Track test collection
are shown in Table 2. On this test collection, we have 12 runs from
6 images. Results from the TREC 2018 Common Core Track test
collection are shown in Table 3: there are 19 runs from 4 images.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5 FUTURE VISION AND ONGOING WORK</title>
      <p>
        Our eforts complement other concurrent activities in the
community. SIGIR has established a task force to implement ACM’s policy
on artifact review and badging [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and our eforts can be viewed as
a technical feasibility study. This workshop also complements the
recent CENTRE evaluation tasks jointly run at CLEF, NTCIR, and
TREC [
        <xref ref-type="bibr" rid="ref11 ref6">6, 11</xref>
        ]. One of the goals of CENTRE is to define appropriate
measures to determine whether and to what extent replicability
and reproducibility have been achieved, while our eforts focus on
how these properties can be demonstrated technically. Thus, the jig
can provide the means to achieve CENTRE goals. Given fortuitous
alignment in schedules, participants of CENTRE@CLEF2019 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
were encouraged to participate in our workshop, and this in fact
led to the contribution of the IRC-CENTRE2019 image.
      </p>
      <p>From the technical perspective, we see two major shortcomings
of the current jig implementation. First, the training hook is not as
well-developed as we would have liked. Second, the jig lacks GPU
support. Both will be remedied in a future iteration.</p>
      <p>
        We have proposed and prototyped a technical solution to the
replicability challenge specifically for the SIGIR community, but
the changes we envision will not occur without a corresponding
cultural shift. Sustained, cumulative empirical progress will only
be made if researchers use our tools in their evaluations, and this
will only be possible if images for the comparison conditions are
available. This means that the community needs to adopt the norm
of associating research papers with source code for replicating
results in those papers. However, as Voorhees et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] reported,
having a link to a repository in a paper is far from suficient. The
jig provides the tools to package ad hoc retrieval experiments in a
standard way, but these tools are useless without broad adoption.
The incentive structures of academic publishing need to adapt to
encourage such behavior, but unfortunately this is beyond the scope
of our workshop.
      </p>
      <p>
        Given appropriate extensions, we believe that the jig can be
augmented to accommodate a range of batch retrieval tasks. One
important future direction is to add support for tasks beyond batch
retrieval, for example, to support interactive retrieval (with real or
simulated user input) and evaluations on private and other sensitive
data. Moreover, our efort represents a first systematic attempt
to embody the Evaluation-as-a-Service paradigm [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] via Docker
containers. We believe that there are many possible paths forward
building on the ideas presented here.
      </p>
      <p>Finally, we view our eforts as a stepping stone toward
reproducibility, and beyond that, generalizability. While these two
important desiderata are not explicit goals of our workshop, we note
that the jig itself can provide the technical vehicle for delivering
reproducibility and generalizability. Some researchers would want
to package their own results in a Docker image. However, there is
nothing that would prevent researchers from reproducing another
team’s results, that is then captured in a Docker image conforming
to our specifications. This would demonstrate reproducibility as
well as replicability of those reproducibility eforts. The jig also
supports mechanisms for evaluations on document collections and
information needs beyond those that an image was originally
designed for. This aligns with intuitive notions of what it means for a
technique to be generalizable.</p>
      <p>Overall, we believe that our eforts have moved the field of
information retrieval forward both in terms of supporting “good science”
as well as sustained, cumulative empirical progress. This work
shows that it is indeed possible to coordinate a large,
communitywide replication exercise in ad hoc retrieval, and that Docker
provides a workable foundation for a common interface and lifecycle
specification. We invite the broader community to join our eforts!</p>
    </sec>
    <sec id="sec-8">
      <title>6 ACKNOWLEDGEMENTS</title>
      <p>We would like to thank all the participant who contributed Docker
images to the workshop. This exercise would not have been possible
without their eforts. Additional thanks to Microsoft for providing
credits on the Azure cloud.</p>
      <p>ANT_r4_100_percent.BM25+.s-stem.RF 0.2184 0.3199 0.4211</p>
      <sec id="sec-8-1">
        <title>Anserini</title>
        <p>Anserini
Anserini
Anserini
Anserini
Anserini</p>
      </sec>
      <sec id="sec-8-2">
        <title>Anserini-bm25prf</title>
        <p>Anserini-bm25prf</p>
      </sec>
      <sec id="sec-8-3">
        <title>ATIRE</title>
      </sec>
      <sec id="sec-8-4">
        <title>Birch</title>
        <p>Birch
Birch
Birch
Birch
Birch
Birch
Birch
Birch
Birch
Birch
Birch</p>
      </sec>
      <sec id="sec-8-5">
        <title>Galago ielab</title>
      </sec>
      <sec id="sec-8-6">
        <title>Indri</title>
        <p>Indri
Indri
Indri
Indri
Indri
Indri
Indri
Indri
Indri</p>
      </sec>
      <sec id="sec-8-7">
        <title>JASS</title>
      </sec>
      <sec id="sec-8-8">
        <title>JASSv2</title>
      </sec>
      <sec id="sec-8-9">
        <title>NVSM</title>
      </sec>
      <sec id="sec-8-10">
        <title>OldDog</title>
        <p>OldDog</p>
      </sec>
      <sec id="sec-8-11">
        <title>PISA</title>
      </sec>
      <sec id="sec-8-12">
        <title>Terrier</title>
        <p>Terrier
Terrier
Terrier
Terrier
Terrier
Terrier
Terrier
Terrier
Terrier
0.2531
0.2903
0.2895
0.2467
0.2747
0.2774
0.2916
0.2928
0.3102
0.3365
0.3333
0.3079
0.3232
0.3229
0.3396
0.3438
JASSv2_c17_10
core17-1000
v0.1.1
v0.1.1
v0.1.1
v0.1.1
v0.1.1
v0.1.1
v0.1.1
v0.1.3
v0.1.3
v0.1.1
v0.1.1
v0.1.3
0.4293
0.5093
0.4980
0.4467
0.4827
0.4953
0.5613
0.6347</p>
        <p>AP
0.2495
0.2920
0.3136
0.2526
0.2966
0.3073
0.1802
0.2381
core18-1000
0.2384 0.3500 0.3927</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jaime</given-names>
            <surname>Arguello</surname>
          </string-name>
          , Matt Crane, Fernando Diaz,
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Lin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Trotman</surname>
          </string-name>
          .
          <source>2015. Report on the SIGIR 2015 Workshop on Reproducibility</source>
          , Inexplicability, and
          <article-title>Generalizability of Results (RIGOR)</article-title>
          .
          <source>SIGIR Forum 49</source>
          ,
          <issue>2</issue>
          (
          <year>2015</year>
          ),
          <fpage>107</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Timothy</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Armstrong</surname>
            , Alistair Mofat,
            <given-names>William</given-names>
          </string-name>
          <string-name>
            <surname>Webber</surname>
            , and
            <given-names>Justin</given-names>
          </string-name>
          <string-name>
            <surname>Zobel</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Improvements That Don't Add Up: Ad-Hoc Retrieval Results Since 1998</article-title>
          .
          <source>In Proceedings of the 18th International Conference on Information and Knowledge Management (CIKM</source>
          <year>2009</year>
          ). Hong Kong, China,
          <fpage>601</fpage>
          -
          <lpage>610</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Monya</given-names>
            <surname>Baker</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Is There a Reproducibility Crisis? Nature 533 (</article-title>
          <year>2016</year>
          ),
          <fpage>452</fpage>
          -
          <lpage>454</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Ferro</surname>
          </string-name>
          , Norbert Fuhr, Maria Maistro, Tetsuya Sakai, and
          <string-name>
            <given-names>Ian</given-names>
            <surname>Soborof</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Overview of CENTRE@CLEF 2019: Sequel in the Systematic Reproducibility Realm. In Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Tenth International Conference of the CLEF Association (CLEF</source>
          <year>2019</year>
          ). Lugano, Switzerland.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Ferro</surname>
          </string-name>
          and
          <string-name>
            <given-names>Diane</given-names>
            <surname>Kelly</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>SIGIR Initiative to Implement ACM Artifact Review and Badging</article-title>
          .
          <source>SIGIR Forum 52</source>
          ,
          <issue>1</issue>
          (
          <year>2018</year>
          ),
          <fpage>4</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Ferro</surname>
          </string-name>
          , Maria Maistro, Tetsuya Sakai, and
          <string-name>
            <given-names>Ian</given-names>
            <surname>Soborof</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Overview of CENTRE@CLEF 2018: A First Tale in the Systematic Reproducibility Realm</article-title>
          .
          <source>In Proceedings of the Ninth International Conference of the CLEF Association (CLEF</source>
          <year>2018</year>
          ). Avignon, France,
          <fpage>239</fpage>
          -
          <lpage>246</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Juliana</given-names>
            <surname>Freire</surname>
          </string-name>
          , Norbert Fuhr, and Andreas Rauber (Eds.).
          <source>2016. Report from Dagstuhl Seminar</source>
          <volume>16041</volume>
          :
          <article-title>Reproducibility of Data-Oriented Experiments in e-Science. Schloss Dagstuhl-Leibniz-Zentrum für Informatik</article-title>
          , Germany.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Frank</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          , Allan Hanbury, Henning Müller, Ivan Eggel, Krisztian Balog, Torben Brodt,
          <string-name>
            <surname>Gordon</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Cormack</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jimmy Lin</surname>
          </string-name>
          , Jayashree
          <string-name>
            <surname>Kalpathy-Cramer</surname>
            , Noriko Kando, Makoto P. Kato, Anastasia Krithara, Tim Gollub, Martin Potthast, Evelyne Viega, and
            <given-names>Simon</given-names>
          </string-name>
          <string-name>
            <surname>Mercer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Evaluation-as-a-Service for the Computational Sciences: Overview and Outlook</article-title>
          .
          <source>ACM Journal of Data and Information Quality (JDIQ) 10</source>
          ,
          <issue>4</issue>
          (November
          <year>2018</year>
          ),
          <volume>15</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          :
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The Neural Hype and Comparisons Against Weak Baselines</article-title>
          .
          <source>SIGIR Forum 52</source>
          ,
          <issue>2</issue>
          (
          <year>2018</year>
          ),
          <fpage>40</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Jimmy</surname>
            <given-names>Lin</given-names>
          </string-name>
          , Matt Crane, Andrew Trotman, Jamie Callan, Ishan Chattopadhyaya, John Foley, Grant Ingersoll, Craig Macdonald, and
          <string-name>
            <given-names>Sebastiano</given-names>
            <surname>Vigna</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Toward Reproducible Baselines: The Open-Source IR Reproducibility Challenge</article-title>
          .
          <source>In Proceedings of the 38th European Conference on Information Retrieval (ECIR</source>
          <year>2016</year>
          ). Padua, Italy,
          <fpage>408</fpage>
          -
          <lpage>420</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Tetsuya</surname>
            <given-names>Sakai</given-names>
          </string-name>
          , Nicola Ferro, Ian Soborof, Zhaohao Zeng, Peng Xiao, and
          <string-name>
            <given-names>Maria</given-names>
            <surname>Maistro</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Overview of the NTCIR-14 CENTRE Task</article-title>
          .
          <source>In Proceedings of the 14th NTCIR Conference on Evaluation of Information Access Technologies</source>
          . Tokyo, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Ellen</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Voorhees</surname>
            , Shahzad Rajput, and
            <given-names>Ian</given-names>
          </string-name>
          <string-name>
            <surname>Soborof</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Promoting Repeatability Through Open Runs</article-title>
          .
          <source>In Proceedings of the 7th International Workshop on Evaluating Information Access (EVIA</source>
          <year>2016</year>
          ). Tokyo, Japan,
          <fpage>17</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Wei</surname>
            <given-names>Yang</given-names>
          </string-name>
          , Kuang Lu,
          <string-name>
            <given-names>Peilin</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Critically Examining the “Neural Hype”: Weak Baselines and the Additivity of Efectiveness Gains from Neural Ranking Models</article-title>
          .
          <source>In Proceedings of the 42nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR</source>
          <year>2019</year>
          ). Paris, France.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Ruifan</surname>
            <given-names>Yu</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Yuhao</given-names>
            <surname>Xie</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <year>2018</year>
          . H2oloo at TREC 2018:
          <article-title>CrossCollection Relevance Transfer for the Common Core Track</article-title>
          .
          <source>In Proceedings of the Twenty-Seventh Text REtrieval Conference (TREC</source>
          <year>2018</year>
          ). Gaithersburg, Maryland.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>