<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>A Literature Review on Reproducibility Studies in Computer Science</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>TobiasHummel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>JohannesManner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Reproducibility, Replicability, Repeatability, Experiment Documentation</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Distributed Systems Group, University of Bamberg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <fpage>54</fpage>
      <lpage>62</lpage>
      <abstract>
        <p>Researchers expect a clear and well-documented experiment from industry experience reports and experimental research papers. All necessary configuration parameters, the source code, the experiment's machine configuration etc. should be documented in such a way that readers can interpret the results of these publication in detail. If this is the case, an interested reader is capable of redoing the experiments stated and verifying the results of others. This ability to properly interpret experiments and even reproduce them is a cornerstone of good scientific practice.</p>
      </abstract>
      <kwd-group>
        <kwd>https</kwd>
        <kwd>//www</kwd>
        <kwd>uni-bamberg</kwd>
        <kwd>de/pi/team/manner-johann(Je</kwd>
        <kwd>sM/anner)</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>“Reproducible Research in all sciences is critical to the advancement of knowledge.</p>
      <p>It is what enables a researcher to build upon, or refute, previous research</p>
      <p>allowing the field to act as a collective of knowledge
rather than as tiny uncommunicated clusters” - [1, Cacho and Taghva p. 519].</p>
      <p>This quote illustrates the importance of reproducibility of scientific work and describes an
ideal state to gain knowledge. Nevertheless, the reality does not reflect this ideal state. In
2016, a survey with 1,576 respondents was published in Nat2u].r9e0[% of them stated that
the scientific community in general is facing a reproducibility crisis. To understand this crisis,
this paper tries to shed some light into this complex topic and states the current situation for</p>
      <p>
        However, in research these terms are sometimes used interchang3e]a.bRleypr[oducibility
and repeatability are even used the other way ar4o].unTdhe[re is also a standard called
Reproducibility Badging and Definitions published by the National Information Standards
Organization (NISO)5[]. There, they define further nuances of the three introduced terms above.
Despite their depth, this standard being around for already three years is only referenced by
27 Google Scholar h2itws hen searching for the title of this publication. For this paper, we
will stick to the ACM definitions introduced above due to their clarity and adaptation in the
community.Collberg and others expected computer science to be in a special role compared
to other discipline“sre:producing the work published in a systems conference or journal should be
as simple as going to the authors’ website, downloading their code and data, typing ‘make,’ and
seeing if the results correspond to the published ones” [4, p. 1]. However, in a 30 minutes time
frame they were able to retrieve and build the source code of only 32.3% of the papers analyzed.
Other secondary studies, e.6g,. 7[
        <xref ref-type="bibr" rid="ref8">, 8</xref>
        ], confirm this issue. Only 3 out of 26 experiments are
reproducible based on an assessment within the early Function as a Service (FaaS) research
domain [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Another study concluded thmatajaority of 122 is not reproducibl7e].[ Couture
and others8[] revealed that only 26% of 315 data projects published raw data.
      </p>
      <p>All these studies should be a wake-up call for the computer science domain to publish raw
data, source code and scripts. For without this background information a correct interpretation
of results is not always possible for others which prevents the correct evaluation of the merit
of a paper. This can be illustrated with an example from FaaS research: Tw9o, S10L]Rs [
revealed papers reporting about unusual performance results when running a function on a
cloud provider. Properly interpreting the original results based on the provider documentation
showed that the measurements where due to a memory setting. It unintentionally assigned
more than one CPU to the function. Without enough details about the experimental setup, it
would not have been possible to falsify the misinterpretations of the original authors. As a
consequence, conference chairs and publishers should incentivise researchers to make their
experiments open to the public and enable others to reproduce experimental results.
1https://www.acm.org/publications/policies/artifact-review-and-badging-current
2Google Scholar search was performed otnho1f9February 2024.</p>
      <p>Since we are interested in the current state of reproducibility studies, the objective of this
paper is not to conduct a further reproducibility study but to analyze the already published
studies by answering the following three research questions:
RQ1 What are publication trends for reproducibility studies in computer science?
RQ2 Why do authors try to reproduce the work of others?
RQ3 How successful are the reproducibility studies and which challenges do the authors
encounter?</p>
      <p>RQ1 focuses on the number of studies and visible trends. The reasons for performing such
studies are questioned RinQ2. And the last questioRnQ,3, reveals success numbers on how
many studies were able to reproduce the original work. One caveat here could be a publication
bias, where reproducibility problems could be more likely to be published than1s1u].ccFeosrs [
upcoming reproducibility studies, we also look at challenges the included publications faced to
share aggregated learnings.</p>
      <p>The agenda of our paper is as follows: In Sec2t,iowne shortly introduce the SLR methodology
and present some numbers and an overview of included papers. Results for our three research
questions are presented as subsections in Sec3t.iSoenction4 concludes the paper with a short
summary and an outlook to future work.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Methodology</title>
      <p>
        We conducted a SLR to understand the current state of reproducibility studies in the computer
science domain. A SLR is well suited for summarizing and synthesizing the current status on a
topic in a fair way11[]. For the search phase, ACM Digital Library, IEEE Digital Library, DBLP
and Google Scholar were used. While the first three libraries have a computer science focus,
Google Scholar was chosen as a more general search engine with a large corpus. Therefore, we
only used the first 50 entries at Google Scholar, sorted by relevance to get a manageable set of
literature as already done by other S1L2,R1s3[
        <xref ref-type="bibr" rid="ref14">, 14</xref>
        ]. Figure1 summarizes the search process.
      </p>
      <p>ACM Digital Library
IEEE Digital Library</p>
      <p>DBLP
Google Scholar</p>
      <p>Initial Search
(198)
12
5
131
50 (limited)</p>
      <p>Filtering
(38)
5
3
16
14</p>
      <p>Merge and
Deduplication</p>
      <p>
        The search term required the phrase “reproducibility study” to be part of the title or abstract.
As the search engines difer in their search options, the exact search string1s5]d. iIfenrt[otal,
198 entries were found in the initial search phase. The inclusion and exclusion criteria are based
on the research questions and were applied to title, abstract and co1n5c]l.uIsniocnase[s of
doubt, the full text was skimmed. The selection criteria were developed upfront and piloted on
some studies as recommended bKyitchenham andCharters [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Nevertheless, they evolved
over the process. The inclusion criteria were as follows:
• Reproducibility studies in the area of computer science, other disciplines were excluded
immediately.
• Publications in which an attempt is made to reproduce the findings of one or several prior
publications.
• Publications by universities, public research institutes or industry.
      </p>
      <p>• In addition to peer-reviewed publications, also preprints.</p>
      <p>
        The following exclusion criteria were used:
• Publications in which the authors do not use or do not mention a specific software
implementation.
• Publications that belong to the field of computational science rather than computer
science, i.e. that are more concerned with the application of computer science in fields
such as medicine.
• Publications in languages other than English (language is often used as an exclusion
criterion1[
        <xref ref-type="bibr" rid="ref11 ref17 ref6">6, 17, 11</xref>
        ]).
      </p>
      <p>All raw data for the SLR are archived at Ze1n8o]d.oIn[clusion and exclusion decisions as
well as all other data related to the SLR can be found there. In case of exclusion, a reason is
stated. Although requiredKbitychenham andCharters, the quality of the studies was not
assessed for every paper11[]. It is assumed that in particular search engines from ACM, IEEE
and DBLP have results with high quality. Otherwise, studies of questionable quality would have
been excluded during the data extraction phase.</p>
      <p>
        After the merge and deduplication step, 20 results remained. During data extraction, one
paper turned out to be a replicability study and was thus ex1c9l]u.Wdeidth[in the remaining, 14
belonged to the information retrieval do2m0a,i2n1,[
        <xref ref-type="bibr" rid="ref22 ref23 ref24 ref25 ref26 ref27 ref28 ref29 ref30 ref31 ref32 ref33">22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33</xref>
        ].
So as not to over-represent this field, out of these 14 papers the first four papers based on the
ifrst author names2[
        <xref ref-type="bibr" rid="ref21 ref22 ref23">0, 21, 22, 23</xref>
        ] were chosen as examples from this computer science domain
to discuss RQ2 and RQ3. Table 1 shows the remaining nine publications and the computer
science sub-field to which they belong.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Results</title>
      <sec id="sec-4-1">
        <title>3.1. RQ1: Publication Facts and Trends</title>
        <p>The most obvious fact in our SLR investigation is that the information retrieval domain is a role
model for reproducibility studies. 74% of our identified papers (14/19) are from this domain.
The implementation of diferent ranking and retrieval algorithms is inherent to this computer
science sub-field. For those an interpretation of results is only feasible by checking the raw data</p>
        <sec id="sec-4-1-1">
          <title>A Comparison between Term-Independence Retrieval Models for [23]</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Ad Hoc Retrieval</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>A Reproducibility Study of Question Retrieval for Clarifying Ques- [21] tions</title>
        </sec>
        <sec id="sec-4-1-4">
          <title>Cross-Domain Retrieval in the Legal and Patent Domains: A Repro- [20] ducibility Study</title>
        </sec>
        <sec id="sec-4-1-5">
          <title>Cyberbullying Detection in Social Networks Using Deep Learning [22]</title>
        </sec>
        <sec id="sec-4-1-6">
          <title>Based Models; A Reproducibility Study</title>
        </sec>
        <sec id="sec-4-1-7">
          <title>Examining the Reproducibility of Using Dynamic Loop Scheduling [36]</title>
        </sec>
        <sec id="sec-4-1-8">
          <title>Techniques in Scientific Applications</title>
        </sec>
        <sec id="sec-4-1-9">
          <title>IPAL: Breaking up Silos of Protocol-Dependent and Domain-Specific [37]</title>
        </sec>
        <sec id="sec-4-1-10">
          <title>Industrial Intrusion Detection Systems</title>
          <p>Ref
Sub-Field</p>
        </sec>
        <sec id="sec-4-1-11">
          <title>Information Retrieval</title>
        </sec>
        <sec id="sec-4-1-12">
          <title>Natural Language</title>
        </sec>
        <sec id="sec-4-1-13">
          <title>Processing</title>
        </sec>
        <sec id="sec-4-1-14">
          <title>Distributed Systems</title>
        </sec>
        <sec id="sec-4-1-15">
          <title>Security and Privacy</title>
        </sec>
        <sec id="sec-4-1-16">
          <title>Machine Learning Based Invariant Generation: A Framework and [38]</title>
        </sec>
        <sec id="sec-4-1-17">
          <title>Reproducibility Study</title>
        </sec>
        <sec id="sec-4-1-18">
          <title>Software tion</title>
        </sec>
        <sec id="sec-4-1-19">
          <title>Verifica</title>
          <p>and experiment details. The same holds true for another little cluster, namely Natural Language
Processing (NLP). Here, the research objectives are also highly dependent on input data and its
processing.</p>
          <p>tsano 456
i
lcPub 23
i
fo 1
# 0
2016
2017
2018
2021
2022</p>
          <p>2023
2019 2020</p>
          <p>Publication Year</p>
          <p>When we searched for literature, we did not limit the publication year. Nevertheless, the
oldest paper in our set of filtered literature was published in 2016.2Fsihgouwres the number
of publications by their publication year. The bar for 2023 is colored orange since the search
phase was conducted in May 2023, so further reproducibility studies might have been published
in the rest of 2023. The distribution of publications over time shows that the number of papers
which target reproducibility concerns has been rising over the last years.</p>
          <p>Another facet of publication trends are the venues where papers get submitted and presented.
The European Conference on Information Retrieval is the top venue for reproducibility studies. A
majority of the information retrieval papers were published there (9/14). One reason for this
high number of papers is a special reproducibility track with a dedicated call3f.oTrhpiaspers
shows that dedicated conference formats can support the reproducibility efort to bring experts
together in one place.</p>
          <p>To summarize the insights foRrQ1: Information retrieval is the top domain for reproduciblity
research due to the domain specific challenges and supported by dedicated conference tracks.
Reproducibility studies are a trending topic when looking at the rising publication numbers.</p>
          <p>For answering the remaining research questions, we focus on the nine papers 1o.f Table</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. RQ2: Reasons for Reproducibility</title>
        <p>
          Five of the nine papers from Tab1l,efocus upon the reproducibility assessme3n5,t2[
          <xref ref-type="bibr" rid="ref1 ref22 ref34 ref36">1, 22, 36, 34</xref>
          ].
In the other cases, the main contribution lies on other aspects but all publications included
implicit or explicit statements why the reproducibility study was conducted. The motivations
are summarized in the following list which ansRwQer2s, Why do authors try to reproduce the
work of others?:
• Creating awareness for reproducibil3i5t,y21[
          <xref ref-type="bibr" rid="ref22 ref34 ref36">, 22, 36, 34</xref>
          ].
• The analyzed paper is important for the research domain but the artifacts of the original
study were not availab2l1e][.
• Starting point for own resear2c0h,2[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
• Follow-up research after introducing a tool or framework to show that it works
correctly2[
          <xref ref-type="bibr" rid="ref3 ref36 ref37 ref38">3, 38, 36, 37</xref>
          ].
        </p>
        <p>
          As an additional finding it could be shown that some studies mix terms and approaches.
As introduced, we distinguish based on the ACM terms reproducibility, repeatability and
replicability. Three of the publications analyzed in the SLR had a replicability study as an
additional par2t0,[
          <xref ref-type="bibr" rid="ref22 ref37">22, 37</xref>
          ]. That is, the authors first conducted a reproducibility study and then
reused the setup for another dataset. There is a trade-of between an experimental setup as
close as possible to the original setup, a reasonable use of resources, and the re-usability for
other research questions. Nevertheless, the distinction is important since the motivation and
results are influenced by the combination of team members and experimental setup.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. RQ3: Success Ratios and Challenges</title>
        <p>The success ratios for the nine included papers reveal a mixed picture. Four studies were
successful, three were partly successful and two were not successful at all. The two biggest
challenges were missing artifacts and lack of documentation.</p>
        <p>
          Studies were considered successful when the measured values were close to the originally
reported ones22[
          <xref ref-type="bibr" rid="ref23 ref34 ref37">, 23, 34, 37</xref>
          ]. For instanceD, advar and others stated that “the majority of
the reproduced results were within the standard deviation of the referenc2e2,rpe.s6u]l.ts” [
Nunzio and others report that they “have quite comparable results in terms of processing time”
3https://ecir2023.org/calls/reproducibility.html?v=3.8
and the “final scores difer from the original values by almost 2 percentage po34in,pt.s”34[].
For three publications, partly successful in this context means that the authors investigated
several primary studies where some reproductions fa35il,e3d8,[
          <xref ref-type="bibr" rid="ref36">36</xref>
          ]. The two remaining studies
failed in their eforts20[
          <xref ref-type="bibr" rid="ref21">, 21</xref>
          ].
        </p>
        <p>One challenge of the studies was the availability of artifacts like source code or raw data. For
three of the publications considered in this SLR the necessary artifacts were2a2v,2a3i,l3a4b]le [
and the reproducibility for these studies succeeded. Regarding the other six, only incomplete
artifacts were available.</p>
        <p>Another big challenge was insuficient documentation of experimental sNetuunpzsio. and
others, for example, identified unclarities regarding the environment of the primary study. The
hardware was not suficiently described in the original publication, especially concerning the
CPU and GPU as both could influence the execution time. In the original paper, there were
two contradictory statements about the used GPU model. Additionally, the original authors
trained one of their models in the cloud because they needed a more powerful environment. In
comparison, the authors of the reproducibility study needed the cloud environment for two
models which prevented the comparison of the training time for the additional model.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion</title>
      <p>Reproducibility is a key requirement for science and has to be considered when starting a new
research project. In reproducibility studies, researchers other than the original authors use a
setup as close as possible and reasonable to reproduce and validate the work of others. Starting
from this, the objective of this paper was to give an overview about reproducibility studies in
computer science.</p>
      <p>The publication trends showed that the information retrieval domain is a role model for other
domains, supporting its researchers with dedicated conference tracks. Based on the number of
publications per year, we saw that the reproducibility topic gains traction. Additionally, we
could confirm other meta studies that showed that a lack of raw data and artifacts as well as an
incomplete or missing documentation are the most serious challenges for good research.</p>
      <p>For future work, we propose a follow-up study on reproducibility which should include
additional search engines and incorporate further SLR tasks like sn4o.wWbaelallisnogplan
to incorporate the feedback from reviewers to include the term “executable paper” which was
a hyped term around 2011 as an additional search term. In addition, we want to broaden our
scope by incorporating similar terms like “analysis” and “survey” paired with “reproducibility”,
“replicability” and “repeatability” to also include publications which use another term but target
“reproducibility studies” in the sense of this paper.
4Snowballing was already performed for this study and the raw results are already available at Zenodo. Due to time
and space constraints, these snowballed publications were not assessed but published for follow-up research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. R. F.</given-names>
            <surname>Cacho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Taghva</surname>
          </string-name>
          ,
          <article-title>The state of reproducible research in computer science</article-title>
          ,
          <source>in: Proc. of ITNG</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Baker</surname>
          </string-name>
          ,
          <volume>1</volume>
          ,
          <string-name>
            <surname>500</surname>
          </string-name>
          <article-title>scientists lift the lid on reproducibility</article-title>
          ,
          <source>Nature</source>
          <volume>533</volume>
          (
          <year>2016</year>
          )
          <fpage>452</fpage>
          -
          <lpage>454</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H. E.</given-names>
            <surname>Plesser</surname>
          </string-name>
          ,
          <article-title>Reproducibility vs. replicability: A brief history of a confused terminology</article-title>
          ,
          <source>Frontiers in Neuroinformatics</source>
          <volume>11</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Collberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Proebsting</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Warren</surname>
          </string-name>
          ,
          <article-title>Repeatability and benefaction in computer systems research - a study and a modest proposal</article-title>
          ,
          <source>University of Arizona Technical Report 14</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] NISO, Reproducibility Badging and Definitions: A Recommended Practice of the National Information Standards Organization</article-title>
          ,
          <source>Technical Report niso-rp-31-2021, National Information Standards Organization (NISO)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kuhlenkamp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Werner</surname>
          </string-name>
          , Benchmarking FaaS Platforms:
          <article-title>Call for Community Participation</article-title>
          ,
          <source>in: Proc. of WoSC</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kalibera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <article-title>Rigorous benchmarking in reasonable time</article-title>
          ,
          <source>in: Proc. of ISMM</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Couture</surname>
          </string-name>
          , et al.,
          <article-title>A funder-imposed data publication requirement seldom inspired data sharing</article-title>
          ,
          <source>PLOS ONE 13</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Manner</surname>
          </string-name>
          , SeMoDe
          <article-title>- simulation and benchmarking pipeline for function as a service</article-title>
          ,
          <source>in: Bamberger Beiträge zur Wirtschaftsinformatik und Angewandten Informatik</source>
          ,
          <volume>105</volume>
          ,
          <string-name>
            <surname>Otto-</surname>
          </string-name>
          Friedrich-University,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Manner</surname>
          </string-name>
          ,
          <article-title>A Simulation Framework for Function as</article-title>
          a Service, University of Bamberg Press,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kitchenham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Charters</surname>
          </string-name>
          ,
          <article-title>Guidelines for performing systematic literature reviews in software engineering</article-title>
          , Keele University and University of Durham
          <source>Technical Report</source>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kolb</surname>
          </string-name>
          ,
          <article-title>On the Portability of Applications in Platform as a Service</article-title>
          ,
          <source>Ph.D. thesis, Bamberg</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Manner</surname>
          </string-name>
          ,
          <article-title>A structured literature review approach to define serverless computing and function as a service</article-title>
          ,
          <source>in: Proc. of CLOUD</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Taibi</surname>
          </string-name>
          , et al.,
          <article-title>Patterns for serverless functions (function-as-a-service): A multivocal literature review</article-title>
          ,
          <source>in: Proc. of CLOSER</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Brereton</surname>
          </string-name>
          , et al.,
          <article-title>Lessons from applying the systematic literature review process within the software engineering domain</article-title>
          ,
          <source>Journal of Systems and Software</source>
          <volume>80</volume>
          (
          <year>2007</year>
          )
          <fpage>571</fpage>
          -
          <lpage>583</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Scheuner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leitner</surname>
          </string-name>
          ,
          <article-title>Function-as-a-service performance evaluation: A multivocal literature review</article-title>
          ,
          <source>Journal of Systems and Software</source>
          <volume>170</volume>
          (
          <year>2020</year>
          )
          <fpage>110708</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Yussupov</surname>
          </string-name>
          , et al.,
          <article-title>A systematic mapping study on engineering function-as-a-service platforms and tools</article-title>
          ,
          <source>in: Proc. of UCC</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hummel</surname>
          </string-name>
          ,
          <string-name>
            <surname>A Literature</surname>
          </string-name>
          <article-title>Review on Reproducibility Studies in Computer Science</article-title>
          : Supporting Material,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R. F. G.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Paixao</surname>
          </string-name>
          , M. de Almeida Maia,
          <article-title>Duplicate question detection in stack overflow: A reproducibility study</article-title>
          ,
          <source>in: Proc. of SANER</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Althammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hofstätter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <article-title>Cross-domain retrieval in the legal and patent domains: A reproducibility study</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Cross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mourad</surname>
          </string-name>
          ,
          <article-title>A reproducibility study of question retrieval for clarifying questions</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dadvar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <article-title>Cyberbullying detection in social networks using deep learning based models; A reproducibility study</article-title>
          , CoRR abs/
          <year>1812</year>
          .08046 (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>E. K. F.</given-names>
            <surname>Dang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W. P.</given-names>
            <surname>Luk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Allan</surname>
          </string-name>
          ,
          <article-title>A comparison between term-independence retrieval models for ad hoc retrieval</article-title>
          ,
          <source>ACM Transactions on Information Systems</source>
          <volume>40</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Färber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Klein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sigloch</surname>
          </string-name>
          ,
          <article-title>Neural citation recommendation: A reproducibility study</article-title>
          ,
          <source>in: Proc. of BIR</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hendriksen</surname>
          </string-name>
          , et al.,
          <article-title>Scene-centric vs. object-centric image-text cross-modal retrieval: A reproducibility study</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          , et al.,
          <article-title>State encoders in reinforcement learning for recommendation: A reproducibility study</article-title>
          ,
          <source>in: Proc. of SIGIR</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>W.</given-names>
            <surname>Lajewska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <article-title>From baseline to top performer: A reproducibility study of approaches at the TREC 2021 conversational assistance track</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          , et al.,
          <article-title>Improving query representations for dense retrieval with pseudo relevance feedback: A reproducibility study</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mackenzie</surname>
          </string-name>
          , et al.,
          <article-title>Compressing inverted indexes with recursive graph bisection: A reproducibility study</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sequiera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Update delivery mechanisms for prospective information needs: A reproducibility study</article-title>
          ,
          <source>in: Proc. of CHIIR</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          , et al.,
          <article-title>Statistical stemmers: A reproducibility study</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          , G. Zuccon,
          <article-title>Federated online learning to rank with evolution strategies: A reproducibility study</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>P.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <article-title>A reproducibility study of information retrieval models</article-title>
          ,
          <source>in: Proc. of ICTIR</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>G. M. D. Nunzio</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Minzoni</surname>
          </string-name>
          ,
          <article-title>A thorough reproducibility study on sentiment classification: Methodology, experimental setting</article-title>
          ,
          <source>results, Information</source>
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>76</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>M.</given-names>
            <surname>Arvan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parde</surname>
          </string-name>
          ,
          <article-title>Reproducibility in computational linguistics: Is source code enough?</article-title>
          ,
          <source>in: Proc. of EMNLP</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hofeins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Ciorba</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Banicescu</surname>
          </string-name>
          ,
          <article-title>Examining the reproducibility of using dynamic loop scheduling techniques in scientific applications</article-title>
          ,
          <source>in: Proc. of IPDPS</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wolsing</surname>
          </string-name>
          , et al.,
          <article-title>Ipal: Breaking up silos of protocol-dependent and domain-specific industrial intrusion detection systems</article-title>
          ,
          <source>in: Proc. of RAID</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Haltermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wehrheim</surname>
          </string-name>
          ,
          <article-title>Machine learning based invariant generation: A framework and reproducibility study</article-title>
          ,
          <source>in: Proc. of ICST</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>