<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Story of an Experiment: A Provenance-based Semantic Approach towards Research Reproducibility</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sheeba Samuel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kathrin Groeneveld</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frank Taubert</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Walther</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tom Kache</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Teresa Langenstuck</string-name>
          <email>teresa.langenstueckg@med.uni-jena.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Birgitta Konig-Ries</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>H. Martin Bucker</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Biskup</string-name>
          <email>christoph.biskupg@uni-jena.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Biomolecular Photonics Group, Jena University Hospital, Friedrich Schiller University</institution>
          ,
          <addr-line>Jena</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Computer Science, Friedrich Schiller University</institution>
          ,
          <addr-line>Jena</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Tom.Kache</institution>
          ,
          <addr-line>christoph.biskup</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>End-to-end reproducibility of scienti c experiments is a key to the foundation of science. Reproducibility of an experiment does not necessarily guarantee the accuracy of its results, but it guarantees that the steps of an experiment can be repeated to a certain level of signi cance to generate similar results. Data provenance plays a key role in telling the story of an experiment which helps one step towards reproducibility. To convey the message of a story, it is essential to provide su cient data and its ow along with its semantics. In this paper, we present a provenance-based semantic approach to explain the story of a scienti c experiment with the primary goal of reproducibility. The REPRODUCE-ME ontology extended from PROV-O and P-Plan is used to represent the whole story of an experiment describing the path it took from its design to result. We visualize and evaluate the provenance lifecycle of a scienti c experiment taking into account the use case of life science experiments.</p>
      </abstract>
      <kwd-group>
        <kwd>Provenance</kwd>
        <kwd>Reproducibility</kwd>
        <kwd>Experiment</kwd>
        <kwd>Story</kwd>
        <kwd>Ontology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A story generally consists of the following components: plot, characters,
background context, settings, events, con icts, climax, and the nal message. It is
essential to know the characters, the context, and the ow of the story to
understand its climax and message. Similarly, to make the story of a scienti c
experiment and its results understandable and reproducible, it is necessary to
present its agents, execution, environmental attributes and work ow in a way
that can be understood by the scienti c community. According to [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], an
experiment performed at time T with the environment setup E (e.g. settings) using
data D (e.g. experiment materials, measurements) consisting of a sequence of
steps S is said to be reproducible, if it can be executed at time T 0 &gt; T with
environment setup E0 (similar or identical to E) with a sequence of steps S0
(modi ed from or equal to S ) using data D0 (similar or identical to D) with
similar results. There are various challenges that hinder reproducibility of
experiments which include integration of data generated from di erent devices,
incomplete and uncertain provenance information, lack of documentation in
digital media, lack of knowledge of the type of data and their formats and most
importantly their semantics.
      </p>
      <p>
        In life-science experiments, the preparation of experimental materials is very
important. Scientists use the methods described in publications to prepare the
specimens. But sometimes, the scientists fail to replicate the methods mentioned
in the publications due to incorrect or incomplete data because of accidental
omission or errors. Critical steps may not be included or fully described or the
order of execution may be missing from the method description. Inconclusive
results are often omitted. But such negative results are sometimes useful for
experiments carried out by other scientists. Sources of reagents can also result in
signi cantly di erent results [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. So it is essential to know the attribution of these
materials to replicate the method. Simple typographical errors or experiments
done with di erent units like using 10 milligrams (mg) instead of 10 micrograms
( g) can make a di erence. Thus, it is important to capture the provenance of
an experiment with these ne details.
      </p>
      <p>
        The aim of our work is to capture the provenance data from multiple resources
of an experiment and provide the ability to visualize the data along with its
semantics. Our contributions of this paper are as follows: (i) Identifying the
components and competency questions needed to present and validate the story
of a scienti c experiment using microscopy experiments as an example. (ii)
Presenting our provenance-based semantic approach using REPRODUCE-ME [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
ontology by extending PROV-O [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and P-Plan [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to represent the di erent
paths taken from an input to an output of an experiment, the steps and the
input and output variables of each step. (iii) Visualization of the provenance data
of an experiment as a dashboard to the scientists in our prototype, CAESAR.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>State of the Art</title>
      <p>
        Missier [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] presents three main challenges for the practical usability of
provenance data, one of which is the role of provenance in the reproducibility of the
scienti c processes. A well-developed approach to capture provenance is based
on Scienti c Work ow Management Systems. These are mature systems which
ensure reproducibility in computational sciences by tracking and recording all
the evolution of scienti c work ows made by the user [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. But the provenance
captured by these systems is coupled to a work ow de nition and does not
include the detailed information of the execution environment (e.g., temperature)
or the standard operating procedures followed in generating the materials of an
experiment. They give more importance to provenance at the execution level of
a work ow than on the entire lifecycle of an experiment. Many ontologies have
also been introduced and developed to describe the work ow of computational
experiments such as OBI [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and CMPO [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. However, these ontologies do not
use the standard PROV model preventing the interoperability of the collected
data.
      </p>
      <p>
        Access to primary research data is important for scientists. Many public
repositories have been created to host and share computational experiments with the aim
of preservation of provenance components. The environment myExperiment [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
is an online web-portal which provides researchers the facility to share their
scienti c work ows. Image Data Resource (IDR) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] is another public
repository for sharing imaging data and links data from public chemical databases
with controlled ontologies. Most of these platforms either focus on publishing
datasets or the work ow part of an experiment.
      </p>
      <p>Capturing provenance of non-computational parts of an experiment when not
using a scienti c work ow management system is challenging. We present our
prototype which is a scienti c data management platform where a user can
document the experimental data like in a laboratory notebook. The semantics of the
experimental data is then represented using an ontology to represent the whole
story of an experiment including the path taken, the steps etc. Our prototype is
di erent because we represent the whole picture of an experiment using
provenance standards like PROV-O and P-Plan as well as give equal importance to
the computational and non-computational provenance part of the experiment,
unlike other systems, which target mostly the computational part. We also give
equal importance to the agents involved directly or indirectly in an experiment
because that may also a ect the reproducibility of experiments.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Provenance-based Semantic Approach</title>
      <p>
        The motivation of our work arises from the Collaborative Research Center (CRC)
ReceptorLight3 where scientists from two universities4 , two university hospitals5
and a non-university research institute6 work together to understand the function
of membrane receptors and develop high-performance microscopy techniques.
Interviews with the scientists in the CRC as well as a workshop conducted to foster
reproducible science7 helped us to understand the di erent scienti c practices
followed in their experiments and their requirements of reproducibility and data
management. We collected a list of competency questions from these oral
interviews from the scientists from various projects performing di erent kinds of
experiments. The relevance of these questions is further supported by their large
overlap with competence questions obtained in other contexts, e.g. the
provenance challenge [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. From the collected questions, we selected the ones which
were commonly told by scientists and generalized them. Here we present the
3 http://www.receptorlight.uni-jena.de/
4 https://www.uni-jena.de/, https://www.uni-wuerzburg.de/
5 http://www.uniklinikum-jena.de/,http://www.ukw.de/
6 http://www.ipht-jena.de/
7 http://fusion.cs.uni-jena.de/bexis2userdevconf2017/workshop/
most common competency questions of which answers are required to describe
and validate the story of an experiment.
1. What are the input and output variables of an experiment?
2. Which are the methods and standard operating procedures used?
3. Which are the les and materials that were used in a particular step?
4. Which are the steps involved in an experiment which used a particular
material?
5. What is the complete path taken by a scientist for an experiment?
6. Which are the instruments that are associated with an experiment and their
settings when the output was generated?
7. Which are the agents directly or indirectly responsible for an experiment?
8. Who created this experiment and when? Who modi ed it and when?
9. Which are the publications or external resources that were referenced in each
step of an experiment?
10. List all the experiments which use growth protocol (EFO 0003789) and
studies on \Homo sapiens" and resulted in phenotype \shorter prophase" which
passed the quality control.
      </p>
      <p>Question 10 is an example query speci c to life science experiments. To answer
these kinds of competency questions, we developed an ontology to represent the
conceptual model of an experiment.
3.1</p>
      <sec id="sec-3-1">
        <title>Development of REPRODUCE-ME ontology</title>
        <p>
          The REPRODUCE-ME ontology8 is extended from PROV-O to represent all
entities, agents, activities and their relationships [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. It also extends P-Plan to
represent the steps of the activities or events involved in an experiment in detail.
Figure 1 shows an excerpt of the classes and properties of the
REPRODUCEME ontology depicting the lifecycle of a scienti c experiment. We explain how
we added classes and properties to the ontology to answer each of the
competency questions.
        </p>
        <p>
          To answer the competency questions from 1 to 5, we describe an Experiment
as p-plan:Plan which in turn also is a prov:Entity by inference rules. The
object property p-plan:isSubPlanOfPlan is used to associate an experiment with
its subplans. Each subplan of an experiment consists of several smaller steps
p-plan:Step which uses input and output variables p-plan:Variable which are
represented using p-plan:hasInputVar and p-plan:hasOutputVar object
properties. For example, the HighContentScreening is a step of Experiment and
ImageAcquisition step has Image as an output variable. The complete path taken
by an experiment is described by ordering these steps using the object property
p-plan:isPrecededBy. For example, the execution order of cells of a Jupyter
Notebook9, a subplan of an experiment, is described using the p-plan:isPrecededBy
property [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
8 https://w3id.org/reproduceme
9 https://jupyter.org/
isPrecededBy
        </p>
        <p>xsd:dateTime
hasOutputVar</p>
        <p>isSubPlanOfPlan</p>
        <p>Output
rdf:type
ExperimentalData</p>
        <p>rdf:type
version
generated
AtTime</p>
        <p>Study
has
Experimental</p>
        <p>Material
hasData
rdf:type</p>
        <p>RawData
ProcessedData
xsd:string
wasRevisionOf</p>
        <p>rdf:type
Library</p>
        <p>Organism
License</p>
        <p>HighContentScreening
isOutputVarOf</p>
        <p>hasOutputVar
Cell
isStepOfPlan</p>
        <p>Notebook
isStepOfPlan</p>
        <p>Publication
wasAttributedTo
Agent
hasRole
Author
ImageAcquisition
hasOutputVar</p>
        <p>Image
correspondsToVariable</p>
        <p>Instrument
rdf:type</p>
        <p>Microscope
rdf:type
isStepOfPlan
Protocol
isSubPlanOfPlan
isSubPlanOfPlan
Experiment
has</p>
        <p>ECxopnedriitmioenntal
Experimental
Material</p>
        <p>ECxopnedriitmioenntal</p>
        <p>Setting
rdf:type</p>
        <p>hasSetting
ManufacturerSpec</p>
        <p>Model
REPRODUCE-ME
P-Plan</p>
        <p>PROV-O
hasInputVar
Source
rdf:type</p>
        <p>
          Variable
Images are an integral part of life-science experiments which involve microscopes
for their acquisition. The acquisition, analysis, and annotation properties of a
biological microscopic image are added as part of the ontology using the OME
Data Model [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The class Image represents all the features of an image. For
question 6, various instruments used in these experiments like microscopes are
added as a subclass of Instrument, while the settings of these instruments are
added as p-plan:Variable. Each of the instruments has ManufacturerSpec which
consists of Manufacturer, Model, SerialNumber and LotNumber. Apart from this
information, these instruments have certain attributes of their own. For example,
Laser has data properties like Wavelength. In addition to the static properties of
an instrument, some properties or settings are changed during an experiment to
capture the image in a particular way. These properties are represented through
Settings.
        </p>
        <p>To answer the competency questions from 7 to 9, the following classes were
added. A story needs characters to proceed. The agents are represented through
the prov:Agent. Each person has a prov:Role in the experiment like Experimenter,
Distributor, or Manufacturer. The resources and devices used in an experiment
are either distributed by Distributor or manufactured by Manufacturer.
Sometimes these resources are produced in the laboratory using a method represented
in a Publication.</p>
        <p>The time is another important factor in the story of an experiment. The data
properties like prov:generatedAtTime, receivedAtTime, and modi edAtTime are
used to describe the time of each event. The ResearchGroup and
ResearchProject represents the plot which describes the group and the project/institute for
which the experiment was performed. The results of an experiment are the main
outcome of an experiment which is represented by Output. The nal message of
the story of an experiment is represented using Rating, which is the rating given
to the experiment and its description represents if any problems occur during
its execution.</p>
        <p>The StandardOperatingProcedure is a p-plan:Plan which describes the procedure
of a method. The File is a variable which is also an experimental data. One
variable is associated with another variable using the object property reference. Some
variables are also added as prov:Entity so that properties associated with agent,
entity and time can also be used. For example, File is a p-plan:Variable as well
as a prov:Entity. Together, all the constructs described so far enable answering
questions like Question 10.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Development of the semantic-based scienti c data management platform</title>
        <p>
          For the experimental data management, we present our prototype, CAESAR
(CollAborative Environment for Scienti c Analysis with Reproducibility), which
is extended from OMERO [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. The OMERO software, developed by the Open
Microscopy Environment Consortium, is an open-source imaging database
platform for experimental biology. The framework has a plugin architecture with a
rich set of features including analyzing and modifying images and supporting
over 140 image le formats using BIO-Formats [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. BIO-Formats is an image
translation library which reads and converts the proprietary microscopy data to
an open standard model which then can be used by other tools. With the help
of BIO-Formats, OMERO automatically extracts the image acquisition data
including the data about devices and their settings.
        </p>
        <p>
          CAESAR extends the OMERO framework so that scientists can document their
experimental data along with their images [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The platform provides a
formbased provenance capture system where a user can record experimental data
like in their laboratory notebook. Each experiment has multiple steps where
each step follows a Standard Operating Procedure (SOP). These procedures can
either be les or Jupyter Notebooks. The user can upload the les and link the
experiment materials used in each step of an experiment at a speci c step.
The form also provides auto-completion of data. For example, if a user enters a
Chemical Abstracts Service (CAS) number of a chemical compound, the
webclient calls the CAS registry web service and auto- lls the chemical formula,
molecular weight, and structural formula of the compound. The form also
provides a virtual keyboard so that the user can insert special symbols, sub and
superscript text since they are widely used in the documentation of the
experimental data which include chemical formula and other structures. The prototype
employs the user and group management provided by the OMERO platform.
Groups in OMERO enable sharing of data between users. Users have role and
permissions to restrict the modi cation of data. A user may belong to one or
many groups. The data are shared between the users in the same group in the
same OMERO server. The data can be made available to members of other
groups based on the permission level of the group. Based on the permission
level, if a user does not have the right to modify other member's data, then
that user can propose changes to the experiment. This request is done using the
proposal feature provided by CAESAR. The members of the current group or
other groups can provide suggestions and propose the experimental data. The
owner of the experiment receives those suggestions as proposals. The user has
two options: First, to accept the proposal and add it to the current
experimental data. Second, to reject the proposal and delete the proposal. CAESAR also
provides a facility for the user to view the version history of an experiment to
see all the changes made in its description.
        </p>
        <p>
          In order to capture the provenance of the computational part of a scienti c
experiment, JupyterHub10 is installed and connected to CAESAR so that users
can create new notebooks, run and share them [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. We capture the provenance
of the execution of interactive notebooks used in a scienti c experiment over the
course of time with the help of ProvBook [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], an extension of Jupyter Notebook.
The stored provenance data, also available as RDF, include the start and end
time of the execution, the total time it took to run the cell and the input and
output of the cell. The provenance di erence feature provided by ProvBook
provides the users to compare their results with the original results of the author of
the notebook and detect which factors a ected their results. In this way, we
capture and link the provenance of both the computational and non-computational
part of a scienti c experiment to represent its story.
        </p>
        <p>
          The data in the OMERO server which include the image metadata and
experimental data are stored in the PostgreSQL relational database. To semantically
represent this data and avoid replicating the data, we used ontology-based data
access techniques to convert the relational database data to RDF data [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. We
created a mapping between the conceptual layer and the relational database
layer. So now we have a mapping between the OMERO database and the
experimental database. Every eld in the database is mapped to the
REPRODUCEME ontology. The ontology-based mapping is done using Ontop11. The mapping
with Ontop created new terms in the ontology based on the mapping. Manual
intervention was needed to remove the unnecessary terms and add the correct
terms from the REPRODUCE-ME ontology.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Visualization of Provenance Data</title>
        <p>Visualization of provenance data is another important feature of CAESAR
provided to the user using a dashboard. Figure 2 shows a part of the project
dashboard. Data are visualized at the project level, where multiple experiments are
evaluated together. The competency questions described in Section 3 were
converted to SPARQL queries and the answers to these questions are represented
as tables in the dashboard. The dashboard provides a panel for each component
of a story. The data are represented as data tables so that users can search and
10 http://jupyter.org/hub
11 http://ontop.inf.unibz.it/</p>
        <p>Fig. 2. The Project Dashboard in CAESAR
lter the data. The plot panel provides the dates, research group, and research
project associated with an experiment. The characters panel provides data about
the agents responsible for an experiment. The materials panel displays the
materials used in an experiment. The external resources and les panel display the
publications and les associated with the materials and each step of an
experiment. The devices panel shows all the devices associated with an experiment
and the settings panel displays all the settings including settings of the devices.
In addition to the dashboard, we also provide a SPARQL query editor with
SPARQL templates so that answers to the questions like 10 can be obtained.
In addition to the existing parameterized SPARQL templates, the user can also
write their own SPARQL queries related to experimental data and get results.
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Evaluation</title>
        <p>We conducted both a user and a data-based evaluation on two di erent datasets
to check whether the data along with the ontology can explain the story of an
experiment. The REPRODUCE-ME ontology, the supplementary materials used
for the evaluation and the results are publicly available12. All the evaluations are
based on the list of competency questions described in Section 3. The metrics
used for the evaluation were the usability of the system and whether the
competency questions were answered. For the user-based evaluation, a group of four
scientists working with high-end light microscopy techniques from the project B1
of the CRC ReceptorLight evaluated the platform. Around ten di erent
experiments of type Fluorescence Resonance Energy Transfer (FRET) and confocal
Patch Clamp Fluorometry (cPCF) were used for the evaluation. Eleven di erent
types of experiment materials like chemicals, proteins, solutions etc. were used
as input of the experiments and around 70 microscopic images generated from
12 https://w3id.org/reproduceme/research
instruments with di erent settings were used for the evaluation. The results of
SPARQL queries in the dashboard were manually compared and their
correctness was evaluated by the domain experts. The evaluation results show that the
dashboard provided them with a complete overview of the experiments. Since
none of them (like most scientists) possesses SPARQL knowledge, such a
complete overview could not have been gained without the dashboard. The ability
to lter the results in each table of the dashboard also helped them to search
their queries.</p>
        <p>
          The other evaluation of the ontology was done with the data from the Image
Data Repository (IDR) which currently consists of around 35 imaging
experiments [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The metadata of each imaging study from the IDR datasets was
extracted and described in RDF using the REPRODUCE-ME ontology with
scripts13. The SPARQL queries generated from the competency questions were
executed also on this data to check whether the ontology can be used to describe
the story of other types of experiments as well. To illustrate the story of such an
experiment, we take an example from IDR. The study \Focused mitotic
chromosome condensation screen using HeLa cells" (idr000213) is an Experiment which
consists of ImagingStudy as a step. There are around 1160 Images which are the
output variables of ImagingStudy. The Publication and the ProcessedData le
are also the output variables of this experiment. Each Image in the experiment is
annotated with the GeneIdenti er and Phenotype. The Experiment is attributed
to several agents who take the role of Submitter of the experiment, Manufacturer
of the Library used and Author of Publication. The Experiment has several
Protocol s, which describe the various instructions followed in the experiment. The
results from data-based evaluations show that the provenance-based semantic
system helps in providing the whole story of an experiment along with all its
dependencies.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>Data provenance is a key factor towards reproducibility of scienti c experiments.
In this paper, we present a provenance-based semantic approach to explain the
story of a biological experiment from its plot to its output. The
REPRODUCEME ontology extended from the existing ontologies PROV-O and P-Plan, is used
to represent a whole picture of an experiment including the plot, characters,
settings, plans, steps, input and output variables. The ontology and the prototype
are validated through answering the competency questions. In future work, we
will focus on the scalability and performance of the system.</p>
      <sec id="sec-4-1">
        <title>Acknowledgements</title>
        <p>This work is supported by the Deutsche Forschungsgemeinschaft (German
Research Foundation, CRC/TRR 166 \High-end light microscopy elucidates
membrane receptor function - ReceptorLight", projects Z2 and B1).
13 The subsets of idr datasets converted to RDF are available here https://github.</p>
        <p>com/Sheeba-Samuel/REPRODUCE-ME/</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Allan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burel</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blackburn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Linkert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loynton</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>MacDonald</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>W.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neves</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patterson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.:
          <article-title>OMERO: exible, modeldriven data management for experimental biology</article-title>
          .
          <source>Nature methods 9</source>
          (
          <issue>3</issue>
          ),
          <volume>245</volume>
          {
          <fpage>253</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Reproducibility crisis: Blame it on the antibodies</article-title>
          .
          <source>Nature</source>
          <volume>521</volume>
          (
          <issue>7552</issue>
          ),
          <volume>274</volume>
          {
          <fpage>276</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bandrowski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brinkman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brochhausen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brush</surname>
            ,
            <given-names>M.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bug</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chibucos</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clancy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courtot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derom</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumontier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>The ontology for biomedical investigations</article-title>
          .
          <source>PloS one 11(4)</source>
          ,
          <year>e0154556</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chirigati</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freire</surname>
          </string-name>
          , J.:
          <source>Provenance and Reproducibility</source>
          , pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          . Springer New York, New York, NY (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Garijo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gil</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Augmenting PROV with plans in P-Plan: scienti c processes as linked data</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhagat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aleksejevs</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruickshank</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michaelides</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borkum</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bechhofer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , et al.:
          <article-title>myExperiment: a repository and social network for the sharing of bioinformatics work ows</article-title>
          .
          <source>Nucleic acids research 38(suppl 2)</source>
          ,
          <source>W677{W682</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>I.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burel</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Creager</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falconi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hochheiser</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mellen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sorger</surname>
            ,
            <given-names>P.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swedlow</surname>
            ,
            <given-names>J.R.:</given-names>
          </string-name>
          <article-title>The open microscopy environment (OME) data model and XML le: open tools for informatics and quantitative analysis in biological imaging</article-title>
          .
          <source>Genome biology</source>
          <volume>6</volume>
          (
          <issue>5</issue>
          ),
          <source>R47</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jupp</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burdett</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heriche</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ellenberg</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parkinson</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustici</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The cellular microscopy phenotype ontology</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          <volume>7</volume>
          (
          <issue>1</issue>
          ),
          <volume>28</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lebo</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sahoo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belhajjame</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheney</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corsar</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garijo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soiland-Reyes</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zednik</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <string-name>
            <surname>PROV-O: The PROV</surname>
          </string-name>
          <article-title>Ontology</article-title>
          .
          <source>W3C Recommendation</source>
          <volume>30</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Missier</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>The lifecycle of provenance metadata and its associated challenges and opportunities</article-title>
          . In: Building Trust in Information, pp.
          <volume>127</volume>
          {
          <fpage>137</fpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Missier</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Woodman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hiden</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Provenance and data di erencing for work ow reproducibility analysis</article-title>
          .
          <source>Concurrency and Computation: Practice and Experience</source>
          <volume>28</volume>
          (
          <issue>4</issue>
          ),
          <volume>995</volume>
          {
          <fpage>1015</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Ludascher,
          <string-name>
            <surname>B.</surname>
          </string-name>
          , et al.:
          <article-title>Special issue: The rst provenance challenge</article-title>
          .
          <source>Concurrency and computation: practice and experience 20(5)</source>
          ,
          <volume>409</volume>
          {
          <fpage>418</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Samuel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <article-title>Konig-</article-title>
          <string-name>
            <surname>Ries</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>REPRODUCE-ME: ontology-based data access for reproducibility of microscopy experiments</article-title>
          . In: The Semantic Web:
          <article-title>ESWC 2017 Satellite Events</article-title>
          , Portoroz, Slovenia. pp.
          <volume>17</volume>
          {
          <issue>20</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Samuel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <article-title>Konig-</article-title>
          <string-name>
            <surname>Ries</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Combining P-Plan and the REPRODUCE-ME ontology to achieve semantic enrichment of scienti c experiments using interactive notebooks</article-title>
          .
          <source>In: The Semantic Web: ESWC 2018 Satellite Events</source>
          . pp.
          <volume>126</volume>
          {
          <issue>130</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Samuel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <article-title>Konig-</article-title>
          <string-name>
            <surname>Ries</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>ProvBook: Provenance-based semantic enrichment of interactive notebooks for reproducibility</article-title>
          .
          <source>In: Proceedings of the ISWC</source>
          <year>2018</year>
          <article-title>Posters &amp; Demonstrations, Industry and Blue Sky Ideas Tracks co-located with ISWC 2018, Monterey</article-title>
          , USA (
          <year>2018</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2180</volume>
          /paper-57.pdf
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Samuel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taubert</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walther</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <article-title>Konig-</article-title>
          <string-name>
            <surname>Ries</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Bucker, H.M.:
          <article-title>Towards reproducibility of microscopy experiments</article-title>
          .
          <source>D-Lib Magazine</source>
          <volume>23</volume>
          (
          <issue>1</issue>
          /2) (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustici</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarkowska</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chessel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antal</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferguson</surname>
            ,
            <given-names>R.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarkans</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          , et al.:
          <article-title>Image data resource: a bioimage data integration and publication platform</article-title>
          .
          <source>Nature methods 14(8)</source>
          ,
          <volume>775</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>