<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>FAIRDOM approach for semantic interoperability of systems biology data and models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga Krebs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katy Wolstencroft</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natalie Stanford</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Norman Morrison</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Golebiewski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rostyk Kuzyakiv</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stuart Owen</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Quyen Nguyen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacky Snoep</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wolfgang Mueller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carole Goble</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Heidelberg Institute for Theoretical Studies</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Leiden Institute of Advanced Computer Science</institution>
          ,
          <addr-line>Leiden, NL</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computer Science, University of Manchester</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Zurich</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Motivation: The ability to collect and interlink heterogeneous data and model collections is essential in systems biology. Effective data exchange and comparison requires sufficient data annotation. This is particularly apparent in systems biology, where data heterogeneity means that multiple community metadata standards are required for the annotation of a whole investigation, including data, models and protocols. Results: FAIRDOM (http://fair-dom.org/) is an initiative to enable the systems biology community to produce and publish FAIR Data, Operating procedures and Models. It allows research assets to be aggregated, interlinked and shared in the context of the systems biology investigations that produced them. Here we present the FAIRDOM strategy in the context of semantic data integration, and how it supports the whole life cycle of data collection, annotation, sharing and reuse of systems biology data and resources. Availability: https://fairdomhub.org * Contact: olga.krebs@h-its.org</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Data integration is an essential part of systems biology. Scientists
need to combine different sources of information in order to model
biological systems, and relate those models to available
experimental data for validation. Metadata is an important aspect of data
management and data sharing. Annotating experimental results with
a consistent set of information allows for easier discovery of relevant
data as well as enabling others to potentially reuse it. Metadata
ranges from simple descriptions about when an experiment was
done to more detailed descriptions of where biological samples
originated, how they were prepared, and what the experimental
conditions were at the time of the experiment. Currently, only a small
fraction of the data and models produced during systems biology
investigations are deposited for reuse by the community, and only a
smaller fraction of that data is standards compliant, semantically
enriched content.</p>
      <p>FAIRDOM project is a joint action of ERA-Net ERASysAPP
(https://www.erasysapp.eu/) and European Research Infrastructure
ISBE (http://project.isbe.eu/) to establish a data and model
management service facility for systems biology. Its prime mission is to
support researchers, students, trainers, funders and publishers by
enabling systems biology projects to make their Data, Operating
procedures and Models, Findable, Accessible, Interoperable and Reusable
(FAIR). FAIRDOM builds on the outcomes of the successful
SysMO-DB and SyBIT data management projects, uniting their tool
and database development as well as their experience serving large
systems biology projects. FAIRDOMHub is a web-based platform
comprising two main components: SEEK (http://seek4science.org)
as a web-based front-end cataloguing and metadata platform and
openBIS as a back-end LIMS for scalable local data collection and
processing (https://sis.id.ethz.ch/software/openbis.html). Here we
present the semantic data integration in SEEK, and how it supports
the whole life cycle of data collection, annotation, sharing, and reuse
of systems biology data and resources.
2</p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>
        The SEEK [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is based on the ISA infrastructure (Investigations,
Studies and Assays), a standard format for describing how
individual experiments (assays) are aggregated into wider studies and
investigations [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The Just Enough Results Model (JERM) describes
the interrelations between assets and the metadata fields required to
describe them. For example, for each dataset uploaded to SEEK, the
JERM describes what type of experiment it was, what was
measured, and what the values in the dataset mean. The JERM captures
the core elements of MIBBI metadata, allowing users to comply
with these standards as well as capturing the information required
for linking in SEEK. The JERM Ontology (available from the
BioPortal, http://bioportal.bioontology.org/ontologies/1488) is an
application ontology designed to describe the relationships between
items in SEEK (for example, data, models, experiment descriptions,
samples, protocols, standard operating procedures and
publications); and to enable these relationships to be expressed with formal
semantics. It is based on the idea of the Minimal Information Models
(https://www.biosharing.org), which have been collected under the
umbrella of MIBBI (Minimum Information for Biological and
Biomedical Investigations).
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>METHODS</title>
      <p>
        The majority of laboratory scientists use spreadsheets for the daily
management and manipulation of data, so the RightField semantic
spreadsheet application [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] (also part of this work) is used to embed
semantic annotation into the data. Individual cells, columns, or rows
in spreadsheets can be restricted to particular ranges of allowed
classes or instances from chosen ontologies. By embedding the JERM
metadata model in a spreadsheet format, and enabling the use of
JERM (and other) vocabulary terms for annotation, the process of
standardized semantic data collection can become part of the
existing data management activities in the laboratory. Bioinformaticians,
with experience in ontologies and data annotation, can prepare
RightField-enabled spreadsheets with embedded ontology term
selection support for distribution across the consortium.
      </p>
      <p>JERM-compliant spreadsheet templates have been developed for a
wide range of experimental data types, their collection is available
from http://docs.seek4science.org/help/templates.html.
By embedding semantic technologies into familiar data management
tools, the SEEK enables semantic annotation of new data and the
generation and querying of Linked Data - compliant datasets, whilst
hiding the complexities of ontologies and metadata from its users.
Underlying semantic web resources additionally extract and serve
SEEK metadata in RDF (Resource Description Format). RDF
enables rich semantic queries, both within SEEK and between related
resources in the web of Linked Open Data.</p>
    </sec>
    <sec id="sec-4">
      <title>ACKNOWLEDGEMENT</title>
      <p>This work was funded by the BBSRC (BBG0102181,
BB/I004637/1, BB/M013189/1), and by the BMBF grants 0315749,
20315781 and 031A525. We would like to thank the FAIRDOM
PALS and users for their valuable feedback, testing and comments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Wolstencroft</surname>
          </string-name>
          et al (
          <year>2015</year>
          )
          <article-title>SEEK: a systems biology data and model management platform</article-title>
          .
          <source>BMC Systems Biology (9)33 DOI:10</source>
          .1186/s12918-015-0174-y
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Rocca-Serra</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brandizi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maguire</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sklyar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
            , C., Be-gley,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Field</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hide</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hofmann</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          et al. (
          <year>2010</year>
          )
          <article-title>ISA software suite: supporting standards-compliant experimental annotation and enabling curation at the community level</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>26</volume>
          ,
          <fpage>2354</fpage>
          -
          <lpage>2356</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Wolstencroft</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Owen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horridge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krebs</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mueller</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoep</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>du Preez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2011</year>
          )
          <article-title>RightField: embedding ontology annotation in spreadsheets</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>27</volume>
          ,
          <fpage>2021</fpage>
          -
          <lpage>2022</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>