<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Capturing Provenance for a Linkset of Convenience</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simon Jupp</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>James Malone</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alasdair J G Gray</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Heriot-Watt University</institution>
          ,
          <addr-line>Edinburgh</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Trust Genome Campus</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Biological interactions such as those between genes and proteins are complex and require intricate OWL models. However, direct links between biological entities can support search and data integration. In this paper we introduce linksets of convenience that capture these direct links. We show the provenance statements required to track the derivation of such linksets; linking them back to the full biological justi cation.</p>
      </abstract>
      <kwd-group>
        <kwd>Data linking</kwd>
        <kwd>Provenance</kwd>
        <kwd>VoID</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Investigating biological systems, such as those implicated in disease, necessitates
the connection of many levels of biology; gene, gene variation, gene expression,
protein structure, signalling pathways, phenotypic, epidemiological data and so
on. The ability to integrate data across these levels relies on links that can be
formed between biological entities, for example, going from a gene to proteins or
proteins to pathways. For each of these links there is some biological justi cation
that may involve several steps (see Section 2 for details). To support tasks such
as search and data integration it is convenient to provide additional shortcuts in
the form of a direct link, e.g. genes to pathways.</p>
      <p>Modeling the true nature of the links using semantic web technologies such
as OWL removes ambiguity when working with data by giving it a well de ned
and precise semantics. However it increases the complexity of interacting with
the data as the OWL model needs to capture the full intricacies of the biological
interactions. As we move to publish biological data as linked open data, there
is an opportunity to describe direct links between di erent types of biological
entities as a shortcut to be made between entities which feature in common
queries, such as gene to protein; capturing the way that biologists often discuss
the domain and enable novel integrations of the data. These direct links provide
a working notion that cuts through the biology but which does not necessitate
capturing (or recapturing) the complex multivariate relationships that can hold
between the two entities. Such linksets are already used to support the Open</p>
      <p>so:has_part
so:gene
so:transcript</p>
      <p>so:polypeptide
Ensembl Gene
so:transcribed_from</p>
      <p>Ensembl Transcript Ensembl Protein
so:translates_to
:ep2upRelation
uniprot:Protein</p>
      <p>
        UniProt Protein
skos:related
PHACTS Discovery Platform [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], although those linksets do not have adequate
provenance.
      </p>
      <p>
        In this paper we propose a mechanism to model these links of convenience
using a combination of VoID linksets [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and PROV [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We avoid
misrepresenting links by applying semantically weaker relationships together with additional
provenance which represents the underlying complexity. We illustrate the model
with an example using data from two popular biological databases.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Linking genes to proteins use case.</title>
      <p>
        We motivate our work with an example mapping between Ensembl [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (a database
of genome annotation) and Uniprot [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (a database of protein sequences). These
databases already contain cross-references between an Ensembl Gene (EG) and
a Uniprot Protein (UP). However to understand how this mapping is generated
you currently need to discover the correct publications and online
documentation; they are not directly discoverable from the data.
      </p>
      <p>Biological theory tells us that a gene encodes for a protein, although this
biological relation only truly holds for the link between the EG and the Ensembl
Protein (EP) entity. There are in fact multiple types of UP to EP mappings, for
instance they can be derived from an exact sequence identity or they might be
based on a percentage sequence identity. Figure 1 illustrates how we model EG
to EP using terminology de ned in the Sequence Ontology, and for illustration
we include a superproperty of the all the EP to UP mappings that we call
ep2upRelation3. We introduce a link of convenience (dashed line) that links
the EG to UP that is there to support queries using the semantically weak
skos:related relation. This schema lacks the provenance to assert that the
related link of convenience is derived from the longer chain of semantically richer
links that hold from a gene to protein.
3 UniProt are currently extending their vocabulary to de ne these relations.
1 # define the ensembl protein partition
2 :ensembl void:classPartition :EPpartition .
3 :EPpartition void:class so:Polypeptide .
4
5 # define the Uniprot protein partition
6 :uniprot void:classPartition :UPpartition .
7 :UPpartition void:class uniprot:Protein .
8
9 # define the linkset that links the two partitions
10 :ensemblProteinToUniprotProteinLinkset a void:Linkset ;
void:linkPredicate :ep2upRelation ;
11
12
13 # define partitions for ensembl gene, gene transcript and
# transcript protein
:ensembl void:classPartition :ensemblGenePartition ;
void:propertyPartition :ensemblGeneTranscriptPartition ;
void:propertyPartition :ensemblTranscriptProteinPartition ;
:ensemblGenePartition void:class so:gene .
:ensemblGeneTranscriptPartition void:property so:transcribed_from .
:ensemblTranscriptProteinPartition void:property so:translates_to .</p>
      <p>The VoID vocabulary of linked datasets allows the description of RDF links
between datasets using VoID linksets. A linkset allows us to describe the links,
captured as a set of triples, between two datasets. We can use VoID to describe
relevant partitions of the datasets based on individual properties or classes, these
form new subsets that can participate in multiple linksets. In our scenario we
need to capture two crucial linksets; the rst is the EP to UP linkset, and the
second is the more convenient EG to UP linkset.</p>
      <p>The EP-UP linkset captures the :ep2upRelation link between types of EP
in the Ensembl dataset, and types of UP in the UniProt dataset (lines 10-11).
We describe two further subsets; the EP partition of all entities that are of type
so:Polypeptide in the Ensembl dataset (lines 2-3) and the UniProt subset of
all entities that are of type uniprot:Protein (lines 6-7).</p>
      <p>The EG to UP link of convenience needs a similar linkset description based
on an EG partition and the previous UP partition, although this time the
relation is skos:related (lines 25-26). We also want to capture that the triples
in this linkset are derived from another set of triples. This captures that the
skos:related is a shortcut relation for a more complex path through the RDF
graph. Again we can use VoID partitioning, but this time using a property based
partition to identify the EG to Ensembl Transcript (ET) and ET to EP links
(lines 15-20) . Finally we use the prov:wasDerivedFrom relation to link the
convenience linkset to the linksets that describe the full path of relations that the
shortcut represents (line 28-30).
4</p>
    </sec>
    <sec id="sec-3">
      <title>Discusion</title>
      <p>It is always important to try and model your data as accurately as possible,
and publishing data with RDF and OWL is well suited for this task. The VoID
vocabulary already provides a mechanism to de ne and attach provenance to
linksets between datasets, and we are proposing the use of PROV to connect
linksets that are derived from other linksets. As a Web of linked biological data
emerges, there is a need to identify links that are there for convenience, and
expose how they relate back to the core biological (OWL) model. In cases where
a link of convenience is derived from a series of other linksets, it is desirable to
be able to spot this and unpack the convenience links using common queries.
The model proposed supports this task but questions remain as to whether VoID
and PROV are enough, so we hope this preliminary work can help motivtate the
discussion.</p>
      <p>Acknowledgements
EBI contribution supported by EU FP7 BioMedBridges Grant 284209.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>A.J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loizou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Askjaer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brenninkmeijer</surname>
            ,
            <given-names>C.Y.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chichester</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evelo</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harland</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waagmeester</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.J.:</given-names>
          </string-name>
          <article-title>Applying linked data approaches to pharmacology: Architectural decisions and implementation</article-title>
          .
          <source>Semant. Web</source>
          <volume>5</volume>
          (
          <year>2014</year>
          )
          <volume>101</volume>
          {
          <fpage>113</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hausenblas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Describing Linked Datasets with the VoID Vocabulary</article-title>
          . Note,
          <source>W3C (March</source>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lebo</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sahoo</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mcguinness</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <string-name>
            <surname>PROV-O: The PROV</surname>
          </string-name>
          <article-title>Ontology</article-title>
          .
          <source>Technical report, W3C Recommendation</source>
          (
          <year>2013</year>
          ) http://www.w3.org/TR/prov-o/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Flicek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amode</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al:
          <year>Ensembl 2014</year>
          .
          <source>Nucleic acids research</source>
          <volume>42</volume>
          (
          <year>2014</year>
          )
          <article-title>D749{D755 doi:</article-title>
          10.1093/nar/gkt1196.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. The UniProt Consortium:
          <article-title>Activities at the universal protein resource (UniProt)</article-title>
          .
          <source>Nucleic acids research</source>
          <volume>42</volume>
          (
          <year>2014</year>
          )
          <article-title>D191{D198 doi:</article-title>
          10.1093/nar/gkt1140.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>