<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scienti c Lenses over Linked Data: An approach to support task speci c views of the data. A vision.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christian Brenninkmeijer</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chris Evelo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carole Goble</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alasdair J G Gray</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul Groth</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Steve Pettifer</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Stevens</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antony J Williams</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Egon L Willighagen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Bioinformatics - BiGCaT, Maastricht University</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, VU University of Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Royal Society of Chemistry</institution>
          ,
          <addr-line>ChemSpider</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>School of Computer Science, University of Manchester</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Within complex scienti c domains such as pharmacology, operational equivalence between two concepts is often context-, user- and task-speci c. Existing Linked Data integration procedures and equivalence services do not take the context and task of the user into account. We present a vision for enabling users to control the notion of operational equivalence by applying scienti c lenses over Linked Data. The scienti c lenses vary the links that are activated between the datasets which a ects the data returned to the user.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Developing an integrated view over pharmacological data is challenging due to
the complexity of the domain. For example, searches for the chemical
\Fluvastatin" on ChemSpider5 and DrugBank6 return di erent compounds: although
their basic chemical structure matches, the compounds di er in their
stereochemistry7. Another challenge is that the domain scientists have di ering opinions on
when data records can be related. For example, when searching for information
about \Protein Kinase C Alpha", the scientist may want information returned
for that protein as it exists in humans8, mice9 or both. Additionally, when
connecting across databases one may want to treat these proteins as equal in order
to bring back all possible information whereas in other cases (e.g. when the
researcher is focused on humans) only information on the particular protein should
be retrieved.</p>
      <p>
        Data integration in Linked Data relies on equality links between resources
across di erent datasets. Relevant existing approaches include Bio2RDF [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
5 http://www.chemspider.com/393587 accessed 7 Sept 2012.
6 http://www.drugbank.ca/drugs/DB01095 accessed 7 Sept 2012.
7 For details see http://www.chemconnector.com/2012/07/29/ accessed 7 Sept 2012.
8 http://www.uniprot.org/uniprot/P17252 accessed 7 Sept 2012
9 http://www.uniprot.org/uniprot/P20444 accessed 7 Sept 2012
      </p>
      <p>
        Chem2Bio2RDF [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and LODD [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. These present a single \ xed" view over
the data which is composed of the links between the data resources.
Applications are supported in discovering the links by identity mapping services e.g.
sameas.org [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Identi ers.org [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and BridgeDB [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. These identity mapping
services do not consider the \meaning" of the link, typically treating everything as
being owl:sameAs equivalent. The semantics of such links between datasets are
often not trivial: as Halpin et al. have shown sameAs, is not always sameAs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Operational equality is often context-speci c, especially in the complex domains
considered in science where opinion can depend on the interpretation applied
to results. Current linked pharmacology datasets and identity mapping services
do not consider the context, task and roles of users. However, it is important to
have the exibility of applying operational equality on a per task basis.
      </p>
      <p>We aim to support users in controlling and varying their view of the data
by applying a scienti c lens which govern the notions of equivalence applied to
the data. Users will be able to change their lens based on the task and role they
are performing rather than having one xed lens. To support this requirement,
we propose an approach that applies context dependent sets of equality links.
These links are stored in a stand-o fashion so that they are not intermingled
with the datasets. This allows for multiple, context-dependent, linksets that can
evolve without impact on the underlying datasets and support di ering opinions
on the relationships between data instances. This exibility is in contrast to
both Linked Data and traditional data integration approaches. We look at the
role personae can play in guiding the nature of relationships between the data
resources and the desired a ects of applying scienti c lenses over Linked Data.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Scienti c Personae</title>
      <p>
        We present two personae representative of the scienti c users of the data
integration system being developed in the Open PHACTS project10 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>Chris the Integrative Systems Biologist. Chris is a researcher in a university
who investigates processes in mice and compares these with processes in humans.
When searching for data he needs to be able to switch between species depending
on the kind of experimental (protein and mRNA) data he wants to compare in
pathways, i.e. sometimes he will be interested in just mice data, sometimes just
human, and at other times the combination of both. However, he always insists
that genes and proteins are kept distinct. When evaluating toxicological e ects
of Fluvastatin he is interested in both the pharmacological active stereoisomer
of Fluvastatin and inactive one, since the latter still might have side e ects. He
therefore allows the link between ChemSpider and DrugBank for Fluvastatin.</p>
      <p>Fiona the Bioinformatician. Fiona works for a pharmaceutical company. She
searches for potential drug-like compounds by mining existing literature for
similar compounds to interact with known targets. For compounds that have been
well studied, she would like to apply strict operational equivalence conditions, i.e.
10 http://www.openphacts.org/ accessed 7 Sept 2012.</p>
      <p>Scienti c Lenses over Linked Data
equating chemical compounds only if their stereochemistry matches and keeping
genes and proteins distinct, in order to work with the most relevant data.
However, for compounds that have not been so widely studied, she is willing to apply
more permissive notions of equivalence to increase the volume of data returned.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Scienti c Lenses and Linked Data</title>
      <p>Within scienti c datasets it is common to nd links to the \equivalent" record
in another dataset. However, there is no declaration of the form of the
relationship. There is a great deal of variation in the notion of equivalence implied by
the links both within a dataset's usage and particularly across datasets, which
degrades the quality of the data. The scienti c user personae have very di erent
needs about the notion of equivalence that should be applied between datasets.
The users need a simple mechanism by which they can change the operational
equivalence applied between datasets. We propose the use of scienti c lenses.</p>
      <p>
        A scienti c lens is de ned in terms of the scienti c notions upon which
operational equivalence can vary, e.g. stereochemistry, gene and protein equivalence,
and cross species matching. Our goal is that the scientists should be able to
focus on their science and not on the links between datasets. We do not propose
to de ne an ontology of relationships; Halpin et al [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] have already proposed a
similarity ontology which captures the nuances of equivalence at a logical level.
However, they also showed that it is di cult to accurately identify the
relationship predicate to use in a given situation. Instead, we propose to capture
the context in which the link holds, and speci cally a justi cation for the link,
e.g. stereochemistry. The metadata and the links are captured as a VoID linkset
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We have speci ed the predicates expected in the VoID header in order to
support scienti c lenses in the Open PHACTS platform [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We propose to
extend BridgeDB [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to interpret the metadata about a linkset and decide under
which scienti c lenses it is active. BridgeDB can then be used by applications
to discover operationally equivalent records under di erent scienti c lenses. It is
intended that a lens will specify all of the operational equivalence criteria that it
applies, i.e. there will be several lenses which match compounds based on their
stereochemistry but which vary in some other category of equivalence.
      </p>
      <p>Consider again the scienti c personae. Chris would have two lenses that he
could switch between: one that keeps species independent of each other and one
that matches proteins across species. Fiona would also have two lenses: one that
applies strict levels of operational equivalence and a more relaxed one. She could
of course de ne intermediary ones to allow her to slide the level of strictness.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>Equivalence of di erent, but for some purposes comparable content from, data
resources depends upon a user's opinion, task, and context. We have proposed
applying scienti c lenses which vary the notion of operational equivalence
between datasets. A scienti c lens varies the links that are active. Thus, a query
over the Linked Data would see a di erence in the data returned when applying
lenses with di erent notions of operational equivalence. Note that although we
are developing our lenses in the context of pharmacology, the same principles
could be applied in other domains. Our ongoing work is concentrating on two
strands. First we are developing mechanisms to de ne scienti c lenses in terms
of scienti c notions which can then be used to control which links are active.
Second we are developing a query infrastructure to support the contextualised
stand-o mappings used by our scienti c lenses approach.</p>
      <p>Acknowledgements
The research leading to these results has received support from the Innovative
Medicines Initiative Joint Undertaking under grant agreement number 115191,
resources of which are composed of nancial contribution from the European
Union's Seventh Framework Programme (FP7/2007- 2013) and EFPIA
companies' in kind contribution.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hausenblas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Describing linked datasets with the VoID vocabulary</article-title>
          . Note,
          <source>W3C (Mar</source>
          <year>2011</year>
          ), http://www.w3.org/TR/void/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Belleau</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nolin</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tourigny</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rigault</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morissette</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Bio2RDF: Towards a mashup to build bioinformatics knowledge systems</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <volume>41</volume>
          (
          <issue>5</issue>
          ),
          <volume>706</volume>
          {
          <fpage>716</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Brenninkmeijer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evelo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>A.J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waagmeester</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willighagen</surname>
          </string-name>
          , E.:
          <article-title>Dataset descriptions for the open pharmacological space</article-title>
          . Working draft,
          <string-name>
            <surname>Open</surname>
            <given-names>PHACTS</given-names>
          </string-name>
          (
          <year>August 2012</year>
          ), http://www.openphacts.org/specs/datadesc/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiao</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wild</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>Chem2Bio2RDF: a semantic framework for linking and data mining chemogenomic and systems chemical biology data</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <volume>255</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Correndo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salvadores</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Millard</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glaser</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shadbolt</surname>
          </string-name>
          , N.:
          <article-title>SPARQL query rewriting for implementing data integration over linked data</article-title>
          . In: EDBT/ICDT Workshops (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Halpin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCusker</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>H.S.:</given-names>
          </string-name>
          <article-title>When owl: sameas isn't the same: An analysis of identity in linked data</article-title>
          .
          <source>In: International Semantic Web Conference (1)</source>
          . pp.
          <volume>305</volume>
          {
          <issue>320</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. van Iersel,
          <string-name>
            <given-names>M.P.</given-names>
            ,
            <surname>Pico</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.R.</given-names>
            ,
            <surname>Kelder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Hanspers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Conklin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.R.</given-names>
            ,
            <surname>Evelo</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.T.A.</surname>
          </string-name>
          :
          <article-title>The BridgeDB framework: standardized access to gene, protein and metabolite identi er mapping services</article-title>
          .
          <source>BMC Bioinformatics 11</source>
          ,
          <issue>5</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Juty</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le Novere</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laibe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Identi ers.org and MIRIAM registry: community resources to provide persistent identi cation</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>40</volume>
          (
          <issue>D1</issue>
          ),
          <source>D580{D586</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Samwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jentzsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouton</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kallesoe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willighagen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajagos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prud</surname>
          </string-name>
          'hommeaux, E.,
          <string-name>
            <surname>Hassanzadeh</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pichler</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stephens</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Linked open drug data for pharmaceutical research and development</article-title>
          .
          <source>Journal of Cheminformatics</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <volume>19</volume>
          + (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harland</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chichester</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willighagen</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evelo</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomberg</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ecker</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mons</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Open</surname>
            <given-names>PHACTS</given-names>
          </string-name>
          :
          <article-title>Semantic interoperability for drug discovery</article-title>
          .
          <source>Drug Discovery Today</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>