<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The YummyData Initiative: How SPARQL-y is your biomedical endpoint?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ivar Andrea Splendiani</string-name>
          <email>andrea.splendiani@deri.org</email>
          <email>andrea.splendiani@intellilaf.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johan Nystrom-Persson</string-name>
          <email>johan@nibio.go.jp</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michel Dumontier</string-name>
          <email>dumontier@carleton.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yasunori Yamamoty</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carleton University</institution>
          ,
          <addr-line>Ottawa</addr-line>
          ,
          <country>Canada michel</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DERI</institution>
          ,
          <addr-line>Galway, IE</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Database Center for Life Science</institution>
          ,
          <addr-line>JP</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The National Institute of Biomedical Innovation</institution>
          ,
          <addr-line>Osaka</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>intelliLeaf ltd</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Although increasing amounts of biomedical data is being provided as structured content on the Semantic Web, there is currently no standardized way to monitor SPARQL endpoints for their availability, reliability or content ux. Importantly, there are additional issues relating to the provision of version-sensitive data republished by third parties or made available as part of a one o research project. All of these aspects have important consequences for users that rely on federated queries across distributed SPARQL endpoints.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>We describe the YummyData initiative to provide monitoring of biomedical
SPARQL endpoints on a variety of factors including availability, reliability,
content summarization, and content evolution. Our prototype website,
yummydata.org, provides simple metrics relating to the availability and response status
of a selected set of endpoints, the size of the data set, etc. In addition to these
fundamental metrics, our long term goal is to compute a SPARKLE score, a
composite metric combining measures such as the number of triples, size and
frequency of updates, the number of links to other datasets, and the capabilities
of the endpoint server. Although the SPARKLE score is not a measure of the
quality of a dataset, it helps indicate whether the dataset changes over time,
and whether these changes are likely to be negative or positive in nature. These
statistics may possibly correlated with the declared update frequency of the
datasets published by the endpoint of a given provider, thus providing an
additional input to our score. Finally, yummydata.org also supports custom queries
that track endpoint/data metrics, allowing new metrics to be computed and
shared amongst participants. Although devising an objective measure of quality
is controversial, we believe that the YummyData initiative will help users better
understand the content that is currently available while also helping providers
understand what kinds of metrics are important to users. We also believe that
emergence of more and more third party SPARQL endpoint rating services like
YummyData will bring about an environment where we can get more objective
evaluation of endpoints, and therefore qualities of the entire RDF-based
biomedical data/services are becoming higher. As yummydata.org is an early prototype,
we welcome suggestions to clarify existing metrics and to help develop additional
metrics to tease out information of interest.</p>
      <p>Acknowledgments
We wish to thanks the organizers of the BioHackathon 2012, during which this
work originated.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>