<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Boosting RDF Adoption in Ruby with Goo</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Manuel Salvadores</string-name>
          <email>manuelso@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul R. Alexander</string-name>
          <email>palexander@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ray W. Fergerson</string-name>
          <email>ray.fergerson@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natalya F. Noy</string-name>
          <email>noy@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark A. Musen</string-name>
          <email>musen@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Stanford Center for Biomedical Informatics Research Stanford University</institution>
          ,
          <country country="US">US</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Why Goo? Why a Framework?</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>For the last year, the BioPortal team has been working on a new iteration that will incorporate major modi cations to the existing services and architecture. As part of this work, we transitioned BioPortal to an architecture where RDF is the main data model and where triple stores are the main database systems. We have a component (called \Goo") that interacts with RDF data using SPARQL, and provides a clean API to perform CRUD operations on RDF stores. Using RDF and SPARQL for a real-world large-scale application creates challenges in terms of both scalability and technology adoption. In BioPortal, Goo helped us overcome that barrier using the technology that developers were familiar with, an ORM-alike API.</p>
      </abstract>
      <kwd-group>
        <kwd>SPARQL</kwd>
        <kwd>RDF</kwd>
        <kwd>ORM</kwd>
        <kwd>Ontologies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>an abstraction over the physical structure of the data and the raw,
underlying SQL queries. They also help to map data between relational models and
object-oriented programming languages. Popular ORMs include Hibernate,
ActiveRecord and SQLAlchemy for Java, Ruby and Python respectively. Goo is
therefore an ORM speci cally designed to work with SPARQL backends. The
Goo library frees developers from thinking about the intricacies of SPARQL,
while still exposing the power of RDF's ability to interconnect data.</p>
      <p>The driving requirements for our design are the following:</p>
      <p>
        A number of libraries for di erent platforms o er ORM-like capabilities for
RDF and SPARQL. Jenabean uses Jena's exible RDF/OWL API to persist
Java Beans [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. But Jenabeans approach is driven by the Java object model
rather than an OWL or RDF schema. A number of tools use OWL schemas to
generate Java classes [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. These tools enable Model Driven Architecture
development, but do not provide support for triple stores. For Python, RDFAlchemy
provides an object-type API to access RDF data from triplestores. It supports
both SPARQL backends and Python RDFLib memory models [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. ActiveRDF
is a library for accessing RDF data from Ruby programs. It can be used as data
layer in Ruby-on-Rails, similar to ActiveRecord, and it provides an API to build
SPARQL queries programmatically [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The SPIRA project, also for Ruby,
provides a useful API for using information in RDF repositories that can be exposed
via the RDF.rb Ruby library [
        <xref ref-type="bibr" rid="ref3 ref4">4, 3</xref>
        ]. Because we use Ruby in the new BioPortal
platform, we considered SPIRA as our ORM. However, SPIRA's query strategy
was not built to handle very large collections of artifacts and the query API did
not allow for the complex query construction that BioPortal needs.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Goo's API in a Nutshell</title>
      <p>This section brie y introduces the Goo API. In this section, and the rest of the
paper, we describe the API using a subset of BioPortal models that are complex
enough to help us outline Goo's capabilities.
1 Model definition
class Person &lt; Goo: Base: Resource
model :person, namespace: :foaf, name_with: :name
attribute :name, enforce: [:unique]
attribute :birth_date, enforce: [:date_time], property: :birthday
attribute :accounts, inverse: [ on: :user_account, property: :person]
end</p>
      <p>Class definition that maps the
model to an RDF vocabulary.</p>
      <p>The following set of models is mentioned in the remainder of the paper: User
(describes a user pro le), Role (describes di erent roles of users in the
application, such as an administrator), Note (describes comments on ontologies provided
by users), Ontology (describes the object that represents an ontology entry in our
repository) and OWLClass. We do not provide a fully detailed schema for these
objects; each example is self-contained and the relations between objects should
be clear to the reader. The full documentation for the Goo API is available at
http://ncbo.github.io/goo/.</p>
      <p>Goo models are regular Ruby classes. To enable RDF support each model
needs to extend the Goo::Base::Resource class. Figure 1 gives a brief
description of how models get de ned. In the same gure it is shown how Ruby
developers can save and validate the object without having to deal with RDF and/or
SPARQL.</p>
      <p>Once objects are de ned, Goo provides a framework for creating, saving,
updating and deleting object instances. Goo assures uniqueness of RDF IDs in
collections and tracks modi ed attributes in persisted objects. The DSL allows
us to provide both validation rules and de ne how objects are interlinked.</p>
      <p>Ruby is a typeless language and thus developers can assign arbitrary value
types to variables and attributes (i.e: the language does not enforce the
assignment of Date values to a property that should only accept Date objects). We
rely on Goo to perform these operations automatically and transparently,
notifying when a validation fails. Goo incorporates multiple built-in validations for
data types like email, URI, integer, oat, date, etc. Moreover, the framework can
be extended by using Ruby lambdas which can be needed to perform custom
validations. For example, a custom ISBN format validation can be included in
the set of validations for a model with the following attribute de nition:
attribute :isbn, enforce: [ lambda { |self| isbn_valid?(self.isbn) } ]
2.1</p>
      <p>Querying: From Graphs of Triples to Graphs of Objects
Goo's most important feature is its exible query API. The API allows retrieving
individual objects, their attributes, collections, and so on.</p>
      <p>Retrieving individual objects One can retrieve single resource instances using
Resource.find. This call is useful when querying by unique attributes or when
the URI that identi es a resource is known.</p>
      <p>Retrieving object attributes By default, none of the query API calls attach any
attribute values to the instance object that they return. If we try to access an
attribute that has not been included (i.e., retrieved from the triplestore), Goo
throws an AttributeNotLoaded exception. Our design always defaults to
strategies that imply minimum data movements. This strategy improves e ciency by
retrieving only the attributes that the application cares about. Data attributes
are loaded into objects by using the include command (Figure 2).
Incremental Object Retrieval We have encountered situations in which one might
not know exactly what attributes need to be loaded in an object. Goo allows
incremental, in-place retrieval of attributes. An array of already loaded objects
can be populated with more attributes. This operation is in-place because Goo
will not create a new array of objects but will populate the objects that are
passed into the query via the models call.</p>
      <p>Users.where.models(users).include(notes: [:content])
Pagination Our REST API outputs large collections of data and in some cases
we have to implement pagination over the responses. Pagination in SPARQL,
with LIMIT and OFFSET, works at the triple level but it is not trivial for non
SPARQL experts to develop the queries that retrieve a paginated collection
of items. Goo provides capabilities that abstract the intricacies of triple level
pagination and leverages this capability to the Ruby objects. Every query in
Goo can be paginated, the underlying SPARQL query uses SPARQL LIMIT and
OFFSET to assure low data transfers and minimun object instantiation. A Goo
API paginated call is shown in Figure 2.4.
1 john = User.find("john")
notes = Note.where(owner: john) First get John and use that user</p>
      <p>object to match the Note graph.</p>
      <p>.include(:content)
2 notes = Note.where(owner: [ :username “john”])
.include(:content)</p>
      <p>Query directly the Note graph to
roewtrnieerveatetrviebruytenothteatthpaotinhtasstoana user
with username John
3 john = User.find("john")</p>
      <p>.include(notes: [:content])
notes = john.notes</p>
      <p>Access the Note graph through User.</p>
      <p>Find John and include his notes with
their content.
4 notes = Note.where(owner: [ :username “john”]) Same as (2) but using pagination
.include(notes: [:content])
.page(1,100)
notes.each do |note|</p>
      <p>puts note.content
end</p>
      <p>Just a plain Ruby loop over
John’s notes printing the
content.
Creating complex queries The API also allows for more complex query de
nitions. One can combine calls with or and join, and with these we internally
construct SPARQL UNION blocks that can be combined with SPARQL joins.
Range queries can be also de ned using the Filter object and the filter
method. All of these operations can be combined to build complex queries.
filter_on_created =
(Goo::Filter.new(:created) &gt; DateTime.parse('2011-01-01'))
.and(Goo::Filter.new(:created) &lt; DateTime.parse('2011-12-31'))
Users.where(notes: [ ontology: [ acronym: "SNOMEDCT"] ])
.or(notes: [ ontology: [ acronym: "NCIT"] ])
.join(notes: [ :visibility [ code: "public"]])
.filter(filter_on_created)
.include(:username, :affiliation)</p>
      <p>The Ruby code above represents a Goo query that retrieves the list of users,
with their :username and :affiliation, that submitted notes to the ontologies
"SNOMEDCT" or "NCIT" and these notes have visibility code "public". In
addition, a lter is created to lter users to just the ones that were created in the
system between a range of dates. This query has an extra complexity, the notes
attribute in User is de ned as an inverse attribute. Goo is able to reverse the
SPARQL query patterns to match the graph with the correct pattern
directionality. The ltering implementation also allows for retrieval of nonexistent graph
patterns. To retrieve the list of users that never submitted a note we simply use
the unbound call in Filter.1
1 See usage of Filter.unbound at http://ncbo.github.io/goo/</p>
    </sec>
    <sec id="sec-3">
      <title>Goo's Query Strategy</title>
      <p>Di erent query strategies can lead to starkly di erent performance in SPARQL.
A triplestore might have an e cient query implementation, but if our
application, for example, moves data around a lot, our queries will not perform well.
Indeed, one often hears complaints about the performance of SPARQL engines
whereas the real issue is the client who is not using SPARQL e ciently. Our
key rationale with implementing query strategies in Goo is to provide a layer
that ensures e cient access and query decomposition for SPARQL without the
developer having to worry about it.</p>
      <p>Goo's strategy navigates the graph of included patterns recursively. The rst
query focusses on constraining the graph and retrieving attributes that are
adjacent to the resource type. To retrieve data attributes located more than one
hop away, Goo runs additional queries|as many of these queries are types of
resources that are involved in the retrieval request. As a result, when a developer
uses Goo to request attributes of dependent resources, Goo will decompose the
request into multiple queries.</p>
      <p>Goo traverses the graph patterns recursively using a Depth First Search
(DFS); it focuses on individual resource types in each step. Figure 3 (right side)
shows how we chain these queries together with SPARQL FILTERs that join
sequences of OR operations. These lters help each subsequent query to retrieve
only attributes for the dependent models and not the entire collection.2</p>
      <p>Consider the following example. In BioPortal, a user can attach notes to
ontologies. A note description links to a user (author of the note) and the
ontology that the note refers to. Our testing dataset contains 400 ontologies, 10K
notes, 900 users and 10 roles. We have intentionally skewed our testing dataset
so that the top 10% ontologies account for 55% of the notes|the distribution
that re ects the actual state of a airs in BioPortal. Figure 3 highlights the
performance gain when retrieving notes for each ontology in BioPortal using Goo.
In the Naive approach performance degrades when we start projecting, into
tabular form, n-ary relations that are hidden in the graph. We have seen that it
is often the case that hand-written SPARQL queries retrieve too much data at
once, and this can cause combinatorial data explosions due to the hidden n-ary
relations.</p>
      <p>As it can be seen in Figure 3, we can demonstrate the di erence in the
performance of di erent query strategies even on this small dataset. Goo's strategy
will recursively navigate di erent types of objects in the retrieval, avoiding
combinatorial explosions.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>Traditional software development and database development has a long history
of reliable frameworks to access backend systems. The majority of Semantic Web
2 For this experiment we used 4store 1.1.5 in a cluster setup with 8 backend nodes.</p>
      <p>Naive
SELECT *
FROM :User FROM :Note FROM :Role
WHERE { ?id a :User .</p>
      <p>?id :username ?username .
?id :email ?email .
?note :owner ?id .
?note :content ?content .
?note :ontology ?ontology .
?ontology :acronym "$ONT_ACRONYM" .
?ontology :name ?name .
?id :roles ?roles . ?roles :code ?description }
3 SELECT ?role ?content</p>
      <p>FROM :Role
WHERE { ?role a :Role .
?role :code ?code .
} FILTER (?role = &lt;&gt; | ?role = &lt;&gt; ..)
4 SELECT ?note ?user ?content</p>
      <p>FROM : Ontology
WHERE { ?ont a :Ontology .
?ont :name ?name .
} FILTER (?ont = &lt;&gt; | ?ont = &lt;&gt; ..)
development today often includes writing SPARQL queries. In many cases,
software developers are unaware of what queries the SPARQL server is optimized
for and what queries they should use. Furthermore, many software-development
teams do not yet have SPARQL experts, which makes relying on triplestores as
components in large software systems problematic. Indeed, even in our team we
have experienced this problem as we were redesigning the BioPortal backend to
use a triplestore and not all of our developers were pro cient in writing e cient
SPARQL queries. The development of Goo enabled our team members to
overcome that barrier using the technology that they were familiar with (Ruby). Goo
is a completely general ORM for SPARQL and therefore other developers can
use it in their projects de ning their models that Goo will validate and relying
on the SPARQL query optimization in Goo to access their triplestores.</p>
      <p>In this paper we show one example of how using naive SPARQL can lead to
unexpected bad performance (see Figure 3). Our preliminary study shows that
the combination of n-ary relations in hand-written queries result in query time
distributions that may not perform well. Figure 3 shows that a simple partioning
strategy can help to alleviate this issue.</p>
      <p>We are proponents of semantic technologies, and thus we were drawn to the
idea of using ontologies to de ne our schemas, including the domains and the
allowed values for attributes. However, we see two major advantages to a DSL
like Goo. First, one of our goals was to make schema de nitions easy to use for
developers who are not familiar with ontologies or OWL, and using ontologies
counteracts that goal. Second, ontologies traditionally entail new information
and are not designed to validate constraints.</p>
      <p>Goo enables software developers without signi cant experience with semantic
technologies, to use SPARQL and RDF naturally and e ciently. The BioPortal
developers have found it easy to start working with the Goo API; as a result,
the transition to RDF and triple store technology within BioPortal was much
faster and smoother than it would have been otherwise. Basic query de nitions
in Goo are intuitive because the combination of hashes and arrays to construct
Goo query patterns resembles JSON structures and developers are very
familiar with them. Knowing that following a few simple restrictions, such as only
querying for attributes that are needed, freed them from having to worry about
the performance of the data storage layer and allowed them just to focus on the
business logic, which sped development time signi cantly.</p>
      <p>Acknowledgments This work was supported by the National Center for
Biomedical Ontology, under grant U54 HG004028 from the National Institutes of Health.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>OWL2Java</given-names>
            <surname>: A Java Code</surname>
          </string-name>
          <article-title>Generator for OWL</article-title>
          . http://www.incunabulum.de/ projects/it/owl2java
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kalyanpur</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimenez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Automatic Mapping of OWL Ontologies into Java</article-title>
          .
          <source>In: Proceedings of Software Engineering and Knowledge Engineering</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Arto</given-names>
            <surname>Bendiken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Gregg</given-names>
            <surname>Kellogg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.L.</given-names>
            ,
            <surname>Borkum</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          : RDF.
          <article-title>rb: Linked Data for Ruby</article-title>
          . http://rdf.rubyforge.org/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Ben</given-names>
            <surname>Lavender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.B.</given-names>
            ,
            <surname>Humfrey</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.J.:</surname>
          </string-name>
          <article-title>Spira: A Linked Data ORM for Ruby</article-title>
          . https: //github.com/datagraph/spira
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Graham</given-names>
            <surname>Higgins</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.C.</surname>
          </string-name>
          : RDF Alchemy, http://www.openvest.com/trac/wiki/ RDFAlchemy
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Oren</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delbru</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerke</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haller</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Activerdf:
          <article-title>Object-oriented Semantic Web Programming</article-title>
          .
          <source>In: Proceedings of the 16th International Conference on World Wide Web</source>
          . pp.
          <volume>817</volume>
          {
          <fpage>824</fpage>
          . WWW '07,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2007</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1242572.1242682
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Salvadores</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horridge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>P.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fergerson</surname>
            ,
            <given-names>R.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noy</surname>
            ,
            <given-names>N.F.</given-names>
          </string-name>
          :
          <article-title>Using SPARQL to Query BioPortal Ontologies and Metadata</article-title>
          .
          <source>In: International Semantic Web Conference (2)</source>
          . pp.
          <volume>180</volume>
          {
          <issue>195</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Vollel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Jenabean: A library for persisting java beans to RDF</article-title>
          . http://code. google.com/p/jenabean/
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Whetzel</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noy</surname>
            ,
            <given-names>N.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>P.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nyulas</surname>
            ,
            <given-names>C.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tudorache</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>BioPortal: Enhanced functionality via new web services from the national center for biomedical ontology to access and use ontologies in software applications</article-title>
          .
          <source>Nucleic Acids Research</source>
          (NAR)
          <volume>39</volume>
          (
          <issue>Web Server issue</issue>
          ),
          <source>W541{5</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>