<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Talia library platform</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michele Nucci</string-name>
          <email>mik.nucci@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Hahn</string-name>
          <email>hahn@netseven.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Barbera</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Semedia Group - 3mediaLabs Università Politecnica delle Marche</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Talia is a web-based distributed digital library and publishing system, designed for scholarly research in philosophy. Talia is based on semantic web technology; it's being developed with the Ruby on Rails web framework. Rails is a relatively new environment, which allows developers to easily create well-structured web applications. By combining its power with semantic web technology and leveraging on existing solutions like ActiveRDF, it is possible to quickly create a full-featured semantic library platform, from scratch, in a short time. Talia does not aim at creating new semantic web technology as such, but at providing practical solutions for embedding the existing technology in modern web applications. This paper focuses on a few select features of Talia to show the possibilities.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Discovery3 is a European project to create an extensive online library for
philosophers. The four content partners of the project will provide tens of thousands
digitised and annotated pages, so that the project will start with a large body
of material.</p>
      <p>
        Philosophers usually work in a traditional, print-and-paper based way.
However, finding all relevant publications on a topic is difficult and acquiring copies
is yet another matter. This is especially true for original manuscripts, which can
be notoriously hard to acquire [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>By creating a comprehensive resource with original manuscripts and
secondary writings, we can provide an invaluable tool to the scholarly research
community.</p>
      <p>Discovery will be based on Philosource, a federation of interlinked online
libraries. The libraries currently use the Hyper platform. Hyper was created for
the preceding HyperNietzsche project (now NietzscheSource4).
3 http://www.discovery-project.eu/
4 http://www.nietzschesource.org/</p>
      <p>The Hyper platform was developed specifically for a Nietzsche library.
Particular assumptions about Nietzsche are hardcoded into the application, and the
codebase is not flexible enough to be a viable long-term solution for the project.
It will be superseded by the Talia platform described in this paper.</p>
      <p>In the humanities, the text itself is the subject of study. Talia aims to provide
an integrated library system that offers tools to work with original content. This
is in marked contrast to citation systems like CiteSeer5, CiteULike6 or even
online archives like arXiv7. These focus on bibliographical metadata, the actual
content is not more than an opaque, downloadable document.</p>
      <p>
        Talia provides a tight integration between a semantic backend store and
a powerful interface toolkit, making it unlike existing “infrastructure” library
systems, such as BRICKS [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or Fedora8. These provide a backend on which an
interface has to be built from scratch. The JeromeDL system [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is somewhat
similar to Talia, but aimed at a different audience.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Requirements and Technology</title>
      <p>Within the Discovery project, Talia does not exist in a vacuum – it’s a means
to an end. There are a number of features and components that must be
implemented in Talia to make the software useful within the project:
– Users expect that the User interface contains all features that are already
present on the Hyper platform.
– Metadata and ontology support is essential. Each group of scholars will
create their own domain ontology to order their material.
– A Remote Federation API has to be implemented to allow automatic
bidirectional references between documents in different libraries.
– Publication and Workflow. Scholars will be able to publish new results
online. Talia must offer a number of workflow models for peer-reviewed
publication.</p>
      <p>A speedy development is also essential, since the existing content has to be
migrated in the second year of the project and Talia must be running for the
project to succeed.</p>
      <p>The rapid development of Talia would not be possible without leveraging
existing technology, both from commercial web development and semantic web
research.</p>
      <p>
        – Ruby on Rails9 is a web development framework that has been picking
up a lot of pace recently. It uses the Model-View-Controller [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] pattern and
allows Developers to create full-featured web applications with a minimal
amount of code. It’s available under the MIT license [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
5 http://citeseer.ist.psu.edu/
6 http://www.citeulike.org/
7 http://arxiv.org/
8 http://www.fedora-commons.org/
9 http://www.rubyonrails.com/
– ActiveRDF10 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a Ruby toolkit that provides an object-RDF mapping
for a number of existing RDF triple stores.
– An RDF store is central to the application. During development the
Redland RDF11 engine is used, but Talia can work with any storage that is
supported by ActiveRDF. The development team will also provide a setup
for the Sesame12, store which supports inferencing.
– Of course Talia will also need “normal” web application features, like user
sign-on and permission management.
      </p>
      <p>As a general rule, Talia tries to avoid any unnecessary complexity. This is
also true for the use of RDF metadata:
– Talia is schema-agnostic. Ontology descriptions may be used by the system;
however the software does not rely on their presence. It also doesn’t attempt
to enforce any kind of schema rules on the RDF data.
– Talia does not attempt to do any inferencing on the RDF data. If this is
needed, it will be the responsibility of the RDF store.
– Talia tries not to use specialised RDF store features, in order to be
compatible with any RDF endpoint.
3</p>
    </sec>
    <sec id="sec-3">
      <title>RDF storage and querying</title>
      <p>Talia uses a hybrid RDBMS/RDF store solution in the storage backend. A
relational database as a highly reliable, transaction-aware storage for the critical
data. It is kept in sync with an RDF datastore for advanced semantic features,
which may range from SPARQL queries to inferencing, depending on the type of
the store. Using a standard RDF store also provides an easy way to interoperate
with semantic web software.</p>
      <p>The hybrid design is feasible because a Talia library has relatively static
content. Data access will mostly be read-only, the few modifications can be
easily synchronised without much overhead for the system.</p>
      <p>Talia provides a simple API that hides most of the internal workings of the
data store. Each document is represented as an object of type Source; the object
provides access provides access to all properties of the document, no matter if
defined as RDF or not. Listing 1.1 shows an example of this API.</p>
      <p>Listing 1.1. Basic Operations on a document
document = T a l i a C o r e : : Source . new ( " http : / / u r l . com/my_document" )
document . wo r k fl o w _s t a te = 2 # non r d f p r o p e r t y
document . dcns : : t i t l e &lt;&lt; "My␣ f i r s t ␣ document " # RDF p r o p e r t y
author . i n v e r s e . dcns : : c r e a t o r # " i n v e r s e "
# Replace t r i p l e
document . dcns : : t i t l e . r e p l a c e ( "My␣ f i r s t ␣ document " , "New␣name" )
document . dcns : : t i t l e . remove # remove t r i p l e s
10 http://www.activerdf.org/
11 http://librdf.org
12 http://www.openrdf.org/</p>
      <p>The interface borrows heavily from the ActiveRDF interface, and ActiveRDF
is used in the backend to connect to the RDF store. However, ActiveRDF was
designed mostly as an easy read interface for web applications. The library was
substantially refactored to improve the data manipulation capabilities (such as
deleting triples). Other modifications were made to make it easier to call the
library indirectly as part of a backend, instead of directly as a part of a script.
3.1</p>
      <sec id="sec-3-1">
        <title>Queries</title>
        <p>Talia provides an unified query interface for both database and RDF
metadata, as shown in listing 1.2. For normal queries, the interface hides most of
the complexity and automatically decides wether to use a RDF/SPARQL or an
RDBMS/SQL query on the backend.</p>
        <sec id="sec-3-1-1">
          <title>Listing 1.2. Querying for documents</title>
          <p># The f o l l o w i n g w i l l do a q u e r y on t h e RDF s t o r e
T a l i a C o r e : : S o u r c e . f i n d ( : a l l , N : : DCNS. C r e a t o r =&gt; " D a n i e l " )
# The f o l l o w i n g w i l l LIMIT a q u e r y t h a t u s e s RDF and DB d a t a
T a l i a C o r e : : S o u r c e . f i n d ( : a l l , N : : DCNS. C r e a t o r =&gt; " D a n i e l " ,
: w o r k f l o w _ s t a t e =&gt; 2 , : l i m i t =&gt; 5 )</p>
          <p>
            During development we found that the SPARQL [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ] query language is not
always best suited for web applications, where large result sets are usually broken
down into individual pages. This is usually done by using the LIMIT, OFFSET
and COUNT operators to retrieve the overall size and a subset of the
(possibly huge) result set. This method requires, however, that whole query can be
executed as a single statement.
          </p>
          <p>SPARQL does not provide an easy way express OR statements in a single
query (for example “subtype OR supertype”), and it’s filter mechanism is highly
inefficient for large result sets. A straightforward COUNT implementation is
also missing from the standard. Another problem is that in different RDF stores
the SPARQL specifications is implemented to various extents, making it difficult
to provide a store-agnostic solution.</p>
          <p>Early versions of Talia also encountered the problem of reconciling “mixed”
queries that use both the database and the RDF store. With the new hybrid
design it will be possible to answer each query either from the RDF data or the
database. This will allow the backend engine to select the query language best
suited for the job.</p>
          <p>The built-in query mechanism is optimised for compatibility with a number
of RDF stores and RDBMS. If this is too limiting the developer has the choice
issue queries (either SQL or SPARQL) directly and access store-specific features.
3.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Ontologies</title>
        <p>As mentioned, Talia’s core functionality does not rely on an ontology description.
Still, if one is needed, it can easily be loaded into the RDF store and queried
from Talia.</p>
        <p>Talia provides a SourceClass abstraction to represent RDF classes and to
navigate the ontology hierarchy, as shown in listing 1.3. The user may also
retrieve metainformation from the ontology, supertypes or subtypes of a class.</p>
        <sec id="sec-3-2-1">
          <title>Listing 1.3. Navigating the ontology</title>
          <p># Get t h e f i r s t r d f t y p e and g e t sub
f i r s t _ c l a s s = document . r d f _ t y p e s . f i r s t
sup = f i r s t _ c l a s s . s u p e r t y p e s
and s u p e r c l a s s e s
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Semantic UI templates</title>
      <p>Source objects can be used in HTML templates to create a web representation a
document. Talia uses Rails’ standard rhtml templates that contain HTML with
embedded ruby code. Listing 1.4 show a simple template that renders a HTML
snippet with some properties from the RDF store.</p>
      <sec id="sec-4-1">
        <title>Listing 1.4. Sample rendering template</title>
        <p>&lt;p&gt; The c u r r e n t document i s &lt;%= document . r d f s : : l a b e l . f i r s t %&gt;
i t ’ s authored by &lt;%= document . dcns : : c r e a t o r . j o i n ( " , ␣" ) %&gt; &lt;/p&gt;</p>
        <p>Talia needs a rich user interface for each document type. Unlike many
semantic web applications that use an “auto-generated” generic interface for semantic
metadata, Talia needs to provide specialised views depending on the RDF type
of a resource.</p>
        <p>A automatic semantic template engine is built into Talia to do just that.
Listing 1.5 shows a simple example; the site developer simply passes the Source
object to the source_snippet UI widget, and the semantic template engine does
the rest.</p>
        <p>Listing 1.5. Rendering a document with a semantic template
&lt;% my_document = T a l i a C o r e : : Source . new ( u r l ,</p>
        <p>N : :MYONT: : the_type ) %&gt;
&lt;%= widget ( : s o u r c e _ s n i p p e t , : s o u r c e =&gt; my_document ) %&gt;</p>
        <p>When the template engine renders a document, it will look at the document’s
RDF type and attempt to find a template which has a name that matches
the namespace and name of one of those types. In the example the document
has the type myont:the_type; if the template engine finds a template named
_myont_the_type.rhtml it will use it to render the document. Otherwise it will
fall back to a default template.</p>
        <p>The template engine allows the UI templates to be easily created by
professional web designers who don’t know semantic web concepts. In the final web
application, template selection happens automatically and it’s very easy to add
new templates. By providing a sensible default template, the engine is still able
to deal with elements of new and unknown types.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This paper showed a quick glimpse of Talia’s semantic web features and
demonstrated how semantic web development can be made a breeze by combining the
power of an existing framework with a dynamic RDF store API.</p>
      <p>More semantic web features will be included in future version, like semantic
links between remote libraries, user-created metadata and integration with the
DBin desktop application.</p>
      <p>Talia is freely available from its home page13, the page also contains some
instructions and additional documentation for the software. There’s also an online
demo of the current development version14.</p>
      <sec id="sec-5-1">
        <title>Acknowledgements</title>
        <p>This work has been supported by Discovery, an ECP 2005 CULT 038206 project
under the EC eContentplus programme.</p>
        <p>The authors wish to thank Eyal Oren and the ActiveRDF team for their
work and support.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>D</given-names>
            <surname>'Iorio</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.:</surname>
          </string-name>
          <article-title>Nietzsche on new paths: The hypernietzsche project and open scholarship on the web</article-title>
          . In Fornari,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Franzese</surname>
          </string-name>
          , S., eds.: Friedrich Nietzsche. Edizioni e interpretazioni.
          <source>Edizioni ETS</source>
          ,
          <string-name>
            <surname>Pisa</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Risse</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Kneˆzevic,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Meghini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Hecht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          :
          <article-title>The bricks infrastructure - an overview</article-title>
          . In: The International Conference EVA,
          <string-name>
            <surname>Moscoww</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kruk</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Woroniecki</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gzella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dabrowski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McDaniel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Anatomy of a social semantic library</article-title>
          .
          <source>In: European Semantic Web Conference. Volume Sematic Digital Library Tutorial</source>
          . (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Reenskaug</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>MVC Xerox Parc 1978-79</article-title>
          . http://heim.ifi.uio.no/~trygver/ themes/mvc/mvc-index.
          <source>html (1979 [accessed March</source>
          <year>2008</year>
          ])
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. MIT: MIT License. http://www.opensource.org/licenses/mit-license.
          <source>php ([accessed March</source>
          <year>2008</year>
          ])
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Oren</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Debru</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerke</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haller</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>ActiveRDF: Object-Oriented Semantic Web Programming</article-title>
          .
          <source>In: 16th International World Wide Web Conference (WWW2007)</source>
          , Banff, Alberta, Canada. (
          <volume>8</volume>
          -12 May,
          <year>2007</year>
          )
          <fpage>817</fpage>
          -
          <lpage>823</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. :
          <article-title>SPARQL Query Language for RDF</article-title>
          . http://www.w3.org/TR/rdf-sparql-query/ (
          <year>January 2008</year>
          [accessed March 2008])
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          13 http://trac.talia.
          <source>discovery-project.eu/</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>14 http://demo.talia.discovery-project.eu/ - note that this site may not always be online</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>