<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Lifting File Systems into the Linked Data Cloud with TripFS</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>University of Vienna, Department of Distributed and Multimedia Systems</string-name>
          <email>bernhard.schandl@univie.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Liebiggasse 4/3-4</institution>
          ,
          <addr-line>1010 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <volume>27</volume>
      <issue>2010</issue>
      <abstract>
        <p>A major fraction of digital information is stored in le systems. File systems organize les usually in labelled directory trees and provide a minimum support for user-driven le annotation, linkage and categorization. Although le systems play a major role in knowledge organization, both in enterprise contexts as well as in the personal information sphere, they have rarely been considered in Web-based information integration. To a large extent, this can be contributed to the limited metadata support of le systems and to the lack of stable identi ers for le and directories, which makes it hard to expose these objects in a global Web. We present TripFS, a lightweight approach for exposing parts of local lesystems as Linked Data. Serving le system objects via dereferenceable HTTP URIs paves the way to integrate them with the Web of Data, and enables new possibilities of exploiting le system data, for example, by linking them with other data sources or by annotating them using Semantic Web technologies.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>le systems</kwd>
        <kwd>le metadata</kwd>
        <kwd>information representation</kwd>
        <kwd>information integration</kwd>
        <kwd>event detection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>File systems store and organize data and documents of
all sorts and of arbitrary complexity, ranging from small
information snippets that can be put into single les, to
large repositories of heterogeneous content that are
organized within deep hierarchical structures. They act as the
storage backbone of many information processing systems
and can be considered as one major fundament of personal
and corporate information management. Since common le
systems do not impose major restrictions on creating,
naming, and arranging directories and les, they support a user's
individual preferences for data organization. File systems do
not only store les that were created or modi ed locally: a
large share of les originates from other sources, like
multimedia devices, other desktops, or the Web. In corporate
environments it is common to store data of collective
interest on shared le servers that enable a simple form of
collaboration.</p>
      <p>Overall, le systems can be considered as one of the
primary information sources both for organizations and
individuals, and it is quite likely that they will remain to be
important in the future. Therefore they are of high interest
for information integration. However, le systems have only
rarely been considered in the eld of Web-based data
integration. This stems mostly from their limited possibilities of
data organization1, limited metadata support, and the lack
of stable identi ers for les and directories.</p>
      <p>
        One promising strategy for Web-based information
integration is the Linked Data paradigm. This term denotes a
set of technologies and best practices that facilitate
information integration and linkage on a global scale. To expose
information as Linked Data means to follow simple principles
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]: rst, identify each resource of interest with a globally
unique, dereferenceable HTTP URI; second, provide useful
information for clients when they access the URI (usually
expressed in RDF and HTML); and third, include links to
other resources so that clients can retrieve more potentially
interesting information.
      </p>
      <p>In this paper, we present TripFS, a lightweight approach
that applies these principles for le systems in order to
expose their contents as Linked Data, and therefore enables
their direct inclusion in Web-based integration scenarios.
It assigns stable, globally valid, dereferenceable URIs to
les and directories, monitors changes in the system, serves
metadata extracted from les as RDF data, and interlinks
les with external data sources. It provides a plug-in
architecture so that it can easily be extended to support
additional le types and linking components, it adapts to the
speci cs of the underlying le system, and it provides a
sophisticated le change tracking component that increases
the stability of le identi ers.</p>
      <p>Because it is easy to set-up, TripFS also facilitates ad-hoc
sharing of le-based resources using standardized (semantic)
1By now it seems commonly accepted that a single
hierarchical scheme is insu cient for the organization of large
amounts of data as we encounter them on today's desktop
environments.</p>
      <p>Web technologies. Moreover, it overcomes shortcomings of
hierarchical organization mechanisms, because its
metadatacentric approach allows to query for descriptive information
instead of le location, and to establish multiple, orthogonal
views on le system data.</p>
      <p>After outlining application scenarios and describing how
users can bene t from exposing le systems as Linked Data
(Section 2), we discuss which steps have to be taken in order
to realize this idea (Section 3). We present details about the
TripFS architecture and implementation (Section 4). After
a discussion of related work (Section 5) we conclude the
paper in Section 6.</p>
    </sec>
    <sec id="sec-3">
      <title>BENEFITS OF LINKED FILE SYSTEMS</title>
      <p>
        The bene ts of exposing data as Linked Data resources
are manifold [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In this section we outline three scenarios
that illustrate how the quality of le system usage can be
increased by exposing les as Linked Data.
      </p>
      <p>
        A) Integrating File Systems into Enterprise Data. A
substantial fraction of enterprise data is available in the form
of le systems. While these data can be accessed in a
distributed context using protocols like CIFS or WebDAV, it
is di cult to integrate them in a global enterprise context
due to the lack of stable identi ers for les and
platformindependent metadata-based le access mechanisms. Linked
Data has been shown to be a viable approach for lightweight
enterprise information integration [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]; therefore, making le
systems part of a global or enterprise-internal Web of Data
enables them to be seamlessly integrated with, and
semantically connected to other data sources.
      </p>
      <p>B) Web-based Ad-hoc Data Sharing. Despite the vast
amount of possibilities for digital communication we have at
our disposal, ad-hoc sharing of meaningful information (e.g.,
the exchange of digital documents between participants'
laptops during face-to-face meetings) is still cumbersome. We
can regularly observe that collaborators use e-mail or
instant messaging to quickly exchange les. This approach,
however, does not allow more complex data to be shared,
or to exchange les together with metadata that describe
their correct context. Linked Data builds on top of common
Web technologies, thus any Linked Data source can be
directly accessed using a common Web browser. A tool that
allows users to temporarily share selected parts of their local
le systems as Linked Data (which implies not only sharing
plain les, but also extracted metadata, annotations, and
links) facilitates e cient information exchange amongst
collaborators.</p>
      <p>
        C) Semantic Web-based File Annotations. Semantic
annotation and interlinking of les is badly supported today:
although modern le systems support the storage,
management, and retrieval of le annotations (e.g., extended
attributes or le forks), these data are not accessible in a
standardized and platform-independent way. This makes
the organization of les into logically connected units di
cult, and reduces the e ciency of le retrieval especially in
distributed environments. If le systems were published as
part of a Web of Data, they could be annotated and
interlinked using tools like the LEMO annotation framework [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
or the Silk framework [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], which would lead to an increased
quality of search and retrieval, as well as linkage with other
relevant data sources. In turn, these Linked Data and
Webbased annotations could be propagated back into the
working context of the le system user, e.g., by being considered
by desktop search engines.
3.
      </p>
    </sec>
    <sec id="sec-4">
      <title>REPRESENTING FILE SYSTEMS AS</title>
    </sec>
    <sec id="sec-5">
      <title>LINKED DATA</title>
      <p>Since the characteristics of le systems and Linked Data
di er signi cantly, a number of steps have to be performed
in order to lift le system data into a Web of Data:
1. Appropriate representations for les and directories
have to be found, which comply to the Linked Data
principles.
2. Vocabularies that convey the characteristics of data
found in le systems have be to be speci ed and aligned
to already existing relevant vocabularies.
3. Descriptive metadata about les have to be extracted
from the le system and transformed into the RDF
data model.
4. Meaningful links to other, external data sources have
to be detected and established.
5. Consistency between the le system and its
corresponding Linked Data representation has to be ensured.
6. Data have to be served according to Linked Data
principles, i.e., in a form that is usable for both, humans
and machines.</p>
      <p>In the following we outline how each of these steps can be
realized.
3.1</p>
    </sec>
    <sec id="sec-6">
      <title>File URIs in the Web of Data?</title>
      <p>
        Within the context of a le system, les and directories
can be uniquely identi ed using their absolute paths, each
of which consists of a sequence of directory names and a le
name. The file: URI scheme [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is a means to directly
reuse these paths to form URIs, which can in turn be used
to access local le resources in a computer system.
      </p>
      <p>However, le URIs are neither globally unique, since they
describe a local path to a resource on a particular host, nor
stable, since the referenced les and directories may be
removed, moved, or renamed. Therefore they are not suitable
for being used in a global Web of Data.</p>
      <p>To solve this identi er problem, we chose to use opaque,
randomly generated UUIDs, and assign them to les and
directories. The usage of random UUIDs in a global
distributed context is assumed to be safe since the
probability of a collision is su ciently low. Further, since UUIDs
are fully opaque, they do not convey information about
the physical location of les and directories, and are
therefore stable even when the underlying le system objects are
changed. However, this requires to maintain a mapping
between stable, UUID-based URIs on the one hand, and
unstable, path-based identi ers on the other hand, to ensure
that modi cations in the le system are properly re ected
in the Linked Data representation. In Section 3.6 we outline
our strategy to accomplish this.</p>
    </sec>
    <sec id="sec-7">
      <title>3.2 Files and Directories as Web Resources</title>
      <p>The parent-child relationships between les and
directories can be represented as RDF triples with appropriate
predicates. Several triples are added to each le or
directory resource that convey data that are directly retrieved
from the le system: the local name (i.e., the actual le or
directory name without the entire path information), the
le size, and the dates of creation and last modi cation. An
example of a le's RDF representation is depicted in
Figure 1. Resources that represent les or directories are
internally identi ed by UUID-based URNs; for serving them as
Linked Data they are dynamically rewritten to HTTP URIs
(cf. Section 3.7).
3.3</p>
    </sec>
    <sec id="sec-8">
      <title>Vocabularies</title>
      <p>In order to describe les, directories, their metadata and
their relations as RDF, we have developed a simple OWL
vocabulary published at http://purl.org/tripfs/2010/02#.
We have derived our vocabulary from existing semantic
vocabularies as much as possible. However, as it is currently
uncommon to expose le resources as Linked Data, we
observed a lack of community-accepted vocabularies for this
purpose. To the best of our knowledge, only the
NEPOMUK File Ontology2 (NFO) has been speci cally de ned
to model the contents of le systems. It provides terms to
describe les, directories, and their properties. Our
vocabulary is aligned with NFO and provides more specialized
terms, according to our system's requirements.</p>
      <p>
        A number of other vocabularies, however, have a general
notion of the concept of documents, and usually align this
concept to the foaf:Document class. On the other hand,
several vocabularies have a notion of collections, which can
be compared to directories in a le system; for instance,
2http://www.semanticdesktop.org/ontologies/nfo/
OAI-ORE [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] or Dublin Core3. The Dublin Core Type
Vocabulary4, as another example, de nes terms for di erent
resource types as well as collections. Additionally, there
exists a large number of vocabularies that can be used to
identify media types and their speci cs; e.g., the MPEG-7
ontology5, the Music Ontology6, or the set of NEPOMUK
ontologies.
      </p>
      <p>To reach a maximum level of interoperability, a data source
should aim to adhere to commonly accepted vocabularies
as much as possible. The RDF semantics allows to
arbitrarily mix di erent, unrelated vocabularies; therefore we
propose|in addition to using a custom vocabulary|to
model le system data using the NFO vocabulary, and to add
type information from popular vocabularies like Dublin Core
and FOAF as they t. By serving data using multiple, even
already aligned vocabularies, we disburden data consumers
from the need to perform additional inference. An example
of such a mixed representation is presented in Figure 2.
3.4</p>
    </sec>
    <sec id="sec-9">
      <title>Extracting Semantic File Metadata</title>
      <p>Current le systems provide only a limited set of
lowlevel metadata attributes associated with les such as name,
owner, size, creation and modi cation date, or permission
attributes. Modern le systems provide additional means
to store higher-level metadata, like extended attributes or
multiple data streams; however these are only useful if they
are actually populated by applications, which is rarely the
case.
3http://dublincore.org/groups/collections/
collection-application-profile/
4http://dublincore.org/documents/
dcmi-type-vocabulary/
5http://metadata.net/mpeg7
6http://musicontology.com
As it is one of the Linked Data principles to \provide
useful information" about a resource when a client
dereferences its URI, it is desirable to extract additional,
descriptive metadata from les and directories and expose them
also as Linked Data. Reconsider, for example, Scenario A
described in Section 2, where the value of le-system level
metadata (like le size, le type, or le permissions) is
limited; higher-level descriptive metadata that can be used for
selective retrieval of les respectively their descriptions, e.g.,
via SPARQL, is required. However, the combination of these
metadata enables sophisticated discovery, retrieval and
access methods based on (i) the parent/child relations of le
system objects, (ii) low-level le system metadata, and (iii)
high-level content-based metadata.</p>
      <p>The problem of extracting metadata from le systems has
been studied for a long time. The biggest challenge in this
eld is the data diversity found in le systems, which is
imposed by the multitude of di erent le types. To illustrate
this, currently more than 51,000 le types are registered at
the popular FILExt service7. Di erent le types exhibit
different internal structures, and consequently di erent
metadata can be extracted. It is therefore impractical to provide
metadata extractors for this large amount of di erent le
types within a single software component. It is instead more
feasible to de ne a generic metadata extraction framework
that allows speci c extraction components for di erent le
types to be plugged-in. By this, the system can be tailored
to the respective application context.</p>
      <p>In our approach, extractors read les and extract an RDF
graph that contains triples representing the extracted
metadata. Multiple extractors can be cascaded into an extractor
pipeline and are sequentially applied to each object. The
resulting RDF graphs are stored in the triple store and are
7http://filext.com
then served as part of the le's and directory's description
via the Linked Data interface. Extractors may extract not
only le metadata (i.e., data about the documents
represented by les), but also entities that are related to les
(e.g., the artist who has performed the music stored in a
MP3 le) and can in turn be linked to external data sources.</p>
      <p>As an example, Figure 3 shows the RDF representation of
metadata that have been extracted from two les; the rst
resource represents a PDF document containing a scienti c
publication, the second represents an MP3 audio le8. The
blank node used to identify the artist in this example (line
9) needs to be dynamically rewritten to a stable,
dereferenceable URI by the Web server (see Section 3.7).
3.5</p>
    </sec>
    <sec id="sec-10">
      <title>Linking Files to External Sources</title>
      <p>Once les and directories are represented as RDF resources
it is possible to link them to other related resources on the
Web. Doing so allows clients to retrieve more, potentially
interesting information about the resource. For instance,
les may be classi ed according to a classi cation scheme
that uses dereferenceable URIs as identi ers; in this case,
clients are enabled to query for les using these terms.</p>
      <p>
        The task of linking les and directories to external
resources can be accomplished by tools that provide this
functionality for generic Web resources, which usually apply
various heuristics to detect semantically related resources (e.g.,
shared identi ers or object similarity [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]). These heuristics
depend on the information that is available for a particular
entity; therefore in the context of a le system they depend
on the data provided by metadata extraction components,
as described in the previous section.
8In this example we have used terms from the
OSCAF/NEPOMUK ontologies (http://www.
semanticdesktop.org/ontologies).
      </p>
      <p>As a consequence, we follow the same strategy as for
metadata extractors and do not provide an all-in-one solution to
the problem of linking les to external resources, but
instead provide a framework that allows specialized linking
components to be plugged in. These linking components
can access not only the raw le data, but also extracted
metadata, and use this information as basis for
interlinking. Like extractors, linking components return RDF triples
which are added to the metadata model and served via the
Linked Data interface.</p>
      <p>As an example, Figure 4 shows to which external sources a
scienti c publication and a music le can be linked, based on
string similarity between the publication title and the
combination of track title and artist name, respectively. In this
example the PDF document from Figure 3 has been linked
to the Linked Data variant of the popular DBLP publication
database, and the MP3 le has been linked to resources of
the MusicBrainz service.
3.6</p>
    </sec>
    <sec id="sec-11">
      <title>Maintaining Consistency</title>
      <p>As described in Section 3.1, it is required to mint a
UUIDbased URI for each le and directory, which can be
considered globally unique from a practical point of view.
However, without further precautions such URIs might be quite
unstable as the mapping between an UUID-based external
URI and a le-based internal URI is invalidated whenever
a referenced le is moved, removed, or renamed. Further,
updating such les may result in inconsistencies between a
le and the metadata that has been previously extracted
and stored. Note that this could lead also to invalid links
between resources if these were automatically created based
on le metadata, as described in Section 3.5.</p>
      <p>
        In order to preserve a stable mapping between these URIs
and the local les and directories they represent, we have to
employ a watcher component that is responsible for
detecting le system events that may result in di erent le URIs or
modi ed le contents of referenced les. Whenever such an
event is detected, appropriate actions have to be taken, and
the RDF model has to be updated. Note that in this sense,
the mapping between stable UUID-based URIs and instable
le and directory paths acts as a kind of translation
service between external, globally valid UUID-based URIs and
corresponding local le URIs, comparable to PURL or DOI
services [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Table 1 summarizes the reactions that have to
be taken after le system events have been detected.
3.7
      </p>
    </sec>
    <sec id="sec-12">
      <title>Serving File Systems as Web Resources</title>
      <p>
        Once the RDF-based representation of les and directories
has been generated and enriched with extracted metadata
and links to external data sources, the resulting RDF graph
can be served according to Linked Data principles. For
this purpose, internal UUID-based URNs are dynamically
rewritten to HTTP-based URIs with a con gurable host
part; e.g., http://example.com:8080/resource/&lt;uuid&gt;. It
is considered good practice [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to serve at least two variants
of the data, an RDF representation for machines and an
HTML representation for human consumption, and to let
clients choose which representation they prefer using HTTP
content negotiation. In addition to serving resources
according to Linked Data principles, it is recommended to provide
a SPARQL endpoint [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to allow clients to search for
resources based on their RDF descriptions. Furthermore, the
actual le data itself can be downloaded to the client. In
the special case where the Linked Data resources are
retrieved locally (i.e., server and client are executed on the
same machine), the Web server can add links to the HTML
interface that allow the user to directly open directories or
launch les from the browser, thus providing a seamless
interaction experience. Figure 5 shows a screenshot of such
an HTML-based interface, which provides these options to
the user.
      </p>
    </sec>
    <sec id="sec-13">
      <title>4. IMPLEMENTATION</title>
      <p>TripFS has been designed as a modular service framework,
which de nes plug-in interfaces that can be used to extend
and adapt the system to the actual needs of the use case,
the le types to be served, and the special characteristics
of the underlying operating system. Such interfaces exist
for RDF storage components, le metadata extractors, le
linkers, and le system crawlers (responsible for crawling
a con gured subtree of the le system) and watchers
(responsible for maintaining the consistency of the mapping
between external UUID-based URIs and internal le-based
URIs). The system's architecture is depicted in Figure 6.</p>
      <p>The TripFS core is a standalone server application, which
has been implemented in pure Java, based on the Jena
Semantic Web framework9. On startup, it crawls a con
gured sub-tree of the local le system, applies extractor and
linker components to crawled les, and stores the resulting
RDF triples in a triple store (either in memory or
persistent). It initializes the watcher component to monitor the
exposed le system sub-tree, which in turn noti es TripFS
upon changes to les or directories. Subsequently, the RDF
model is updated accordingly, and extractors and linkers are
re-applied to the modi ed objects.</p>
      <p>Metadata Extraction and Linking. We have
implemented simple extractors that extract low-level le
metadata, such as name, le size or a hash sum that could for
example be used to identify and link equal les across
different TripFS instances.</p>
      <p>Further, we have implemented extractor components based
on the Aperture metadata extraction framework10, which
provides a multitude of extractors for many di erent le
types, including O ce documents and multimedia data. As
a proof of concept, we have also implemented several linker
components: one that links documents, based on their
titles, to resources in the DBLP data set; one that links
audio les to MusicBrainz by analyzing track title and artist
9An evaluation version of TripFS can be obtained from
http://www.cs.univie.ac.at/tripfs.
10http://aperture.sourceforge.net
Path-based
navigation
Direct file access
Metadata access</p>
      <p>Link-based
navigation
Extracted
metadata
name, and one that links les to potentially interesting
DBpedia resources via the DBpedia lookup service. Both, the
set of extractors and linkers are to be understood as
proofof-concept; by far they do not leverage the full potential of
the presented approach. However, as described before, more
extractors and linkers can be integrated easily according to
the needs of an actual use case.</p>
      <p>
        Maintaining Consistency. We have used DSNotify [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]
as an implementation for the watcher component.
DSNotify is a change detection add-on for datasources, supporting
them in maintaining link integrity in their data. At its core,
DSNotify extracts feature vectors from considered data
entities that are used in heuristic comparisons to determine
whether items that are no longer found at their original
locations were in fact removed or moved to another location.
DSNotify can easily be extended by implementing custom
crawlers, feature extractors, and comparison heuristics.
      </p>
      <p>We have implemented a generic le-feature extractor for
DSNotify that extracts low-level features from local les (cf.
Table 2)11. Further, we have developed a simple
heuristic that calculates the plausibility that a le (described by
the feature vector X) was moved to another location (the
le there being described by the feature vector Y ). This
heuristic consists of two parts: rst, plausibility checks are
performed. For example, if the last modi cation date of le
Y is before the one of le X, it cannot be a successor of X.
Another example is that a le cannot become a directory
or vice versa (checked by the isDirectory feature). Second,
a similarity metric between the remaining features is
calculated by using the strategies listed in Table 2. The resulting
11The set of extracted features used by DSNotify is
overlapping but not equal to the set of metadata attributes
extracted and exposed by the TripFS. In the current
implementation, these latter metadata are stored in the RDF
graph while DSNotify stores features in its own indices.
Feature
Last access
Last modi cation
IsDirectory
Checksum
Name
Extension
Path
Size
Permissions</p>
      <p>Datatype</p>
      <p>Similarity</p>
      <p>Weight
Date
Date
Bool
Integer
String
String
String
Long
Bitstring</p>
      <p>
        Plausibility
Plausibility
Plausibility
Plausibility
Levensthein
Major MIME
type equality
Levensthein
Equality
Equality
3.0
1.0
0.5
0.1
0.1
similarities are weighted12 (e.g., the name similarity is
considered more important than equal le sizes), summed up,
and normalized. These similarities are then used by
DSNotify to detect move, remove and create events. Furthermore,
DSNotify reports update events based on changes in the
extracted feature vectors (cf. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]).
      </p>
      <p>DSNotify periodically monitors the le subtree that is
exposed by TripFS, extracts feature vectors based on the le
attributes described before, and stores these vectors in an
index. DSNotify uses a native C++ component for e
ciently monitoring the local lesystem that makes use of the
Windows API FindNextChangeNoti cation() method. We
12The selection of features as well as their weight was our
own subjective choice based on several test-runs with the
system. We consider an extensive evaluation of DSNotify as
a tool for detecting le system events as future work.
HTTP</p>
      <p>FILE/CIFS/SMB/NFS/...</p>
      <p>SPARQL Linked Data Interface</p>
      <p>Watcher</p>
      <p>/
Crawler
Extractors</p>
      <p>Linkers
TripFS</p>
      <p>Local Filesystem
have also implemented a generic, yet less e cient Java-based
monitor component that should work on all common
platforms. This allows us to re-crawl the respective subdirectory
tree only if there were actual changes reported by the
operating system. The detected events are then forwarded to
TripFS; the le's path is updated in the RDF model, and
extractors and linkers are re-applied.</p>
      <p>Linked Data Interface. TripFS includes an
embedded Jetty Web server, which serves data from the triple
store, as described in Section 3.7. It dynamically rewrites
the internally used UUIDs and blank nodes to
dereferenceable HTTP URIS, and provides XHTML+RDFa and pure
RDF representations of le and directory resources, as well
as a SPARQL endpoint. It further allows clients to directly
download le contents and, in the case of local requests, to
directly launch these les.</p>
      <p>Neither component of TripFS makes any changes to the
exposed le system; i.e., no special les or directories (like
needed e.g., for SVN) are created. Currently, TripFS also
does not provide means to modify le systems via the Linked
Data interface.</p>
    </sec>
    <sec id="sec-14">
      <title>RELATED WORK</title>
      <p>
        Although modern le systems support the creation,
storage, management, and retrieval of le-related metadata (e.g.,
using extended attributes or le forks), they remain mostly
isolated from Web-based information integration and
exchange contexts. Even le systems that provide
sophisticated support for le annotations or links (e.g., LiFS [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
or AttrFS [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]) do not consider a global Web context but
restrict their features often to objects within the local
system. On the other hand, Web-based le systems usually
focus on performance (e.g., [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]) or security (e.g., [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), but
not on semantically rich le descriptions or metadata
interoperability. In this respect, TripFS can be seen as
complementary to metadata-rich or highly scalable le systems
in order to bridge the gap between le systems and Web
environments. In combination with other works that
represent Web resources as virtual le systems (e.g., [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]), local
le systems and remote Web resources can be seamlessly
integrated, providing uni ed programming interfaces and a
consistent user experience.
      </p>
      <p>
        As described before, le system contents are highly diverse
and heterogeneous, and contain information that is valuable
in many scenarios. TripFS presents a generic framework to
RDF
expose these contents as Linked Data, but does not by itself
extract higher-level metadata from les. For this, it relies on
additional components, of which a wide variety exists. The
Aperture metadata extraction framework was already
mentioned before; it is based on the Gnowsis adapter framework
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and is capable of extracting RDF descriptions from a
wide range of les and other data sources. For most le types
there exist extractors that return RDF descriptions of the
le content, ranging from BibTeX les over calendar data
to JPEG images; a list of these extractors is maintained at
the W3C ESW Wiki13. Such conversion or extraction
components exist also for Web sources, e.g., PiggyBank [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] or
Virtuoso Sponger technology14, which create RDF
descriptions from a multitude of Web sources on the y.
      </p>
      <p>
        TripFS is in line with a number of other generic
frameworks that allow one to expose Linked Data based on a
different underlying data representation. Frameworks in this
area include D2R [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Triplify [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for relational data
bases, SparqPlug [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for DOM-based sources, OAI2LOD
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for OAI-PMH repositories, and XLWrap [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] for
spreadsheet data. With TripFS, le system contents can likewise
be made \ rst-class citizens" of the Web of Data and can
be seamlessly integrated with all these other data sources.
6.
      </p>
    </sec>
    <sec id="sec-15">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>In this paper we have presented and discussed TripFS, a
service that exposes local le systems according to Linked
Data principles. This approach potentially brings bene t to
a range of application scenarios (cf. Section 2). In an
enterprise information integration scenario (Scenario A), les
are assigned stable, globally unique URIs and can therefore
be referenced from external systems. Metadata that are
extracted from les can be indexed by Semantic Web search
engines, and links to other (enterprise-internal or external)
data sources can increase the quality of information
organization and data retrieval.</p>
      <p>A lightweight component like TripFS can also be used in
ad-hoc le sharing situations (Scenario B): participants in a
face-to-face meeting can easily set up and start the sharing
server, which exposes a certain sub-tree of their le system
as Linked Data. This enables collaborators in the same
network to access and retrieve these les, based not only on
lowlevel characteristics like le name, but also using extracted
semantic metadata and links. Using additional components,
more intuitive approaches like faceted navigation can be
performed on top of extracted data, and more experienced users
are enabled to issue complex SPARQL queries over the le
system.</p>
      <p>A Linked Data representation of le systems also
facilitates the application of Web-based annotation services
(Scenario C), which overcomes the limitations of the hierarchical
directory metaphor for le organization. Such annotations
can refer to single les or even parts thereof, and can range
from simple text-based comments to complex descriptions
that may refer to external entities and concepts. TripFS
makes le systems a part of a global, uniform Web of Data
and therefore allows one to apply Web-based annotation
techniques immediately to le system objects.</p>
      <p>In future work, we plan an extensive evaluation of TripFS,
13http://esw.w3.org/topic/ConverterToRdf
14http://docs.openlinksw.com/virtuoso/
virtuososponger.html
in particular regarding the performance and scalability of
our approach. For this purpose, we aim to apply TripFS in
a concrete enterprise information integration setting, and we
plan to develop a simple user interface that allows end users
to more easily share their les using Linked Data
technologies. Further, we plan to improve and evaluate the accuracy
of the DSNotify component for detecting le system events.</p>
      <p>Additionally, we plan to introduce a more ne-grained
model for selecting what le system objects are exposed via
TripFS (currently one can select only a single subtree of the
le system) and implement a secure HTTPS version that
takes privacy considerations into account.</p>
    </sec>
    <sec id="sec-16">
      <title>Acknowledgements</title>
      <p>Parts of this work have been funded by FIT-IT grants 812513
and 815133 from Austrian Federal Ministry of Transport,
Innovation, and Technology.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Sasha</given-names>
            <surname>Ames</surname>
          </string-name>
          , Nikhil Bobb,
          <string-name>
            <surname>Kevin M. Greenan</surname>
          </string-name>
          , Owen S. Hofmann, Mark W. Storer, Carlos Maltzahn, Ethan L.
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>and Scott A.</given-names>
          </string-name>
          <string-name>
            <surname>Brandt. LiFS</surname>
          </string-name>
          :
          <article-title>An Attribute-Rich File System for Storage Class Memories</article-title>
          .
          <source>In Proceedings of the 23rd IEEE / 14th NASA Goddard Conference on Mass Storage Systems and Technologies</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>William</surname>
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Arms</surname>
          </string-name>
          . Uniform Resource Names: Handles, PURLs, and
          <article-title>Digital Object Identi ers</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>44</volume>
          (
          <issue>5</issue>
          ):
          <fpage>68</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] Soren Auer, Sebastian Dietzold, Jens Lehmann,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Hellmann</surname>
          </string-name>
          , and David Aumueller. Triplify:
          <article-title>Light-weight Linked Data Publication from Relational Databases</article-title>
          .
          <source>In WWW '09: Proceedings of the 18th international conference on World wide web</source>
          , pages
          <volume>621</volume>
          {
          <fpage>630</fpage>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Arati</given-names>
            <surname>Baliga</surname>
          </string-name>
          , Joe Kilian, and
          <string-name>
            <given-names>Liviu</given-names>
            <surname>Iftode</surname>
          </string-name>
          .
          <article-title>A Web-based Covert File System</article-title>
          .
          <source>In Proceedings of the 11th Workshop on Hot Topics in Operating Systems</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Masinter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>McCahill. Uniform Resource</surname>
          </string-name>
          <article-title>Locators (URL) (RFC 1738)</article-title>
          . Network Working Group,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Tim</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <source>Linked Data. World Wide Web Consortium</source>
          ,
          <year>2006</year>
          . Available at http://www.w3.org/DesignIssues/LinkedData.html, retrieved 08-Aug-
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Chris</given-names>
            <surname>Bizer</surname>
          </string-name>
          , Richard Cyganiak, and Tom Heath. How to Publish
          <source>Linked Data on the Web</source>
          ,
          <year>2007</year>
          . Available at http://www4.wiwiss.fu-berlin.de/bizer/pub/ LinkedDataTutorial/, retrieved 02-Dec-
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Chris</given-names>
            <surname>Bizer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andy</given-names>
            <surname>Seaborne</surname>
          </string-name>
          .
          <article-title>D2RQ - Treating Non-RDF Databases as Virtual RDF Graphs</article-title>
          .
          <source>In Poster at the 3rd International Semantic Web Conference (ISWC2004)</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          , Tom Heath, and
          <string-name>
            <surname>Tim</surname>
          </string-name>
          Berners-Lee.
          <article-title>Linked Data | The Story So Far</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Kendall</given-names>
            <surname>Grant</surname>
          </string-name>
          <string-name>
            <surname>Clark</surname>
          </string-name>
          ,
          <article-title>Lee Feigenbaum, and Elias Torres. SPARQL Protocol for RDF (W3C Recommendation 15 January</article-title>
          <year>2008</year>
          ).
          <source>World Wide Web Consortium</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Coetzee</surname>
          </string-name>
          , Tom Heath, and Enrico Motta.
          <article-title>SparqPlug: Generationg Linked Data from Legacy HTML, SPARQL and the DOM</article-title>
          .
          <source>In Proceedings of the First International Workshop on Linked Data on the Web (LDOW)</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Sanjay</surname>
            <given-names>Ghemawat</given-names>
          </string-name>
          , Howard Gobio , and
          <string-name>
            <surname>Shun-Tak Leung</surname>
          </string-name>
          .
          <article-title>The Google File System</article-title>
          .
          <source>In 19th ACM Symposium on Operating Systems Principles</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Bernhard</surname>
            <given-names>Haslhofer</given-names>
          </string-name>
          , Wolfgang Jochum,
          <string-name>
            <surname>Ross King</surname>
            ,
            <given-names>Christian</given-names>
          </string-name>
          <string-name>
            <surname>Sadilek</surname>
            , and
            <given-names>Karin</given-names>
          </string-name>
          <string-name>
            <surname>Schellner</surname>
          </string-name>
          .
          <article-title>The LEMO Annotation Framework: Weaving Multimedia Annotations with the Web</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          ,
          <volume>10</volume>
          (
          <issue>1</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Haslhofer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Schandl</surname>
          </string-name>
          .
          <article-title>The OAI2LOD Server: Exposing OAI-PMH Metadata as Linked Data</article-title>
          .
          <source>In International Workshop on Linked Data on the Web (LDOW2008)</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>David</given-names>
            <surname>Huynh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Mazzocchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and David R.</given-names>
            <surname>Karger</surname>
          </string-name>
          .
          <article-title>Piggy Bank: Experience the Semantic Web Inside Your Web Browser</article-title>
          . In International Semantic Web Conference, volume
          <volume>3729</volume>
          of Lecture Notes in Computer Science, pages
          <volume>413</volume>
          {
          <fpage>430</fpage>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Georgi</surname>
            <given-names>Kobilarov</given-names>
          </string-name>
          , Tom Scott, Yves Raimond, Silver Oliver, Chris Sizemore,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Smethurst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Robert</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Media Meets Semantic Web | How the BBC Uses DBpedia and Linked Data to Make Connections</article-title>
          .
          <source>In Proceedings of the 6th European Semantic Web Conference</source>
          , pages
          <volume>723</volume>
          {
          <fpage>737</fpage>
          , Berlin, Heidelberg,
          <year>2009</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Carl</given-names>
            <surname>Lagoze and Herbert Van de Sompel. ORE</surname>
          </string-name>
          <article-title>Speci cation | Abstract Data Model</article-title>
          . Open Archives Initiative,
          <year>2008</year>
          . Available at http://www.openarchives.
          <source>org/ore/1</source>
          .0/datamodel.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Langegger</surname>
          </string-name>
          and
          <article-title>Wolfram Wo . XLWrap - Querying and Integrating Arbitrary Spreadsheets with SPARQL</article-title>
          .
          <source>In International Semantic Web Conference</source>
          . Springer,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Niko</given-names>
            <surname>Popitsch</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Haslhofer</surname>
          </string-name>
          .
          <article-title>DSNotify: Handling Broken Links in the Web of Data</article-title>
          .
          <source>In 19th International WWW Conference (WWW2010)</source>
          , Raleigh,
          <string-name>
            <surname>NC</surname>
          </string-name>
          , USA,
          <fpage>2</fpage>
          <lpage>2010</lpage>
          . ACM. to be published.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Leo</given-names>
            <surname>Sauermann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sven</given-names>
            <surname>Schwarz</surname>
          </string-name>
          .
          <article-title>Gnowsis Adapter Framework: Treating Structured Data Sources as Virtual RDF Graphs</article-title>
          .
          <source>In Proceedings of the 4th International Semantic Web Conference (ISWC</source>
          <year>2005</year>
          ), pages
          <fpage>1016</fpage>
          {
          <fpage>1028</fpage>
          .
          <string-name>
            <surname>Springer-Verlag</surname>
            <given-names>GmbH</given-names>
          </string-name>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Schandl</surname>
          </string-name>
          .
          <article-title>Representing Linked Data as Virtual File Systems</article-title>
          .
          <source>In Proceedings of the 2nd International Workshop on Linked Data on the Web (LDOW)</source>
          , Madrid, Spain,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Julius</surname>
            <given-names>Volz</given-names>
          </string-name>
          , Christian Bizer,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Gaedke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Georgi</given-names>
            <surname>Kobilarov</surname>
          </string-name>
          .
          <article-title>Discovering and Maintaining Links on the Web of Data</article-title>
          .
          <source>In Proceedings of the 8th International Semantic Web Conference (ISWC</source>
          <year>2009</year>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>C.E.</given-names>
            <surname>Wills</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giampaolo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.S.</given-names>
            <surname>Mackovitch</surname>
          </string-name>
          .
          <article-title>Experience with an Interactive Attribute-based User Information Environment</article-title>
          .
          <source>In Computers and Communications</source>
          ,
          <year>1995</year>
          .
          <source>Conference Proceedings of the 1995 IEEE Fourteenth Annual International Phoenix Conference on</source>
          , pages
          <volume>359</volume>
          {
          <fpage>365</fpage>
          ,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>