<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Optique 1.0: Semantic Access to Big Data?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>E. Kharlamov</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Giese</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E. Jimenez-Ruiz</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. G. Skj veland</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Soylu</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Zheleznyakov</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>T. Bagosi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Console</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>P. Haase</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>I. Horrocks</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>S. Marciuska</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>C. Pinkel</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Rodriguez-Muro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Ruzzi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V. Santarelli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. F. Savo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K. Sengupta</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Schmidt</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E. Thorstensen</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Trame</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Waaler</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Free University of Bozen-Bolzano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Sapienza Universita di Roma</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Oxford</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>uid Operations AG</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Optique project aims at developing an end-to-end system for semantic data access to Big Data in industries such as Statoil ASA and Siemens AG. In our demonstration we present the rst version of the Optique system customised for the Norwegian Petroleum Directorate's FactPages, a publicly available dataset relevant for engineers at Statoil ASA. The system provides di erent options, including visual, to formulate queries over ontologies and to display query answers. Optique 1.0 o ers installation wizards that allow to extract ontologies from relational schemata, extract and de ne mappings connecting ontologies and schemata, and align and approximate ontologies. Moreover, the system o ers highly optimised techniques for query answering.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Accessing the relevant data in Big Data scenarios is increasingly di cult both
for end-user and IT-experts, due to the volume, variety, velocity, and complexity
dimensions of Big Data. This brings a high cost overhead in data access for large
enterprises. For instance, in the oil and gas industry, engineers spend 30{70% of
their time gathering and assessing the quality of data. The Optique project1 [
        <xref ref-type="bibr" rid="ref1 ref2">1,
2</xref>
        ] advocates for a next generation of the well known Ontology-Based Data Access
(OBDA) approach to address the data access problem. The project aims at
solutions that reduce the cost of data access dramatically. In our demonstration
we present the rst version of the Optique system which we customised for the
Norwegian Petroleum Directorate's (NPD) FactPages.2
      </p>
      <p>OBDA systems address the data access problem by presenting a general
ontology-based and end-user oriented query interface over heterogeneous data
sources. The core elements in a classical OBDA systems are an ontology, describing
? The research was supported by the FP7 grant Optique (n. 318338).
?? Corresponding author: evgeny.kharlamov@cs.ox.ac.uk
1 http://www.optique-project.eu/
2 http://factpages.npd.no
Application
Layer</p>
      <p>Visual Query
Formulation</p>
      <p>Reasoner
Data Layer</p>
      <p>SPARQL
Editor</p>
      <p>Visualisation</p>
      <p>System Interface</p>
      <p>Answer
Visualisation</p>
      <p>Ontology</p>
      <p>Visualisation
Triple Store</p>
      <p>Query
Answering
NPD</p>
      <p>NPD
FaFcatcPtaPgaegses</p>
      <p>Ontology and</p>
      <p>Mapping
Management</p>
      <p>Reasoner
Expert
users</p>
      <p>End
users</p>
      <p>Basic</p>
      <p>Import
metadata</p>
      <p>Installation Wizards</p>
      <p>Advanced</p>
      <p>Import onto.
vocabulary
&amp; metadata
Automatic extract: Semi-automat.</p>
      <p>ontology &amp; extract:
Direct Mappings R2RML Mapps</p>
      <p>Saturate ontology
from metadata
Add external ontology</p>
      <p>Load
ontology</p>
      <p>Align
ontology
Approximate
ontology
out
the application domain, and a set of mappings, relating the ontological terms
with the schemata of the underlying data sources. End-users formulate queries
using the ontological terms and thus they are not required to understand the
structure of the data sources. These queries are then automatically translated
using the ontology and mappings into an executable code over the data sources.</p>
      <p>State of the art OBDA systems, however, have shown among others the
following limitations:
{ The usability of OBDA systems is hampered by the need to use a formal
query language. Even if the users know the ontological vocabulary, they may
nd di cult to formulate queries with several concepts and relationships.
{ The prerequisites of OBDA, i.e., ontology and mappings, are in practice
expensive to obtain. Additionally, they are not static artefacts and should
evolve according to the new end-users' information requirements.
{ The e ciency of the translation process and the execution of the queries is
usually not su ciently addressed in OBDA systems.</p>
      <p>The rst version of the Optique system, i.e., Optique 1.0, aims at partially
overcoming the above limitations. Demonstration videos are available at following
address: http://www.cs.ox.ac.uk/isg/projects/Optique/demos/iswc2013/.
2</p>
    </sec>
    <sec id="sec-2">
      <title>System Overview</title>
      <p>A general three-layer architecture of the Optique system is depicted in
Figure 1 (Left ). The current version of the system o ers two main functionalities:
to query/visualise data and install/maintain the ontology and the mappings. At
the backend, the system also o ers an e cient query processing mechanism.</p>
      <p>Optique 1.0 allows to pose queries via a visual query formulation (VQF)
interface, a SPARQL editor, or from a query catalog. VQF exploits reasoning
in order to show both explicit and implicit domain knowledge to guide the
formulation of the query.</p>
      <p>Queries are executed by the Query Answering module based on Ontop system.3
Ontop provides functionalities for rewriting SPARQL queries using the system's
ontology and mappings, syntactic and semantic query optimisation, and query
unfolding. Thus, high e ciency of query answering is guaranteed. Rewritten and
unfolded queries are in SQL and they are executed over the NPD FactPages data,
which is stored in a relational database. The query answers are converted into
triples in order to con rm the format of the system's ontology, temporally stored
in the system's triple store, and displayed to the user in a tabular way or on
maps (using OpenStreetMap).</p>
      <p>The installation and maintenance of the ontology and the mappings is done via
the Ontology and Mapping Management component. Currently, this component
includes two installation wizards: basic and advanced. In Figure 1 (Right ) we
depict work ows of the wizards. The basic wizard exploits the relational database
metadata and automatically extracts an initial version of the ontology and direct
mappings4 to the ontology entities. The advanced wizard, unlike the basic one,
requires the user intervention and an ontology vocabulary as input in order to
(manually) create and edit R2RML mappings.5 Both the basic and advanced
wizards provide functionalities to align the bootstrapped ontology with a state of
the art domain ontology and approximate the resulting ontology if it is outside
the desired OWL 2 QL pro le.6 Alignment is performed using the ontology
matching system LogMap,7 which has shown to work well in practice and also
includes mapping repair facilities.</p>
      <p>Optique 1.0 is built on top of the Information Workbench8 (IWB), a generic
platform for semantic data management. The IWB provides a shared triple
store for managing the assets of Optique 1.0, such as, ontologies, mappings,
query logs, (excerpts of) query answers, database metadata, etc. The IWB also
provides generic interfaces and APIs for semantic data management, e.g., ontology
processing APIs. In addition to these backend data management capabilities, the
IWB provides a exible user interface which follows a semantic wiki approach,
based on a rich, extensible pool of widgets for visualisation, interaction, mashup,
and collaboration.</p>
      <p>
        Finally, Optique 1.0 is customised for the NPD FactPages, which is a public,
freely available dataset created to regulate and overlook the petroleum activities
on the Norwegian Continental Shelf (NCS) and contains information collected
from a wide range of activities on the NCS, e.g., operating companies, elds,
discoveries, facilities, pipelines, and seismic surveys|both historic and current
data. Its data has been converted and published as semantic web data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], of
which parts have been fed into the Optique 1.0 system.
3 http://ontop.inf.unibz.it/
4 http://www.w3.org/TR/rdb-direct-mapping/
5 http://www.w3.org/2001/sw/rdb2rdf/r2rml/
6 http://www.w3.org/TR/owl2-profiles/
7 http://code.google.com/p/logmap-matcher/
8 http://www.fluidops.com/information-workbench/
      </p>
    </sec>
    <sec id="sec-3">
      <title>Demonstration Details</title>
      <p>During the demonstration we will describe the NPD FactPages and present
functionalities of the Optique 1.0 system, with the focus on the following aspects:
query formulation and execution, and system installation. These aspects will
be illustrated on the NPD FactPages data. For the query formulation we will
stress our visual query formulation tool that currently supports construction of
tree-shaped conjunctive SPARQL queries. The demonstrated queries will be from
the oil industry domain. An example query is: \Find all elds that are operated
by 'Statoil Petroleum AS' and which have a facility that produces oil" ; it can be
seen in the screenshot of the VQF in Figure 2. We will run queries and present
results both in tables and maps, e.g., the location of \Fields" and \Oil facilities"
will be displayed on maps. Regarding the system's installation, we will present
both basic and advanced wizards and guide through their steps, that is, loading
metadata, extraction of an ontology and mappings, alignment with the domain
ontology, and approximation of the integrated ontology. We will also show how
to edit extracted direct mappings and de ne new R2RML mappings.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Giese</surname>
          </string-name>
          et al. \
          <article-title>Scalable End-user Access to Big Data"</article-title>
          . In: Big Data Computing. Ed. by
          <string-name>
            <given-names>R.</given-names>
            <surname>Akerkar</surname>
          </string-name>
          . Chapman and Hall/CRC,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>E.</given-names>
            <surname>Kharlamov</surname>
          </string-name>
          et al. \
          <article-title>Optique: Towards OBDA Systems for Industry"</article-title>
          .
          <source>In: ESWC postproceedings volume: Best Workshop Papers</source>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>M. G. Skj veland</surname>
            ,
            <given-names>E. H.</given-names>
          </string-name>
          <string-name>
            <surname>Lian</surname>
            ,
            <given-names>and I. Horrocks.</given-names>
          </string-name>
          \
          <article-title>Publishing the Norwegian Petroleum Directorate's FactPages as Semantic Web Data"</article-title>
          .
          <source>In: The Semantic Web { ISWC</source>
          <year>2013</year>
          . Ed. by
          <string-name>
            <given-names>H.</given-names>
            <surname>Alani</surname>
          </string-name>
          et al. Vol.
          <volume>8219</volume>
          .
          <string-name>
            <surname>LNCS</surname>
          </string-name>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>