<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>OWLIM Reasoning over FactForge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Barry Bishop</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Atanas Kiryakov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zdravko Tashev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariana Damova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kiril Simov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ontotext AD</institution>
          ,
          <addr-line>135 Tsarigradsko Chaussee, Sofia 1784</addr-line>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present the reasoning mechanism in the OWLIM family of semantic repositories, which is based on materialization. This mech anism is evaluated using a combination of datasets from the Linked Open Data cloud in a public service called FactForge, where the benefits of materialization are manifested in improved SPARQL query performance. In this paper we present the reasoning mechanism employed in the OWLIM family of semantic repositories. These native RDF databases are implemented in Java and comprise storage components, inference-engine and query-answering engine. They are available in three editions: OWLIM-Lite, an in-memory and very fast RDF database that can load data at over 50,000 statements per second on a 1,000 USD machine using non-trivial inference; OWLIM-SE, that uses file-based, paged indices and data structures to be able to process tens of billions of RDF statements on standard desktop hardware; and OWLIM-Enterprise, a replication cluster based on OWLIM-SE that provides resilience and linearly scalable parallel query performance. OWLIM-Lite is free-for-use, whereas OWLIM-SE and OWLIM-Enterprise are the commercial editions licensed per CPU core. OWLIM-SE and OWLIM-Enterprise use a number of storage and query optimizations that allow it to sustain outstanding insert and delete performance even when managing tens of billions of statements of linked open data. The experiments conducted were performed using datasets from the Linked Open Data cloud, see section 3, that constitute a reason-able view [2] named FactForge1. Entities described in more than one dataset are unified via owl:SameAs statements and a common ontology PROTON2, called a unification ontology [1] for FactForge used for querying and data integration. PROTON is mapped to DBPedia3, FreeBase4 and Geonames5. Query performance with such dataset sizes is like-wise good, with sub-second response times for all the example queries found on the FactForge site.</p>
      </abstract>
      <kwd-group>
        <kwd>LOD</kwd>
        <kwd>materialization</kwd>
        <kwd>OWLIM</kwd>
        <kwd>RDF</kwd>
        <kwd>semantic repository</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>http://factforge.net/
http://www.ontotext.com/proton-ontology
http://dbpedia.org/About
http://www.freebase.com/
http://www.geonames.org/</p>
    </sec>
    <sec id="sec-2">
      <title>Reasoning in OWLIM</title>
      <p>The inferencing strategy in OWLIM is one of materialization based on
R-Entailment as defined by ter Horst [3], where Datalog like rules with inequality constraints
operate directly on a single ternary relation that represents all triples. In addition, free
variables in rule heads are treated as blank nodes. Materialization involves computing
all the entailed statements at load time. While this introduces additional reasoning
cost when loading statements into a repository, the desirable consequence is that
query evaluation can proceed extremely quickly. Several standard rule sets are
included in all editions of OWLIM and these include:
empty – no inference;
rdfs6 – RDFS semantics using rule entailment, but without data-type
reasoning, i.e. without the literal generalization and related rules;
owl-horst – equivalent to pD*, again without data-type reasoning;
owl-max – RDFS and OWL-Lite (that can be captured in rules);
owl2-ql – a fragment of OWL2 Full based on DL-LiteR, a variant of
DL</p>
      <p>Lite that does not require the unique name assumption;
owl2-rl7 – the OWL2 RL profile, a fragment of OWL2 Full amenable to
implementation on rule-engines, but without data-type reasoning.</p>
      <p>In addition to the standard semantics, user-defined rule-sets can be used. In this
case the user provides the full pathname to a custom rule file that contains definitions
of axiomatic triples, rules and consistency checks. For ease of use, the rule files for
the standard rule-sets are included in the distribution and users can modify or extend
these for their specific purposes.</p>
      <p>Consistency checks are used to ensure that the data model is in a consistent state
and are applied whenever an update transaction is committed, for example to ensure
that owl:Nothing has no members or that no pair of individuals have both
owl:sameAs and owl:differentFrom relationships.</p>
      <p>During loading, all inferred statements are materialized, except those generated as
a result of the semantics of owl:sameAs. OWLIM-SE uses special data structures to
maintain equivalence classes and uses the URI of the first asserted resource in each
equivalence class in the statement indices. This allows for the correct expansion of
results during query-answering while keeping the index sizes manageable. This
technique has the further advantage that it can be switched off during query answering in
order to limit the number of ‘duplicate’ results.</p>
    </sec>
    <sec id="sec-3">
      <title>FactForge - a Reason-able View on LOD</title>
      <p>FactForge is a reason-able view [] to the Web of Linked Data, made up of 11 of the
central LOD datasets, which have been selected and refined in order to serve as a use
ful index and entry point to the LOD cloud and to present a good use-case for
largescale reasoning and data integration. The compound dataset of FactForge is the largest
http://www.w3.org/TR/rdf-schema/
http://www.w3.org/TR/owl2-profiles/
3
body of heterogeneous general knowledge on which inference has been performed. It
counts 1.7 billion explicit statements; 15 billion retrievable statements available after
inference and owl:sameAs expansion (cf. section 2); including 1.4 billion inferred
statements. The datasets combined in FactForge are:
• DBPedia - an RDF dataset derived from Wikipedia, designed to provide as full
as possible coverage of the factual knowledge that can be extracted from the
InfoBoxes of Wikipedia with a high level of precision;
• Freebase - a dataset containing information about 11 million things, including
movies, books, locations, companies and more, with underlying schema based
on properties, and not ontologies, which exploits user generated categories;
• Geonames - a geographic database that covers 6 million of the most significant
geographical features on Earth, characterised by coordinates and relations to
other features (e.g. ‘parent feature’ in which the feature is nested);
• CIA World Factbook8 - a collection of structured data, including statistical,
geographic, political, and other information about all countries;
• Lingvoj9 - providing descriptions of the most popular human languages; cur
rently it contains information about more than 500 languages;
• MusicBrainz10 (RDF from Zitgist) – a comprehensive music information
suitable for browsing or useful for tagging;
• WordNet11 - a lexical database of English. Nouns, verbs, adjectives and
adverbs that are grouped into sets of cognitive synonyms (synsets).</p>
      <p>The interlinking of datasets is facilitated by DBPedia, which provides link-sets of
owl:sameAs links of DBpedia with GeoNames, Lingvoj, Freebase, MusicBrainz,
UMBEL and Wordnet. These link-sets are also loaded into FactForge along with the fol
lowing ontologies and schemata:
• DCMI Metadata Terms12 (Dublin Core - DC) - a relatively small, but very
popular metadata schema. It defines attributes that can be used to describe
information resources;
• SKOS13 (Simple Knowledge Organization System) - a relatively simple RDF
schema for describing taxonomies of concepts linked to each other by any sort
of subsumption hierarchy;
• RSS - an RDF schema designed to enable syndication of machine-readable
information about updates from Web sites;
• FOAF - an ontology for defining and linking personal profiles on the Web.</p>
      <p>FactForge provides several methods to explore the combined dataset that exploits
some of the advanced features of OWLIM-SE. Firstly, ‘RDF Search and Explore’
allows entities to be searched by keyword with a real-time auto-suggest feature ordered
by ‘RDF Rank’ (similar to Google’s Page Rank). The results page shows all triples
where the selected node appears as the subject, predicate or object, together with the
8 http://www4.wiwiss.fu-berlin.de/factbook/
9 http://lingvoj.org/
10 http://musicbrainz.org/
11 http://wordnet.princeton.edu/
12 http://dublincore.org/
13 http://www.w3.org/2004/02/skos/
preferred label, RDF Rank indicator, etc. Secondly, a SPARQL page allows users to
write their own queries with clickable options to add each of the known namespaces.
The results are presented in a conveniently formatted table with the option to down
load results in various formats (SPARQL/XML, JSON, etc). Lastly, a graphical
search facility called ‘RelFinder’ [4] that discovers paths between selected nodes.
This is a computationally intensive activity and the results are displayed and updated
dynamically during each iteration. The resulting graph can be reshaped by the user
with simple click and drag operations. Entities within the emerging graph can be
selected and a properties box provides links to the sources of information.
4</p>
    </sec>
    <sec id="sec-4">
      <title>PROTON - Unification Ontology for FactForge</title>
      <p>
        In addition to the above, FactForge also uses an ontology called PROTON
(developed by Ontotext) to unify concepts in the main datasets. The PROTON ontology
is a lightweight, upper-level ontology serving as a modelling basis for a number of
tasks in different domains. PROTON is meant to serve as a seed for ontology genera
tion, i.e. new ontologies constructed by extending PROTON. It can also be used for
automatic entity recognition and more generally Information Extraction (IE) from text
for the purpose of semantic annotation (metadata generation). The PROTON ontology
contains about 500 classes and 150 properties, providing coverage of the general con
cepts necessary for a wide range of tasks. The design principles can be summarized as
follows: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) domain-independence; (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) light-weight logical definitions; (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) alignment
with popular metadata standards; (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) good coverage of named entity types and
concrete domains, e.g. people, organizations, locations, numbers, etc.; and (5) good
coverage of instance data in Linked Open Data Reason-able views.
      </p>
      <p>
        The ontology is encoded in a fragment of OWL Lite and split into two modules:
Top and Extent. Top module is an upper ontology covering some basic philosophical
distinctions between entity types, such as: Object – existing entities (agents,
locations, vehicles); Happening – events and situations; Abstract – abstractions that
are neither objects nor happenings. The Top module also contains the main classes for
each of these types. The Extent module contains more domain and application
oriented classes. In FactForge PROTON is used to join the ontological classes and prop
erties of the main datasets. The mapping between PROTON and a given dataset onto
logy is done in three different ways: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) using rdfs:subClassOf statements
between classes in both ontologies and rdfs:subPropertyOf for properties; (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
using OWL expressions in the mappings where there is a difference in the
conceptualization in both ontologies; and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) using inference rules in cases where additional
individuals are necessary in the repository in order to support the mapping. Only the
PROTON ontology is loaded in FactForge. In this way the conceptual structure
implied by the particular dataset ontologies is ignored and only the PROTON definitions
are presented.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>The effects of the materialization and the owl:sameAs optimization described in
section 2 above result in 79% index compression, which means that close to 12 billion
triples that are not indexed are available for querying, making the total of retrievable
triples of FactForge close to 15 billion. The loading speed amounts to 8 032 explicit
indexed statements per second, and 14 630 indexed statements per second on a CPU
2 x Intel Xeon X5690, 3.46GHz, 12MB cache, 6 Core, RAM - 144 GB machine.</p>
      <p>The utility of reasoning becomes apparent during the evaluation of SPARQL
queries. For example the following SPARQL query about Mass media companies in
Europe which uses PROTON predicates only:</p>
      <p>PREFIX ptop: &lt;http://proton.semanticweb.org/protontop#&gt;
PREFIX pext: &lt;http://proton.semanticweb.org/protonext#&gt;
PREFIX dbpedia: &lt;http://dbpedia.org/resource/&gt;
?Company ptop:locatedIn ?Place ;</p>
      <p>pext:industryOf dbpedia:Mass_media .</p>
      <p>?Place ptop:subRegionOf dbpedia:Europe.</p>
      <p>returns answers indicating that “Associated Newspapers” is a media company
located not only in the United Kingdom, but also in England and in London based on the
materialization of the transitive relation ptop:subRegionOf.</p>
      <p>Furthermore, the queries with formulated with PROTON only return results much
faster than queries combining predicates and concepts from different LOD datasets in
FactForge, which is due to optimization of the joins traversed.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper we presented the inference mechanisms implemented in the OWLIM
semantic repositories and their application to a dataset formed by several LOD data
sets. The materialization of statements in the closure of the inference rules provides a
sound basis for extracting inferred information at query time.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Damova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiryakov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grinberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giasson</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simov</surname>
            ,
            <given-names>K..</given-names>
          </string-name>
          <article-title>Creation and Integration of Reference Ontologies for Efficient LOD Management</article-title>
          .
          <source>In: Semi-Automatic Ontology Development: Processes and Resources</source>
          , IGI Global, USA, Armando Stellato and Maria Teresa Pazienza (Eds.)
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kiryakov</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          ; Ognyanov,
          <string-name>
            <surname>D</surname>
          </string-name>
          ; Velkov,
          <string-name>
            <surname>R</surname>
          </string-name>
          ; Tashev,
          <string-name>
            <surname>Z</surname>
          </string-name>
          ; Peikov,
          <string-name>
            <surname>I;</surname>
          </string-name>
          <article-title>LDSR: a Reason-able View to the Web of Linked Data</article-title>
          ,
          <source>in: SW Challenge (ISWC2009)</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. ter Horst,
          <string-name>
            <surname>H. J. Combining RDF</surname>
          </string-name>
          and
          <article-title>Part of OWL with Rules: Semantics, Decidability, Complexity</article-title>
          .
          <source>In Proceedings of The Semantic Web ISWC</source>
          <year>2005</year>
          , LNCS volume
          <volume>3729</volume>
          pp.
          <fpage>668</fpage>
          -
          <lpage>684</lpage>
          . Springer Berlin / Heidelberg,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Heim</surname>
            ,
            <given-names>P</given-names>
          </string-name>
          ; Hellmann,
          <string-name>
            <surname>S</surname>
          </string-name>
          ; Lehmann,
          <string-name>
            <surname>J</surname>
          </string-name>
          ; Lohmann,
          <string-name>
            <surname>S</surname>
          </string-name>
          ; Stegemann,
          <string-name>
            <surname>T</surname>
          </string-name>
          ; (
          <year>2009</year>
          )
          <article-title>RelFinder: Revealing Relationships in RDF Knowledge Bases</article-title>
          .
          <source>In Semantic Multimedia</source>
          , volume
          <volume>5887</volume>
          <source>of LNCS</source>
          , pp.
          <fpage>182</fpage>
          -
          <lpage>187</lpage>
          . Springer Berlin/Heidelberg,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>