<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Pay-as-you-go Methodology for Ontology and Mapping Engineering in Ontology Based Data Access</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Juan F. Sequeda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel P. Miranker</string-name>
          <email>mirankerg@capsenta.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Capsenta Inc</institution>
        </aff>
      </contrib-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The Ontology Based Data Access (OBDA) paradigm enables (1) a clean separation
between a conceptual business view from heterogeneous data sources and (2) ability to ask
questions in terms of the business view, independent of how and where the data is
physically stored. Two key components are ontologies and mappings. Capsenta is applying
the OBDA paradigm for Business Intelligence applications. Even though OBDA has
been widely researched theoretically, there is still need to understand how to effectively
implement OBDA systems in the real world. From an practical point of view, this begs
the question: where does the ontology and the mapping come from? We present our
ongoing work of a pay-as-you-go methodology for ontology and mapping engineering
focused on the Business Intelligence (BI) questions that need to be answered.</p>
      <p>Consider the following real-world Business Intelligence example: Executives of a
large e-commerce company need to know how many orders were placed in a given
month and the corresponding net sales. Depending on whom they ask they get different
answers. The IT department managing the website records an order when a customer
has checked out. The fulfillment department records an order when it has shipped. Yet
the accounting department records an order when the funds charged against the credit
card are actually transferred to the company’s bank account, regardless of the shipping
status. Unaware of the source of the problem, the executives have inconsistencies across
their business reports.</p>
      <p>This is precisely where the use of ontologies to be the bridge between IT
developers and business users is valuable. Ontologies serve as a uniform conceptual federated
model describing the domain of interest. We are experiencing an increase of Ontology
Based Data Access (OBDA) systems being deployed in industrial applications. In the
OBDA paradigm, the ontology provides a logical abstraction, independent of how and
where the data is physically stored. The ontology serves as a business view, using
business terminology, which is then connected to data sources. Thus, providing a foundation
for comfortable communication between business users and IT developers.</p>
      <p>
        The common definition of OBDA states that given a source relational database, a
target ontology and a mapping from the relational database to the ontology, the goal is to
answer queries over the target ontology using these three components. From a practical
point of view, this begs the question: where does the target ontology and the mappings
come from?
Ontology Challenges Ontology engineering is a challenge by itself. In order to
create the target ontology, users can follow traditional ontology engineering
methodologies [
        <xref ref-type="bibr" rid="ref2 ref8">2, 8</xref>
        ], using competency questions [
        <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
        ], test driven development [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], ontology
design patterns [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], etc. Additionally, per standard practices, it is recommended to reuse
and extend existing ontologies in domains of interest such as Good Relations1 for
ecommerce, FIBO2 for finance, Gist3 for general business concepts, Schema.org 4, etc.
In OBDA, the challenge increases because the source database schemas can be
considered as additional inputs to the ontology engineering process. Common enterprise
application’s database schema commonly consist of thousands of tables and tens of
thousands of attributes. A common approach is to bootstrap ontologies derived from
the source database schemas, known also as putative ontologies[
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. The putative
ontologies can gradually be transformed into target ontologies, using existing ontology
engineering methodologies.
      </p>
      <p>Mapping Challenges Once the Target ontology has been created, the source databases
can be mapped. The W3C Direct Mapping5 standard can be used to bootstrap
mappings . The declarative nature of W3C R2RML6 mapping language enables users to
state which elements from the source database are connected to the target ontology,
instead of writing procedural code. Given that source database schemas are very large, the
OBDA mapping challenge is suggestive of an ontology matching problem: the putative
ontology of the source database and the target ontology. In addition to 1-1
correspondences between classes and properties, mappings can be complex involving calculations
and rules that are part of business logic. For example, the notion of net sales of an order
is defined as gross sales minus taxes, discounts given, etc. The discount can be
different depending on the type of user. Therefore, a business user needs to provide these
definitions before hand. That is why it is hard to automate this process.</p>
      <p>Addressing these challenges is crucial for the success of OBDA in practice.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Pay-as-you-go Methodology for OBDA</title>
      <p>We present our on-going work of a methodology to create the target ontology and
mappings for an OBDA system, driven by a prioritized list of business questions. The
objective is to create a target ontology and mappings, that enable answers to list of business
questions, in an incremental manner. After a minimal set of business questions have
been successfully modeled, mapped, answered and made into dashboards, then the set
of business questions can be extended. The new questions, in turn, may extend the
target ontology and new mappings incrementally added. With this methodology, the target
ontology and mappings are developed in an iterative pay-as-you-go approach. The
result is an agile methodology for BI using the OBDA paradigm because the focus is to
provide early and continuous delivery of answers to the business users.</p>
      <p>We identify three actors involved throughout the process: 1) Business user: subject
matter expert who has knowledge of the business and can identify the list of prioritized
business questions, 2) IT developer: has knowledge of databases and knows how the
1 http://www.heppnetz.de/projects/goodrelations/
2 https://spec.edmcouncil.org/fibo/
3 https://semanticarts.com/gist/
4 http://schema.org/
5 https://www.w3.org/TR/rdb-direct-mapping/
6 https://www.w3.org/TR/r2rml/
data are interconnected and 3) Knowledge engineer: communication bridge between
business users and IT developers, and has expertise in modeling data using ontologies.</p>
      <p>Our methodology is divided into two phases: knowledge capture and
implementation. Figure 1 provides an overview of the methodology.</p>
      <p>Knowledge Capture: Discovery-Vocabulary-Ontology The goal of the knowledge
capture phase is first to extract key concepts and relationships from the set of
prioritized business questions and second to identify which source database(s) contains data
relating to the extracted concepts and relationships. The steps are: 1) Discovery:
knowledge engineer works with business users to discover the concepts and relationships from
the input set of prioritized business questions in order to eliminate ambiguity.
Furthermore, the knowledge engineer takes what has been extracted with the business users and
works with IT developers to identify which tables and attributes from the database(s)
are required. 2) Vocabulary: knowledge engineer works with business users to identify
the business terminology such as preferred labels, alternative labels, and natural
language definitions for the concepts and relationships. 3) Ontology: knowledge engineer
formalizes the ontology in OWL such that it covers the business questions.
Implementation: Mapping-Query-Validation The goal of the implementation phase
is to enable answering the business questions by connecting the ontology with the data.
The steps are: 1) Mapping: knowledge engineer takes what was learned from the
Discovery and Ontology steps and implements the mapping in R2RML. The mapping is
then used to setup the OBDA system. 2) Query: knowledge engineer implements the
business questions as SPARQL queries. 3) Validation: knowledge engineer confirms
with the business users that the SPARQL queries return the correct answers.</p>
      <p>However, this leads to another question: where do the business questions come
from? In practice, we observe that is is often the case that business questions are
currently being answered by a small set of expert users by running multiple SQL queries to
manually generate BI reports. The problem is that the answers to these questions
usually take a long time to be generated and they are not always trusted by the executives.
Consider the following common scenario: Business users asks IT developers to answer
a business question. SQL queries are initially created by IT developers who are
knowledgeable of the large database schema. Developers come and go within an organization.
Queries get shared, altered, extended and combined. After time, business users are
executing SQL queries without any understanding of what the queries actually do. Users
rely on a description of what the SQL query is supposed to be returning.</p>
      <p>Our hypothesis is that we should be able to extract valuable information from SQL
queries which are being used by small amount of expert users to manually create BI
reports. Specifically, by valuable information we mean, the possibility to generate an
Ontology and Mapping from a query. This Ontology and Mapping is the starting point
to implement an OBDA system for BI.</p>
    </sec>
    <sec id="sec-3">
      <title>3 Conclusion</title>
      <p>Based on Capsenta’s real world experiences of deploying ODBA systems, the main
challenges that we encounter is the engineering of ontologies and mappings. In this
poster we present our ongoing work towards tackling this challenge. Our hypothesis
is that ontologies and mappings for OBDA can be generated from business questions
through a pay-as-you-go methodology. Our focus is on business questions coming from
existing SQL queries used to manually generate existing BI reports.</p>
      <p>To the best of our knowledge, the engineering of ontology and mappings for OBDA
is still open grounds for research. There are several challenges going forward, such as:
Automation: Given a SQL query, how can we automatically generate an OWL
ontology and R2RML mappings? Iteration: Manage new business questions that extend the
ontology and mappings. What happens if a new query contradicts the current ontology
and/or mappings, hence it is non-monotonic? Tools: There is a need for tools that can
manage large database schemas at scale.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Kamal</given-names>
            <surname>Azzaoui</surname>
          </string-name>
          et al.
          <article-title>Scientific competency questions as the basis for semantically enriched open pharmacological space development</article-title>
          .
          <source>Drug Discovery Today</source>
          <volume>18</volume>
          (
          <fpage>17</fpage>
          -
          <lpage>18</lpage>
          ) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Oscar</given-names>
            <surname>Corcho</surname>
          </string-name>
          , Mariano Fernndez-Lopez,
          <article-title>Asuncin Gmez-Prez. Methodologies, tools and languages for building ontologies: Where is their meeting point? Data Knowl</article-title>
          .
          <source>Eng</source>
          .
          <volume>46</volume>
          (
          <issue>1</issue>
          ) (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Pascal</given-names>
            <surname>Hitzler</surname>
          </string-name>
          , Aldo Gangemi, Krzysztof Janowicz, Adila Krisnadhi, Valentina Presutti (eds.),
          <article-title>Ontology Engineering with Ontology Design Patterns: Foundations and Applications</article-title>
          .
          <source>Studies on the Semantic Web 25</source>
          , IOS Press/AKA,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Maria Keet</surname>
          </string-name>
          , Agnieszka Lawrynowicz:
          <article-title>Test-Driven Development of Ontologies</article-title>
          . ESWC 2016
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Yuan</given-names>
            <surname>Ren</surname>
          </string-name>
          , Artemis Parvizi, Chris Mellish, Jeff
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          , Kees van Deemter,
          <string-name>
            <given-names>Robert</given-names>
            <surname>Stevens</surname>
          </string-name>
          .
          <article-title>Towards Competency Question-Driven Ontology Authoring</article-title>
          .
          <source>ESWC 2014</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Juan</given-names>
            <surname>Sequeda</surname>
          </string-name>
          , Marcelo Arenas, Daniel P. Miranker.
          <article-title>On directly mapping relational databases to RDF and OWL</article-title>
          . WWW 2012
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Juan</surname>
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Sequeda</surname>
          </string-name>
          , Syed Hamid Tirmizi, Oscar Corcho, Daniel P. Miranker.
          <article-title>Survey of directly mapping SQL databases to the Semantic Web</article-title>
          .
          <source>Knowledge Eng. Review</source>
          <volume>26</volume>
          (
          <issue>4</issue>
          ):
          <fpage>445</fpage>
          -
          <lpage>486</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Mike</given-names>
            <surname>Uschold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Gruninger</surname>
          </string-name>
          .
          <article-title>Ontologies: principles, methods and applications</article-title>
          .
          <source>Knowledge Eng. Review</source>
          <volume>11</volume>
          (
          <issue>2</issue>
          ):
          <fpage>93</fpage>
          -
          <lpage>136</lpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>