<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Business Driven Insight via a Schema-Centric Data Fabric</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dougal Watt</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brad Bebee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Schmidt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amazon Web Services</institution>
          ,
          <addr-line>Seattle, WA 98101</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Meaningful Technology</institution>
          ,
          <addr-line>Auckland, NZ</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Research shows that Analytics and Business Intelligence (BI) projects have high failure rates of up to 80 percent of projects, and suffer from low uptake of big data and AI tooling. Our internal research indicates two core problems that customers face when generating business outcomes from analytics. Firstly, business customers are often totally overwhelmed in knowing where to start. They report that even the simplest modern BI tools provide little or no help in analyzing their data and do not reflect the needs of their business. Secondly, when they do generate insights from their data they often have no trust in these, as they cannot understand provenance information about where the data came from, what was done to it by systems and individuals in their organization. We present a new approach to solving these analytics challenges using a new semantic 'data fabric' platform for business computing, based on the Amazon Neptune graph database service from Amazon Web Services (AWS) and other AWS cloud technologies. This architecture is purpose built around a semantic schema and custom tools that model and manage business data in a form understandable by business users. As a schemacentric platform, all insight is grounded from the definitions and relationships in actual business data, and provenance data captured throughout the platform and across all data flows.</p>
      </abstract>
      <kwd-group>
        <kwd>Business Model</kwd>
        <kwd>Analytics</kwd>
        <kwd>Insight</kwd>
        <kwd>Business Intelligence</kwd>
        <kwd>Cloud</kwd>
        <kwd>Neptune</kwd>
        <kwd>RDF</kwd>
        <kwd>SPARQL</kwd>
        <kwd>Architecture</kwd>
        <kwd>Schema</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Recent research has highlighted problems with organizations deploying analytics
and business intelligence tools and processes to gain business insights, and use these
to effect business performance improvements. Gartner [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] estimates that until 2022,
around 80 percent of analytics insights will not deliver business outcomes, while a
survey by NewVantage [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] showed that 77 percent of businesses report challenges in
the adoption of big data and AI, with many companies failing to use deployed tools.
1.1 Solving the Analytics Challenge
Existing BI and AI tools typically require moderate to significant domain knowledge
to operate. For business users in the SME segment, these issues are often a
considerable barrier as they typically do not have staff with the necessary skills and experience,
and limited funds to afford consultants with the requisite skills. Similarly, traditional
tools require significant effort in data preparation for data loads for machine learning,
or data extraction, transformation and loading into data warehouses. The complex
steps required frequently result in loss of context for business users, such that the link
between business information and insight becomes lost and insights remain unused.
      </p>
      <p>To solve these issues, we created a novel approach to analytics using a
‘schemacentric data fabric’ architecture. A schema directly models business structures and
terminology, and is used to drive the user experience of selecting, integrating,
preparing, and shipping data to external analytics tools integrated with the platform. This
removes the need for specialist, time consuming data preparation and tooling
integration. All interactions with data are conformed to this schema, so any externally
generated insights are immediately understandable by business users in their own language.</p>
      <p>Our architecture addresses five key use cases: 1. manage a business model; 2.
integrate data from cloud applications and data stores; 3. store/conform data into a
semantic database according to the schema; 4. generate business insight from this integrated
data; and 5. simplify tooling and operations via an integrated suite of interfaces,
custom tools, cloud deployment architectures, and pre-built cloud services.</p>
      <p>In our presentation we will outline our use of RDF/OWL to model business
models, SHACL as a constraint and query language for integrating cloud apps, a full
integrated suite of Amazon Web Services including the Neptune Database to store and
manage conformed data in the form of a semantic RDF graph, and a set of custom
components and user interfaces to generate insight. We will also showcase user
interfaces used to manage these artifacts, with reference to specific customer use cases.</p>
      <p>We highlight key aspects of Amazon Neptune that are relevant to our use case,
including (i) high availability through data replication and automated failover, (ii) its
transaction semantics, which facilitates parallel updates from independent sources
with transactional guarantees, and (iii) Neptune’s support for lightweight in-database
analytics. The flexibility to scale the graph database horizontally with read replicas on
demand and using them for analytical workloads and large data exports allows us to
scale compute resources and to separate resource intensive queries from the
continuous OLTP workload, thus avoiding interference of analytics with regular operations.</p>
      <p>Early work extending this architecture in novel directions includes the ability of
Neptune to integrate with other data analytics and ML frameworks in the AWS
ecosystem such as the Amazon Forecast service, a fully managed ML service to deliver
forecasts that we will feed back into our insight tools. We will also showcase
extensions to automate the operation of traditional data warehouses, using our insight
tooling and Neptune’s Change Data Capture feature to subscribe to changes in graph data.</p>
      <p>We are currently applying our approach with several businesses in New Zealand,
with results showing rapid creation of analytics dashboards, simplified
whole-ofcompany performance reporting, and feeding existing data warehouses with
harmonized graph data. Use cases and lessons learnt from these customers will be discussed.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Gartner, https://blogs.gartner.com/andrew_white/
          <year>2019</year>
          /01/03/our-top
          <article-title>-data-and-analyticspredicts-for-</article-title>
          <year>2019</year>
          /, last accessed
          <year>2019</year>
          /06/25.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. NewVantage, https://newvantage.com/wp-content/uploads/2018/12/
          <string-name>
            <surname>Big-Data-ExecutiveSurvey-</surname>
          </string-name>
          2019-Findings.pdf,
          <source>last accessed</source>
          <year>2019</year>
          /06/25.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>