<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SW at the department 'Environment' of the Flemish Government</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Paul Hermans</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ProXML</institution>
          ,
          <addr-line>3140 Keerbergen</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This project proves that the semantic technology stack and related tooling allow to do data integration in a fast and agile way. Several technologies have been utilised: Ontology-Based Database Access, federated SPARQL, federation middleware. The solution is using the RDF Data Cube vocabulary for capturing the emission observations done. The additional 5 star LOD publishing was easily achieved at a minimal cost.</p>
      </abstract>
      <kwd-group>
        <kwd>Data integration</kwd>
        <kwd>OBDA</kwd>
        <kwd>federated search</kwd>
        <kwd>LO(S)D</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background</title>
      <p>The ‘Omgeving’ Department is the environmental administration of the government of
Flanders. It is responsible for preparing, following up and evaluating the Flemish
environmental policy. In Belgium, companies that want to emit polluting substances in the
air or water must have an environmental permit. Some of them are also obliged to report
annually about the emissions of the previous year. In fact, data regarding the amount of
substances per location have been collected since 2004.</p>
      <p>The idea was to integrate these data with all kinds of other data sources (internal and
external) to offer:
• the general public an application showing the emitted substances over the years in
the area that they live
• companies the ability to benchmark themselves against other, similar, organisations
• public servants analytics dashboards to gain insights to steer the policy making.
The ultimate aim was that by integrating the data (research) questions can be addressed
not being able to be answered on the separate data silos.</p>
      <p>The solution preferably needed to be based on open source software or using open
source libraries.</p>
      <p>
        The project serves also as a pilot site of the OpenGovIntelligence project[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
receiving funding from the European Union’s Horizon 2020 research and innovation
programme under grant agreement No 693849, which aims to modernize Public
Administration by connecting it to Civil Society through the innovative application of Linked
Open Statistical Data (LOSD).
      </p>
    </sec>
    <sec id="sec-2">
      <title>Solution</title>
      <sec id="sec-2-1">
        <title>Solution outline</title>
        <p>Dataintegration. The actual status is that more than 10 data silos have been integrated.
These datasets contain data managed by the department itself (archival systems,
RDBMS) and datasets published by other departments (Flemish addresses database,
Belgian company register) together with well-known classification systems such as
NACE (economical activities) and NIS (administrative geographical entities).</p>
        <p>
          The solution takes a hybrid approach. For some datasets, the data are transformed
via dedicated ETL processes into triples and these are loaded in a triple store. Other
datasets, mainly from existing relational databases, are virtualized via OBDA[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
(Ontology-Based Data Access). The own ontology and controlled vocabularies will be
managed with VocBench V3[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. There are also connections to external SPARQL
endpoints. In front of all these different SPARQL endpoints we offer a federation layer.
Data model. The most important datasets are in fact observations. Hence our use of the
RDF Data Cube vocabulary[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]: a vocabulary explicitly made to capture statistical data
to allow OLAP operations such as slice and dice, roll-up and drill-down. The code lists
are encoded in XKOS[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], an extension of SKOS for statistical classifications.
LOD Publishing. All entities have their subject pages published using
dereferenceable URL’s. Next to this there are also public SPARQL endpoints available. But the
important point is that the LOD publishing was not the first aim of the project: it was a
nice free add-on.
        </p>
        <p>
          General Public Dedicated Application. An application has been build showing the
emitted substances over the years in the area that they live. This is a traditional web
application using Polymer [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] web components, which is the standard in the
department. This application is talking to the integrated set of data via the available SPARQL
endpoints.
        </p>
        <p>
          Business Intelligence. A connector has been developed for exploratory [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], a Data
Science environment based on RStats [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], to connect with triple stores, so that BI reporting
can be done using the integrated datasets.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Business Benefits</title>
        <p>The approach used has clearly proven that semantic technologies allow to do data
integration in a fast and agile way. Answers can now be given to questions which involve
data from multiple datasets. The fact that 5 star LOD publishing is only one step away
is a nice add-on. The system is used internally for the moment and will be become open
to the public early November.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. OpenGovIntelligence, http://www.opengovintelligence.eu/</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Diego Calvanese, Giuseppe De Giacomo, Domenico Lembo, Maurizio Lenzerini, Antonella Poggi, Riccardo Rosati,
          <article-title>Ontology-based database access</article-title>
          , http://www.dis.uniroma1.it/~degiacom/papers/2007/sebd07.pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>3. VocBench V3, http://vocbench.uniroma2.it/</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>RDF</given-names>
            <surname>Data Cube</surname>
          </string-name>
          <string-name>
            <surname>Vocabulary</surname>
          </string-name>
          , https://www.w3.org/TR/vocab
          <article-title>-data-cube/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. XKOS,
          <article-title>Extended Knowledge Organization System</article-title>
          , http://www.ddialliance.org/Specification/RDF/XKOS
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Polymer</given-names>
            <surname>Project</surname>
          </string-name>
          , https://www.polymer-project.org/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>7. Exploratory, https://www.exploratory.io/</mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. RStats, https://www.r-project.org/</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>