<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Use of Semantic Technologies to Inform Progress Toward Zero-Carbon Economy⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Extended Abstract</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Oxford</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mechanical Engineering, University of Bath</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>To investigate the efect of possible changes to decarbonise the economy, a detailed picture of the current production system is needed. Material/energy flow analysis (MEFA) allows for building such a model. There are, however, prohibitive barriers to the integration and use of the diverse datasets necessary for a system-wide yet technically-detailed MEFA study. Here we present a demonstration of the “Physical Resources Observatory” (PRObs) system, which exploits Semantic Web technologies to integrate and reason on top of this diverse production system data. We demonstrate how this system supports domain experts to load datasets, map them to the ontology, and query the extended data to meet several modelling use cases.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic Technology</kwd>
        <kwd>Resource eficiency</kwd>
        <kwd>Rule-based approach</kwd>
        <kwd>Data integration</kwd>
        <kwd>Material Flow Analysis</kwd>
        <kwd>Ontology</kwd>
        <kwd>Decision Support System</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A whole-systems understanding of production systems is essential to navigating
the necessary rapid transition to a zero-carbon economy. Understanding how
we produce and consume physical resources is fundamental to understanding
the impacts human activity has, and opportunities to increase eficiency. This
understanding relies on access to data about the production and consumption of
physical resources (materials, energy, etc.) and their associated environmental
impacts. Due to the economy-wide yet detailed nature of these questions, they
must be based on many pieces of data from diverse sources. Moreover, this data is
incomplete, and defined using inconsistent categorisations of the types of resource
and activities. In addition, the lack of well-defined data models for this type of
data is limiting to data reuse and holding back academic research [
        <xref ref-type="bibr" rid="ref3 ref7">7,3</xref>
        ].
      </p>
      <p>We propose and develop a solution using a domain ontology and the RDFox
triple store to eficiently implement Datalog rules integrating diverse data points
into a consistent structure. This forms part of the “Physical Resources
Observatory” (PRObs) system, being developed within the UK FIRES research
programme3, where it supports a wider research agenda on resource eficiency
and decarbonisation in UK industrial strategy.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Modelling and reasoning in PRObs</title>
      <p>
        To allow quantified data points on resource use to be expressed in RDF, we
build on the data model proposed by Pauliuk et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. More details about our
conceptual data model can be found in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Diferent data points have been
defined with respect to diferent classification systems, and must be easily and
transparently retrieved for reuse in new analyses. More broadly, new information
needs to be inferred from the raw data using rules. Generally, this might involve
converting data between diferent definitions of time, location, activity, and object
type. We use the Datalog language with stratified negation and aggregates to
perform these computations. This allows the complex behaviours required to be
expressed in simple rules, while benefiting from the eficient solvers available for
evaluating Datalog programs.
      </p>
      <p>
        The ontology and the RDFox scripts implementing the rules and algorithm
are available online at [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Implementation in PRObs</title>
      <p>The PRObs system is intended to be used to support modelling by domain
experts who are not familiar with semantic web technologies. As such, they
should be supported to enter information (e.g. about materials of interest and
their hierarchical structure) and retrieve results without becoming experts in
RDF and complex SPARQL queries. The system should be usable (as far as
possible) on typical researchers’ computers, and integrate with typical modelling
workflows involving e.g. Python scripts and notebooks. Further, because defining
the system is subjective, the implementation should support users in clearly
documenting this input information.
3 https://ukfires.org
Raw
Data
PRObs
TBox</p>
      <p>Pre-processing</p>
      <p>Ontology
Conversion</p>
      <p>Clean
Data</p>
      <p>OFN
Ontology</p>
      <p>External
Ontologies</p>
      <p>PRObs
Ontology</p>
      <p>Queries
Conversion</p>
      <p>Reasoning</p>
      <p>Answers</p>
      <p>
        Specifically, our system is composed of a frontend interface for defining
and documenting the system definitions as input RDF data, and the backend
implementation based on RDFox [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Our backend pipeline is shown in Figure 1.
More details on the full pipeline of the PRObs system are available in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>To make the use of the PRObs system accessible to domain experts we adopt
a literate programming approach to produce code (RDF) and documentation
(HTML) from a single source.</p>
      <p>This allows for full documentation-writing features within Python executable
notebooks. They are easy-to-use, interactive, and widely used by scientists in
various fields; therefore, they are the ideal platform for a system that aims to
support the needs of domain experts. To streamline use in a MEFA analysis,
we have developed Python wrappers that manage the process of setting up the
RDFox scripting language commands to load the relevant datasets, and running
RDFox as part of a wider workflow to answer queries accessing the relevant
information which form the input to subsequent modelling and analysis steps.
The PRObs system runs RDFox scripts to load the input data and answer queries,
supported by Python utilities to embed this within a testing or analysis workflow.
The input data consists of system definitions in RDF, with external datasets
provided in the form of tabular data files and mapping scripts which are read
during processing by RDFox.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Case study and demo</title>
      <p>
        To illustrate the use of the system, we describe a case study of mapping flows
through the UK production system. This case study will form the main part of
our demo. In particular, we will demonstrate several use cases within this case
study displaying particular intricacies of the PRObs system. A working example
of the PRObs system containing this case study and the material that will be
used for our demo is available online at [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
4.1
      </p>
      <p>UK production system
This case study forms part of the ongoing research within the UK FIRES
programme, motivated by seeking opportunities for innovation in manufacturing
processes. The goal is to obtain a detailed understanding of how supply chains
are dependent on diferent manufacturing processes, and where scrap is currently
arising within the system, in order to quantify the benefits of innovation in
diferent areas . To this end, a MEFA model is used to define the structure of
supply chains and estimate the pattern of flows through the system which best
matches the available measurements. The role of the PRObs ontology described
here is to provide access to data from a diverse set of external datasets in a
coherent structure aligned to the required inputs of the optimisation model.</p>
      <p>
        Since there is no standard system definition of UK manufacturing supply
chains at the level of technical detail required for this analysis, a major element
of the project is to describe a suitable set of Processes and Objects to which the
available data can be mapped, and which describe entities of relevance to the
study’s research questions. These are defined and documented using the system
described in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. At the moment, our whole project includes 701 processes and
617 object types, but in our demo we are going to show only a subset of them.
The datasets used include Prodcom4, Comtrade5, and BGS6.
      </p>
      <p>Example use cases will be demonstrated through the Jupyter Book7, with the
opportunity to demonstrate live queries. During our demo we are going to show
the full pipeline of the PRObs system, from loading the source data to the
reasoning and querying phases. The demo will encompass sample datasets loaded
and mapped into the ontology, as well as use cases which include: retrieving data
points with their system context, retrieving inferred observations at aggregated
object level, and identifying provenance of observations. Concrete details about
the challenges of our scenario and the lessons learned during this work will be
provided as well.</p>
      <p>Specifically, we will introduce the data sources mentioned above, and we
will describe loading and retrieval details for some objects in the BGS dataset
using the PRObs system. Then, we will show why retrieving the original data
points in not enough in this scenario and how we can derive new information
about the observations thanks to the modelling and reasoning capabilities of our
system. We will focus on (i) the aggregation of observations, useful when we
want to derive all possible measurements of an object that can be classified into
further components; (ii) the propagation of observations, useful when we want to
analyse equivalent objects; and (iii) the provenance of observations, useful for
explainability and quality of information. Finally, we will present more complex
real-world examples and explain why this inferred knowledge is useful for domain
experts and for further analyses.
4 Statistics on the production of goods and materials within the EU (Eurostat).</p>
      <p>https://ec.europa.eu/eurostat/web/prodcom
5 Statistics on international trade (United Nations). https://comtrade.un.org
6 British Geological Survey Minerals Yearbook : data tables for production of various
minerals. https://www2.bgs.ac.uk/mineralsuk/download/ukmy/UKMY2015.pdf
7 https://jupyterbook.org</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Germano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saunders</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horrocks</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lupton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Use of semantic technologies to inform progress toward zero-carbon economy</article-title>
          .
          <source>In: International Semantic Web Conference (ISWC)</source>
          (
          <year>2021</year>
          ), [forthcoming]
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Germano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saunders</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lupton</surname>
          </string-name>
          , R.: ukfires/probs-ontology
          <source>: probs-ontology v1.5.2 (Jul</source>
          <year>2021</year>
          ). https://doi.org/10.5281/zenodo.5052739
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hertwich</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heeren</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuczenski</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majeau-Bettez</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Myers</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pauliuk</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stadler</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lifset</surname>
          </string-name>
          , R.:
          <article-title>Nullius in Verba: Advancing Data Transparency in Industrial Ecology</article-title>
          .
          <source>Journal of Industrial Ecology</source>
          <volume>22</volume>
          (
          <issue>1</issue>
          ) (
          <year>2018</year>
          ). https://doi.org/10.1111/jiec.12738
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Lupton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Germano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saunders</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>: ukfires/probs-ISWC2021-example: Initial release for ISWC2021 paper</article-title>
          (
          <year>Apr 2021</year>
          ). https://doi.org/10.5281/zenodo.5052758
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Nenov</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piro</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horrocks</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>RDFox: A highly-scalable RDF store</article-title>
          .
          <source>In: The Semantic Web - ISWC 2015 - 14th International Semantic Web Conference</source>
          , Bethlehem, PA, USA, October
          <volume>11</volume>
          -
          <issue>15</issue>
          ,
          <year>2015</year>
          , Proceedings,
          <source>Part II. Lecture Notes in Computer Science</source>
          , vol.
          <volume>9367</volume>
          . Springer (
          <year>2015</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -25010-
          <issue>6</issue>
          _
          <fpage>1</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Pauliuk</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heeren</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
          </string-name>
          , D.B.:
          <article-title>A general data model for socioeconomic metabolism and its implementation in an industrial ecology data commons prototype</article-title>
          .
          <source>Journal of Industrial Ecology</source>
          <volume>23</volume>
          (
          <issue>5</issue>
          ) (
          <year>2019</year>
          ). https://doi.org/10.1111/jiec.12890
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pauliuk</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majeau-Bettez</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mutel</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steubing</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stadler</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Lifting Industrial Ecology Modeling to a New Level of Quality and Transparency: A Call for More Transparent Publications and a Collaborative Open Source Software Framework</article-title>
          .
          <source>Journal of Industrial Ecology</source>
          <volume>19</volume>
          (
          <issue>6</issue>
          ),
          <fpage>937</fpage>
          -
          <lpage>949</lpage>
          (
          <year>Dec 2015</year>
          ). https://doi.org/10.1111/jiec.12316
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>