<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>The CitySPIN Platform: A CPSS Environment for City-Wide Infrastructures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amr Azzam</string-name>
          <email>aazzam@wu.ac.at</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Di Ciccio</string-name>
          <email>claudio.di.ciccio@ai.wu.ac.at</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sotiris Karampatakis</string-name>
          <email>sotiris.karampatakis@semantic-</email>
          <email>sotiris.karampatakis@semanticweb.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marta Sabou</string-name>
          <email>marta.sabou@tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peb R. Aryan</string-name>
          <email>peb.aryan@tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fajar J. Ekaputra</string-name>
          <email>fajar.ekaputra@tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elmar Kiesling</string-name>
          <email>elmar.kiesling@tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pujan Shadlau</string-name>
          <email>Pujan.Shadlau@wienerstadtwerke.at</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessio Cecconi</string-name>
          <email>cecconi@ai.wu.ac.at</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Javier Fernández</string-name>
          <email>jfernand@wu.ac.at</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Angelika Musil</string-name>
          <email>angelika.musil@tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Thurner</string-name>
          <email>t.thurner@semantic-web.at</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ISE Institute, TU Wien</institution>
          ,
          <addr-line>1040 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Semantic Web Company</institution>
          ,
          <addr-line>1070 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>WU Vienna</institution>
          ,
          <addr-line>1020 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Wiener Stadtwerke Holding AG</institution>
          ,
          <addr-line>1030 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>22</volume>
      <issue>2019</issue>
      <fpage>57</fpage>
      <lpage>64</lpage>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Cyber-physical Social System (CPSS) are complex systems that span
the boundaries of the cyber, physical and social spheres. They play
an important role in a variety of domains ranging from industry
to smart city applications. As such, these systems necessarily need
to take into account, combine and make sense of heterogeneous
data sources from legacy systems, from the physical layer and also
the social groups that are part of/use the system. The collection,
cleansing and integration of these data sources represents a major
efort not only during the operation of the system, but also
during its engineering and design. Indeed, while ongoing eforts are
concerned primarily with the operation of such systems, limited
focus has been put on supporting the engineering phase of CPSS.
To address this shortcoming, within the CitySPIN project we aim to
create a platform that supports stakeholders involved in the design
of these systems especially in terms of support for data
management. To that end, we develop methods and techniques based on
Semantic Web and Linked Data technologies for the acquisition
and integration of heterogeneous data from disparate structured,
semi-structured and unstructured sources, including open data and
social data. In this paper we present the overall system
architecturewith a core focus on data acquisition and integration.We
demon-strate our approach through a prototypical implementation
of an adaptive planning use case for public transportation
scheduling.
1</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>Cyber-physical Systems (CPSs) are systems that span the physical
and cyber-world by linking objects and process from these spaces.
A typical CPS collects data from the physical world via sensors and
applies computation resources from the cyber-space to integrate
and analyze this data in order to decide on optimal feedback
processes that can be put in place by physical actuators. CPSs have
started to diffuse into many areas, including mission-critical public
transportation, energy services, and industrial production and
manufacturing processes.</p>
      <p>
        The results of a recent study about adaptation in CPS [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] revealed
an emerging trend to add an additional social layer in a CPS
architecture to address human and social factors and evolve these
systems into CPSSs [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The resulting systems consist not only of
software and raw sensing and actuating hardware, but are
fundamentally grounded in the behaviour of human actors,
who both generate data and make informed decisions based on data
[
        <xref ref-type="bibr" rid="ref12 ref22 ref5">5, 12, 22</xref>
        ].
      </p>
      <p>The CitySPIN1 project aims to lay a foundation for the
development of CPSSs in the context of Smart City infrastructure services.
To this end, we develop both theoretical and conceptual foundations,
as well as a set of innovative components — illustrated in Figure 1
— that support a CPSS design process in a uniform platform. This
platform supports key stakeholders involved in the design process
through a prototyping environment that provides a visual interface
which allows them to (i) access a wide range of data sources from
sensors, social channels, and legacy systems; (ii) integrate and
analyze heterogeneous data; and (iii) visualise results. This platform
is made possible by methods and tools that make use of Semantic
Web and Linked Data technologies to support the collection and
integration of heterogeneous data sources.</p>
      <p>In this paper, after a brief overview of the CitySPIN arechitecture
in Section 2, we focus on the two core aspects of this technology
stack: the knowledge graph construction, covered in Section 3, and
the prototyping environment, described in Section 4. Furthermore,
we discuss the prototypical implementation and illustrate the
application of the platform by means of an example use case involving
Vienna’s largest public transport provider in Section 5. Finally, we
briefly review related work in Section 6 and conclude the paper
with an outlook on future research in Section 7.
2</p>
    </sec>
    <sec id="sec-3">
      <title>CITYSPIN ARCHITECTURE OVERVIEW</title>
      <p>The design of cyber-physical social systems raises challenges due
to high complexity introduced by social systems in terms of:
(i) the number and heterogeneity of data sources that need to be
integrated: CPSSs involve large amounts of heterogeneous,
polystructured data from a variety of sources, ranging from legacy
databases to highly dynamic sensor data. To create CPSS
applications and services, it is paramount to eficiently integrate
not just the data produced by individual processes within the
organization, but to achieve integration across processes,
departments, organizational boundaries, and domains. Finally,
external data, such as, for instance, social media streams, are
also of pivotal importance in the context of CPSSs. Hence, a
major challenge is to develop flexible data integration
infrastructures that are responsive to the varying needs of CPSSs.
(ii) privacy concerns associated with the processing of sensitive social
data: Adequate privacy protection is a fundamental
requirement in the context of CPSSs, which often make use of and
integrate sensitive information from various sources. Additionally,
the new EU General Data Protection regulation imposes new
demands in terms of transparency of data processing and also
in terms of allowing data subjects to revoke or change their
consent in parts, which calls for more flexible and dynamic
compliance checking. This represents a significant barrier
towards the development and provision of integrated smart city
services and hinders product and process innovation.
(iii) uncertainty due to social dynamics: CPSS designers need a
better understanding of the social dynamics of the groups
involved in the CPSS, both at the design time and the run-time
of the system (e.g., for on-the-fly adaptation).</p>
      <p>All these challenges are amply reflected in a CitySPIN use case
that aims to improve the daily schedule planning for the
Viennese public transport network. In particular, this use case aims
to support planners in their work by allowing them to treat the
transportation system as a CPSS and accounting for the
dynamics of the involved travelers (especially during large-scale events).
This requires, amongst others, the integration of data from various
sources including data internal to the organization (e.g., historic
data about event attendance), open data (e.g., expected events), as
well as real-time data from mobility operators. Some of these data
sources can raise privacy concerns (e.g., when harvested by apps
installed on individual mobiles) and therefore user consent about
the use of this data needs to be appropriately captured and
considered during data processing. Finally, network planners would
like to understand recurring social behaviors and patterns – for
example, the typical routes followed by participants of an event.</p>
      <p>CitySPIN tackles these challenges in the design and prototyping
phases of a CPSS and aims to ofer support to key stakeholders
involved in these stages including decision makers, project
managers, software architects, and software engineers, as depicted in
Figure 1. These stakeholders are provided with a CPSS Prototyping
Environment that adopts a mashup-based paradigm to allow them
to easily acquire, explore, combine and visualise a variety of data
sources (e.g., legacy data, streaming data, social media data, open
data). The CPSS Prototyping Environment relies on and is made
possible by three key components, as described next.</p>
      <p>Scalable Linked Data Integration. We adopt Linked Data
technologies to address the integration of multiple, heterogeneous data
sources. To this end, we developed dedicated components for the
acquisition and semantic enrichment of data as well as the
integration into a CPSS Knowledge Graph. The next sections of this paper
will focus on CitySPIN’s data integration architecture primarily.</p>
      <p>Secure Data Access and Privacy. To deal with privacy concerns
typically associated with social data, we develop components for
capturing user consent and making use of this consent during the
entire data integration chain.</p>
      <p>
        Process Mining on Linked Data. Finally, to support stakeholders in
gaining a better insight into group dynamics, we develop a Process
Mining &amp; Analytics component that can be used to analyze
behavioral patterns and make predictions based on the CPSS knowledge
graph. Process mining is the discipline connecting data science and
business process management that aims at discovering, checking,
and enhancing business processes based on data logged by
information systems [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. In the context of this project, we resort in
particular on declarative process mining to cater for the flexibility
of the processes considered in this project [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Declarative
process models specify dynamic systems through temporal-logic-based
rules that establish the constraints with which the execution must
comply. Therefore, we resort on the expression of those constraints
as queries over the CPSS knowledge graph to monitor and analyse
the behavior of ongoing processes [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The query answers are thus
routed to the CPSS Mashup Platform to allow for further complex
analytics and refinements.
      </p>
    </sec>
    <sec id="sec-4">
      <title>CPSS KNOWLEDGE GRAPH</title>
    </sec>
    <sec id="sec-5">
      <title>CONSTRUCTION</title>
      <p>The broad scope of CPSSs and the large variety of technical
infrastructure and data involved in them give rise to unique
interoperability challenges when it comes to acquiring, enriching, integrating,
managing, and processing data from various sources pertaining to
the social, physical, and cyber dimensions of CPSSs. In CitySPIN
a CPSS knowledge graph acts as an integration hub for all this
heterogeneous data. In this section we describe the technologies
used to construct this knowledge graph.</p>
      <p>We rely on UnifiedViews 2, developed at Semantic Web
Company (SWC), as a core building block for the knowledge graph
construction. Specifically, data sources are aggregated and
transformed using so-called Data Processing Units (DPU)s, which are
assembled into data integration pipelines3. All input data is
available in a structured format for further processing by subsequent
elements of the pipeline. The pipelines transform data from various
source formats and lift them into Resource Description Framework
(RDF) format, a semantically explicit format standardized by the
World Wide Web Consortium (W3C). This results in a knowledge
graph that expresses the data using common standard vocabularies
as well as vocabularies tailored to the use cases. Table 1 provides an
overview of the key vocabularies used for the semantic alignment
of the various datasets which underlie the public transportation
planning use case used as an illustrative example in this paper.</p>
      <p>The knowledge graph is stored into an RDF triple store –
specifically Ontotext GraphDB4. Using the standard RDF query language
2https://unifiedviews.eu
3cf. https://help.poolparty.biz/display/UDDOC/Basic+Concepts+for+DPU+developers
for an introduction to the core concepts
4http://graphdb.ontotext.com</p>
      <p>SPARQL5, input data is queried and further transformed or
aggregated as needed by other components of the CitySPIN platform.
The following paragraphs discuss the concepts applied here in more
detail.</p>
      <p>Data Integration Lifecycle. Heterogeneous sources such as social
media data, sensor data and business intelligence data have to be
made available to the CPSS for further processing. Connectors to the
source systems hook into APIs, CSV repositories or direct database
calls (data acquisition). Various steps follow to remove outliers and
noise from data (data cleansing) as well as to refine their structure
and align their content (data preparation). Finally data are merged,
transformed and saved into pre-processable formats (data storage).
The consolidated data are then available for subsequent analysis
and reuse. Therefore, those data are fed back to the acquisition
stage and the integration cycle restarts.</p>
      <p>Data Acquisition and Enrichment. As we consistently follow an
ontology-based data integration approach, we extract data
according to a CPSS-wide ontology and transform the data into RDF. The
RDF is, in turn, an interchange format which is used as the
canonical one for further processing. By following the W3C standards for
the Semantic Web, our approach ensures compatibility with a wide
range of tools used in the CPSS stack.</p>
      <p>Semantic Alignment for Data Integration. Aligning contents and
data alongside an ontology enables the CPSS to access enriched
contextual knowledge. This additional information forms a critical
part of an integrated view on CPSS data and is essential for realizing
the integrated user interface presenting the planning dashboard.</p>
      <p>Data Cleansing. All gathered data is integrated and enriched by
a processing pipeline, which lies at the functional core of the CPSS
(e.g.: prediction, analysis, decision). Based on domain knowledge
5https://www.w3.org/TR/sparql11-query/</p>
      <p>Name
Time Ontology
Geovocab geometry
Geovocab spatial
wgs84
Event
Cellular data
SPECIAL-CPSS
Transport
and process knowledge, data are consolidated and made available
for extraction and further processing by actuators, visualization
and re-feeds into the learning pipeline.</p>
      <p>Knowledge Graph Storage. The central processing pipeline acts as
an interface to other algorithms, further user-driven explorations,
or visual representations of the output. The loop-back to the Data
Acquisition stage of the CPSS is realized through interim storage
in a central triple store and actuation of external triggers.</p>
    </sec>
    <sec id="sec-6">
      <title>4 CITYSPIN PROTOTYPING ENVIRONMENT</title>
      <p>
        For the implementation of the prototyping environment, the CitySPIN
project proposes an architecture inspired by the Presentation
Abstraction Control (PAC) architectural pattern [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In this section,
we adapt the PAC architecture to the CPSS needs and integrate it
with the modular approach of Linked Widgets [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and
UnifiedViews [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] to develop a CPSS prototyping environment. In the
following subsections, we illustrate the longitudinal section of the
software architecture to describe the associations and information
lfows between the main logical components at large. In line with
the PAC pattern, three-layered architecture of the CitySPIN CPSS
prototyping environment consists of:
• the Back-end layer, in which data are loaded, pre-processed,
and aggregated (abstraction) - details on this layer are
previously discussed in Section 3, together with the CPSS
Knowledge Graph construction and therefore will not explained
further in this Section;
• the Service layer, in which those data are queried and
analyzed to infer additional knowledge and later on generate
prediction models (control) - cf. Section 4.1;
• the Front-end layer (presentation), from which users can
access the prediction models and data analysis reports to
monitor the current status of the infrastructure, explore the
historic performance, and make informed decisions on the
future settings (Section 4.2).
      </p>
    </sec>
    <sec id="sec-7">
      <title>4.1 Service: Querying and Prediction</title>
      <p>There are a wide range of services required in the CPSS context
due to the diversity of application domains, use cases and scenarios.
In our CPSS Prototyping environment, we focused on two main
services: (i) Querying, and (ii) Prediction.</p>
      <p>To cater for the reporting and predicting needs of a CPSS, our
architecture includes an intermediate layer in which data are
extracted from the Querying component and fed to the Prediction
component or directly to the Dashboard of the frontend layer. The
Querying component of our prototyping environment relies on
the data endpoint provided by the back-end module for the
execution of queries. In this component, we are using the W3C-standard
SPARQL query language6 for querying the integrated data.
Furthermore, we can also use SPARQL Construct queries to encode rules
for inferring new knowledge.</p>
      <p>
        The Prediction component is designed to allow for the
application of Machine Learning (ML) techniques aimed to derive
prediction models that – based on historical data – can be used as a
decision support system for CPSS stakeholders to react ahead of
time to predicted arising situations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Example prediction results
include the forecast of numerical trends of variables under
analysis, the identification of changes in the classification of recently
collected data to raise alerts in case of anomalies, or the
recommendation on the next operation to undergo in light of the recent
developments of the data under observation.
      </p>
      <p>
        ML algorithms require learning, validation, and testing phases
on historical data, prior to, or alternated with, run-time processing
or reinforcement on live data. To cater for these requirements, our
architecture binds the Querying and Prediction components with
data-flow associations that proceed in both ways: (i) from Querying
to Prediction for data feed, and (ii) from Prediction to Querying for
updates on the classifications and predictions made. Notice that
this architectural choice allows for the marshalling and storage of
models learned from the Prediction component for further reuse.
This is the basis through which ex-post data analyses conducted via
process mining can be readily available for decision support and
monitoring via successive queries, as suggested in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Finally, we
emphasize that both the Querying and Prediction components are
containers for diverse ML modules that can be used alternatively
in multiple use cases, e.g., as a plugin for Linked Widgets.
4.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>Frontend: Visualization and Decision</title>
    </sec>
    <sec id="sec-9">
      <title>Support</title>
      <p>
        The Frontend layer allows users to interact with, get informed
about, and interactively explore the integrated knowledge acquired
from the data and augmented by the Prediction component. To this
end, the use of Linked Widget Platform (LWP) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] provides the
necessary high degree of flexibility and customizability for CPSS
prototyping. LWP combines semantic web and mashup concepts to
support non-expert users in eficiently making use of various open
and non-open data sources. In particular, the platform allows users
to collaboratively and interactively integrate data in an ad-hoc and
distributed manner. Each stakeholder can contribute their data and
computing resources to a shared data processing flow in a shared
interface that allows them to orchestrate the interaction among
components within a CPSS.
      </p>
      <p>Depending on their needs, users can directly construct
analytical data flows, fine-tune queries, ML parameters, and visualization
parameters within a single graphical interface. Bi-directional
information flows between Querying and Dashboard components allow
users to save their preferences and potentially store the relevant
facts that they may have discovered in the Data Store. This would
be crucial, for instance, to enable reinforcement learning for future
projects building upon the CPSS Prototyping Environment, e.g.,
application developments based on the prototype results.
5</p>
    </sec>
    <sec id="sec-10">
      <title>USE CASE</title>
      <p>In this section, we introduce one of our real-world use cases in
the public transport domain (Section 5.1), discuss data exploration
(Section 5.2), describe the construction of the knowledge graph for
the use case in Section 5.3 and illustrate the prototypical
implementation within the CitySPIN platform (Section 5.4).
5.1</p>
    </sec>
    <sec id="sec-11">
      <title>Mobility Use Case Description</title>
      <p>The goal of the CitySPIN project is to deliver a generic platform for
CPSS development that can support a wide variety of use cases in
the context of city infrastructure services. To develop and prototype
this platform, we chose use cases that cover a broad spectrum of
smart city services (viz. public transportation and district heating
network control) while, and on the other hand, exhibiting synergies
in terms of data and component requirements.</p>
      <p>In this paper, we focus on the CitySPIN Event-Aware Mobility
Planning (CaMP) use case, which allows planners at Wiener Linien
(WL) to estimate mobility demands of large-scale events in order to
tailor the mobility planning accordingly. To cater for the needs of
participants of such large-scale events, WL already actively adapts
its transportation network schedule. In particular, the types,
capacities and frequencies of vehicles in service during such events
are currently decided by planners based on historic data about the
number of attendants to recurring events, which are recorded in
event planning protocols saved as .pdf files.</p>
      <p>This current approach makes it dificult to plan for new or
nonrecurring events for which no planning protocols exist.
Additionally, the current planning process does not take into account any
feedback from social sources, e.g., such as event attendant profiles.</p>
      <p>The CitySPIN project addresses this use case with the concept
of Cyber-Physical Social Systems (CPSS), where citizens are seen
as parts of city-wide infrastructures. Therefore, relevant data is
collected from social sensors and data sources that act as proxies
for human behavior (e.g., ticket sales). The relevant data is collected
from a multitude of data sources (e.g., ticket sales, open government
data, mobility data). The resulting Event-Aware Mobility Planner
(CaMP) system enables WL planners to inspect attendance specific
information for a large number of events drawn from a variety of
data sources. It allows integrated and visual access to attendance
data (i) from legacy (historic) sources, (ii) open data sources and
(iii) social data.
5.2</p>
    </sec>
    <sec id="sec-12">
      <title>Data exploration</title>
      <p>To elicit requirements and how they could be addressed with
available data, several workshops were held to (i) review the
organizational and technical context of the real-world use case, (ii) conduct
a high-level survey of available data sources within the use case
partner’s organization as well as externally available data, (iii)
prioritize available data sources and the required data acquisition
methods, (iv) evaluate design alternatives for data acquisition and
semantic enrichment, (v) explore architectural options for a
platform environment that supports integration of large-scale batch
and high-frequency data flows.</p>
      <p>This resulted in a set of preliminary data models, vocabularies,
and guidelines used in the extraction, transformation and
enrichment steps of the knowledge graph construction, as described next.
5.3</p>
    </sec>
    <sec id="sec-13">
      <title>Mobility Knowledge Graph Construction</title>
      <p>The knowledge graph constructed for the mobility use case
covers (i) public transportation infrastructure (e.g., agencies, lines,
schedule), (ii) internal planning protocols from WL, and (iii) event
information.</p>
      <p>Public Transportation Data. The first part of the mobility
knowledge graph covers public transportation data in Vienna.
Transportation data are often available as open data in GTFS format, which
is widely used by Google for their online services7. This data
format covers transport agencies/operators, the routes and the stop
locations, trip schedule, and rules to describe the operation/service.</p>
      <p>
        In the context of our prototype, we rely on the existing GTFS
ontology8 and GTFS CSV converter9 to transform the original GTFS
data provided by the City of Vienna10 to produce our GTFS
transportation KG. In total, the resulted KG contains more than 20
million triples, which is now available online as a SPARQL endpoint11
hosted in an HDT[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] server.
      </p>
      <p>Event Information. In addition to public transportation data, we
include the event information from the Wien-Ticket open data
API12 as the second part of the mobility knowledge graph. The
Wien-Ticket data contains general event information in Vienna,
e.g., event name, address of the event location, and performer’s
name.
7https://gtfs.org
8https://github.com/OpenTransport/linked-gtfs
9https://github.com/OpenTransport/gtfs-csv2rdf
10https://www.data.gv.at/katalog/dataset/wiener-linien-fahrplandaten-gtfs-wien
11http://triple.ai.wu.ac.at
12http://data.opendataportal.at/dataset/wien-ticket-vorverkauf
To extract this data, we implement a data extraction workflow
as a pipeline in UnifiedViews – depicted in Figure 2. The pipeline
consists of three main stages: (i) The first step is the event data
extraction from the open data API, which is originally provided in
CSV format. This process downloads the original data from the API
(#1), translates into RDF (#2 &amp; #3), and merges it with a namespace
graph (#4). (ii) After the event data is transformed into RDF format,
the second step of the extraction performs the linking of the events
and address dataset (#5). (iii) Finally, in the last step, the resulted
linked graph is merged with the original event dataset (#6) and
inserted into a triple store via SPARQL (#7).</p>
      <p>As a result of this process, we extracted more than 2,2 million
triples of event data. We do not yet provide the resulting data as
open data due to server limitations, but we are investigating options
for opening the dataset for public access in the future.</p>
      <p>WL Planning Protocols. The third part of the mobility KG is
transportation planning data, which originated from the internal
transport planning protocols. The data is extracted from WL planning
protocol documents, which are used internally to document
mobility planners’ measures taken in response to demand expectations,
including those due to special event. Such measures include
increasing the frequency of transportation lines that have stations in the
vicinity of the event in a time interval covering the event’s duration.
The planning protocols are typically stored as Word or PDF
documents, which makes automatic data extraction dificult. Parts of the
challenges on this task includes dealing with various irregularities
and inconsistencies in document layouts and extracting
locallyused codes and abbreviations which are embedded within
written comments. To address this issue, we employ a semi-automatic
information extraction pipeline, using a combination of Natural</p>
      <p>Language Processing (NLP) techniques and human computation to
extract the necessary information. In the end, we are able to extract
information from more than 250 out of a set of 300 test planning
documents, which accounts for more than 8,700 triples in total.
We do not plan to make the raw information about this planning
protocol public, as it may contain sensitive internal information.
5.4</p>
    </sec>
    <sec id="sec-14">
      <title>Interactive Planning Support</title>
      <p>To support the mobility use-case, we developed an interactive
planning support tool (cf. Screenshot in Figure 3) by instantiating the
CPSS prototyping environment. The intended user of this tool is
the operation planning department at WL. In particular, the
system is designed to support decisions on measures to optimize the
transportation network in anticipation of a certain event, especially
by taking into account historic records of such measures for the
same type of events or for events that happened at the same or
neighbouring venues.</p>
      <p>We aim to support scenarios in which a transportation planner
needs to decide trafic adjustment measures for an upcoming event.
In this scenario, the planner will start by browsing a list of upcoming
events - as shown by the top left widget in Figure 3 based on second
part of the knowledge graph on event information. From this list,
they then choose a focus event (e.g., "Cirque du Soleil - TOTEM")
for planning adjustments. Based on the selected event, a geo-map
mashup will visualise the location of this event as a green pin on
the map.</p>
      <p>From this point, there are several possibilities for the planner to
choose as follows:
• inspect the list of events that took place in nearby locations
and for which a planning protocol has been produced ("event
planning protocol" widget). This widget draws on the data
extracted from historic planning protocols –which is the
third part of the knowledge graph on event information–
and allows planners to easily access the decisions taken for
the nearby events (e.g., for event "8" the frequency of 3 lines
has been set to 5, 15, 15 minutes respectively).
• identify the public transportation stops in the immediate
geographic vicinity of the focus event’s location ("Nearby
Stops" widget) based on first part of the knowledge graph on
GTFS public transportation data. In our example scenario,
Hermine-Jursa-Gasse and Maria-Jacobi-Gasse are two nearest
stops to the event location.
• browse social media messages related to the event in order
to identify any additional information from social signals
(e.g., general satisfaction with the transportation support
etc.)</p>
      <p>These functionalities for the CaMP use case are made available
by the underlying infrastructures which (i) ensures that data from
various data sources is loaded and semantically integrated so that
it can be (ii) visualised using a visual widget-based platform where
various widget types can be combined into mashups in order to
support the exploration of relevant planning information by the
transportation planners.</p>
      <p>We plan to continue the development of the current CaMP
prototype13 with the integration of additional social data, in particular
data from mobile operators and results from the process mining
components that should also allow planners to get a better
understanding of the social aspect of the CPSS.
6</p>
    </sec>
    <sec id="sec-15">
      <title>RELATED WORK</title>
      <p>
        A number of vision papers explored the applicability of CPSS in
given domains. In the military domain, the CPSS concept fits
naturally by spanning the boundaries of and connecting physical
networks, the cyberspace, mental space and social networks that are
the main components of command and control systems [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. By
integrating these spaces, CPSS bring benefits such as
synchronization across the spaces, self-adaptation and “chaotic control" as an
alternative to precise control in order to deal with inherent
uncertainties in the domain. In manufacturing [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], a new industrial
revolution is emerging enabled by socio-cyber-physical system
(SCPS) which combine social elements with smart manufacturing
thanks to the four technical pillars of Internet of Things (IoT at the
physical layer), Internet of Knowledge (IoK) and Internet of Services
(IoS) at the cyber level, and Internet of People (IoP). A vision of
Physical-Cyber-Social computing enabled by knowledge
technologies and illustrated with an application in the medical domain is
discussed in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Smart City applications inherently subscribe to
the concept of CPSS [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] as we also demonstrate in our own project
with a transportation and a sustainable energy related use case.
      </p>
      <p>Common to CPSS eforts in all domains is that they primarily
focus on describing concrete systems, and how they function. In
CitySPIN, on the contrary, we aim to support the engineering phase
of these systems. A particular focus is on the ETL and data
integration process which takes up considerable efort. Similarly to our
projects, the QROWD project14 also develops semantics based data
integration approaches. However, these do not support
privacyaware data integration as has been done in CitySPIN.</p>
    </sec>
    <sec id="sec-16">
      <title>CONCLUSION AND OUTLOOK</title>
      <p>In this paper, we provided an overview of the CitySPIN CPSSs
platform and development approach focusing mainly on a data
engineering perspective. Using multiple use cases developed with
stakeholders in a city-scale context as a lense to explore challenges
of heterogeneity, privacy, and process dynamics, we motivated the
design of the CitySPIN architecture described in this paper. We
illustrated the prototypical implementation of this architecture by
means of a real-world use case in public transportation planning.</p>
      <p>In future work, we will investigate the integration of more
realtime sensing and actuation components into the platform, which
will enable CPSS developers to integrate additional social
components into the CPSS loop. In the long term, this could facilitate the
implementation of adaptive strategies in various use cases in the
mobility and energy domains.</p>
    </sec>
    <sec id="sec-17">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was funded by the Austrian Research Promotion Agency
FFG under grant 861213 (CitySPIN).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <fpage>2003</fpage>
          .
          <article-title>Basic Geo (WGS84 lat/long) Vocabulary</article-title>
          . (
          <year>2003</year>
          ). https://www.w3.org/
          <year>2003</year>
          / 01/geo/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <year>2012</year>
          .
          <article-title>NeoGeo Geometry Ontology</article-title>
          . (
          <year>2012</year>
          ). http://geovocab.org/geometry
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <year>2012</year>
          .
          <article-title>NeoGeo Spatial Ontology</article-title>
          . (
          <year>2012</year>
          ). http://geovocab.org/spatial
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Ethem</given-names>
            <surname>Alpaydin</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Introduction to machine learning</article-title>
          . MIT press.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Christos</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Cassandras</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Smart Cities as Cyber-Physical Social Systems</article-title>
          . Engineering 2,
          <issue>2</issue>
          (
          <year>2016</year>
          ),
          <fpage>156</fpage>
          -
          <lpage>158</lpage>
          . https://doi.org/10.1016/J.ENG.
          <year>2016</year>
          .
          <volume>02</volume>
          .012
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>David</given-names>
            <surname>Corsar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Milan</given-names>
            <surname>Markovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Edwards</surname>
          </string-name>
          , and
          <string-name>
            <given-names>John D.</given-names>
            <surname>Nelson</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The Transport Disruption Ontology</article-title>
          .
          <source>In The Semantic Web - ISWC</source>
          <year>2015</year>
          ,
          <string-name>
            <given-names>Marcelo</given-names>
            <surname>Arenas</surname>
          </string-name>
          , Oscar Corcho, Elena Simperl, Markus Strohmaier, Mathieu d'Aquin,
          <string-name>
            <given-names>Kavitha</given-names>
            <surname>Srinivas</surname>
          </string-name>
          , Paul Groth, Michel Dumontier, Jef Heflin,
          <source>Krishnaprasad Thirunarayan, and Stefen Staab (Eds.)</source>
          . Vol.
          <volume>9367</volume>
          . Springer International Publishing, Cham,
          <fpage>329</fpage>
          -
          <lpage>336</lpage>
          . https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -25010-6_
          <fpage>22</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Simon</given-names>
            <surname>Cox</surname>
          </string-name>
          , Chris Little,
          <string-name>
            <surname>Jerry R. Hobbs</surname>
            , and
            <given-names>Feng</given-names>
          </string-name>
          <string-name>
            <surname>Pan</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Time Ontology in OWL. W3C Recommendation. W3C</article-title>
          . https://www.w3.org/TR/owl-time/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Richard</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          , Dave Reynolds, and
          <string-name>
            <given-names>Jeni</given-names>
            <surname>Tennison</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The RDF data cube vocabulary</article-title>
          .
          <source>W3C Recommendation. W3C</source>
          . https://www.w3.org/TR/vocab-datacube/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Claudio</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Ciccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Fajar J.</given-names>
            <surname>Ekaputra</surname>
          </string-name>
          , Alessio Cecconi, Andreas Ekelhart, and
          <string-name>
            <given-names>Elmar</given-names>
            <surname>Kiesling</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Finding Non-compliances with Declarative Process Constraints through Semantic Technologies</article-title>
          . In CAiSE Forum. Springer,
          <fpage>60</fpage>
          -
          <lpage>74</lpage>
          . https://doi. org/10.1007/978-3-
          <fpage>030</fpage>
          -21297-
          <issue>1</issue>
          _
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A</given-names>
            <surname>Dix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Finlay</surname>
          </string-name>
          , GD Abowd, and
          <string-name>
            <given-names>R</given-names>
            <surname>Beale</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Human-computer interaction: Pearson prentice hall</article-title>
          . Inc,
          <string-name>
            <surname>England</surname>
          </string-name>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Javier</surname>
            <given-names>D</given-names>
          </string-name>
          <string-name>
            <surname>Fernández</surname>
          </string-name>
          ,
          <string-name>
            <surname>Miguel A Martínez-Prieto</surname>
            , Claudio Gutiérrez, Axel Polleres, and
            <given-names>Mario</given-names>
          </string-name>
          <string-name>
            <surname>Arias</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Binary RDF representation for publication and exchange (HDT)</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>19</volume>
          (
          <year>2013</year>
          ),
          <fpage>22</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>W.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The integration of CPS, CPSS, and ITS: A focus on data</article-title>
          .
          <source>Tsinghua Science and Technology 20</source>
          ,
          <issue>4</issue>
          (
          <year>August 2015</year>
          ),
          <fpage>327</fpage>
          -
          <lpage>335</lpage>
          . https://doi.org/10.1109/TST.
          <year>2015</year>
          .7173449
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Knap</given-names>
          </string-name>
          , Petr Skoda, Jakub Klímek, and Martin Necasky`.
          <year>2015</year>
          .
          <article-title>UnifiedViews: Towards ETL Tool for Simple yet Powerfull RDF Data Management.</article-title>
          .
          <source>In DATESO</source>
          .
          <volume>111</volume>
          -
          <fpage>120</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Mao</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Cyber-Physical-Social Systems for Command and Control</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          <volume>26</volume>
          ,
          <issue>4</issue>
          (
          <year>2011</year>
          ),
          <fpage>92</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Maria</surname>
          </string-name>
          <string-name>
            <surname>Maggi</surname>
          </string-name>
          , Claudio Di Ciccio, Chiara Di Francescomarino, and
          <string-name>
            <given-names>Taavi</given-names>
            <surname>Kala</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Parallel algorithms for the automated discovery of declarative process models</article-title>
          .
          <source>Inf. Syst. 74, Part</source>
          <volume>2</volume>
          (
          <year>2018</year>
          ),
          <fpage>136</fpage>
          -
          <lpage>152</lpage>
          . https://doi.org/10.1016/j.is.
          <year>2017</year>
          .
          <volume>12</volume>
          .002
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Angelika</surname>
            <given-names>Musil</given-names>
          </string-name>
          , Juergen Musil, Danny Weyns, Tomas Bures,
          <string-name>
            <given-names>Henry</given-names>
            <surname>Muccini</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mohammad</given-names>
            <surname>Sharaf</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Patterns for Self-Adaptation in Cyber-Physical Systems. In Multi-Disciplinary Engineering for Cyber-Physical Production Systems</article-title>
          , Stefan Bifl, Arndt Lüder, and Detlef Gerhard (Eds.). Springer International Publishing, Chapter
          <volume>13</volume>
          ,
          <fpage>331</fpage>
          -
          <lpage>368</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Barry</surname>
            <given-names>Norton</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luis M. Vilches</surname>
          </string-name>
          , Alexander De León, John Goodwin, Claus Stadler, Suchith Anand, Dominic Harries, Boris Villazón-Terrazas, and
          <string-name>
            <surname>Ghislain</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Atemezing</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>NeoGeo Vocabulary Specification</article-title>
          . (
          <year>2012</year>
          ). http: //geovocab.org/doc/neogeo/
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sheth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Anantharam</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Henson</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Physical-Cyber-Social Computing: An Early 21st Century Approach</article-title>
          .
          <source>IEEE Intelligent Systems 28, 1 (Jan</source>
          <year>2013</year>
          ),
          <fpage>78</fpage>
          -
          <lpage>82</lpage>
          . https://doi.org/10.1109/MIS.
          <year>2013</year>
          .20
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Tuan-Dat</surname>
            <given-names>Trinh</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Wetz</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ba-Lam</surname>
            <given-names>Do</given-names>
          </string-name>
          , Elmar Kiesling, and
          <string-name>
            <given-names>A Min</given-names>
            <surname>Tjoa</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Distributed mashups: a collaborative approach to data integration</article-title>
          .
          <source>International Journal of Web Information Systems</source>
          <volume>11</volume>
          ,
          <issue>3</issue>
          (
          <year>2015</year>
          ),
          <fpage>370</fpage>
          -
          <lpage>396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Wil</surname>
            <given-names>M. P. van der Aalst. 2016. Process</given-names>
          </string-name>
          <string-name>
            <surname>Mining - Data Science</surname>
          </string-name>
          in Action,
          <source>Second Edition</source>
          . Springer. https://doi.org/10.1007/978-3-
          <fpage>662</fpage>
          -49851-4
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Ma, and
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Vasilakos</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Word of Mouth Mobile Crowdsourcing: Increasing Awareness of Physical, Cyber, and Social Interactions</article-title>
          .
          <source>IEEE MultiMedia 24</source>
          ,
          <issue>4</issue>
          (
          <year>October 2017</year>
          ),
          <fpage>26</fpage>
          -
          <lpage>37</lpage>
          . https://doi.org/10. 1109/MMUL.
          <year>2017</year>
          .4031317
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>G.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Cyberphysical-social system in intelligent transportation</article-title>
          .
          <source>IEEE/CAA Journal of Automatica Sinica</source>
          <volume>2</volume>
          ,
          <issue>3</issue>
          (
          <year>July 2015</year>
          ),
          <fpage>320</fpage>
          -
          <lpage>333</lpage>
          . https://doi.org/10.1109/JAS.
          <year>2015</year>
          .7152667
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Xifan</given-names>
            <surname>Yao</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yingzi</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Emerging manufacturing paradigm shifts for the incoming industrial revolution</article-title>
          .
          <source>The International Journal of Advanced Manufacturing Technology</source>
          <volume>85</volume>
          ,
          <issue>5</issue>
          (
          <issue>01</issue>
          <year>Jul 2016</year>
          ),
          <fpage>1665</fpage>
          -
          <lpage>1676</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>