<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Linking of open and private data in dataspace: A case study of air quality monitoring and forecasting</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alex Acquier</string-name>
          <email>alex.acquier@universityofgalway.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andy Donald</string-name>
          <email>andy.donald@universityofgalway.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edward Curry</string-name>
          <email>edward.curry@universityofgalway.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ihsan Ullah</string-name>
          <email>ihsan.ullah@universityofgalway.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Umair ul Hassan</string-name>
          <email>umair.ulhassan@universityofgalway.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight SFI Research Centre for Data Analytics, University of Galway</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>JE Cairnes School of Business and Economics, University of Galway</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computer Science, University of Galway</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Air pollutant monitoring and its eficient visualisation can support accurately assessing the air quality and harmful emissions; and it can guide us towards potential mitigation strategies to reduce its impact on public health and our environment. This paper presents a case study of employing dataspaces and proposing an ontology for modelling mobility to address the challenges posed by the heterogeneity of data sources in environmental monitoring, as well as using machine learning for forecasting pollutants. We employ Linked data as a powerful paradigm for harmonising and interlinking diverse and publicly available environmental data with private company data to create a dataspace for environmental monitoring. By applying semantic technologies and ontological modelling to integrate heterogeneous data, our approach fosters data interoperability and facilitates enhanced data exploration and decision support. For decision support, we demonstrate the utility of integrated data for forecasting air pollutants with the help of models developed using machine learning. Finally, a spatio-temporal visualisation platform harnesses the power of semantic relationships and contextual enrichment to support data exploration.</p>
      </abstract>
      <kwd-group>
        <kwd>linked data</kwd>
        <kwd>dataspaces</kwd>
        <kwd>machine learning</kwd>
        <kwd>spatiotemporal data</kwd>
        <kwd>environmental monitoring</kwd>
        <kwd>air pollution</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Advances in digital technologies have ushered in an era of unprecedented data generation
across various domains. Environmental monitoring, in particular, has significantly improved
data collection from diverse sources such as remote sensors, satellite imagery, social media
platforms, commercial databases and government databases [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This wealth of information
provides invaluable insights into our environment and weather. It helps in creating information
systems that aid in decision-making processes, policy formulation, and resource allocation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
The breakthroughs in machine learning and artificial intelligence have led to renewed interest
in forecasting environmental and weather conditions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, the heterogeneity of data
CEUR
Workshop
Proceedings
sources, characterised by diferences in formats, structures, and semantics, poses significant
challenges to their efective integration and utilisation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        To address these challenges, the concept of Linked data has emerged as a powerful paradigm
for harmonising and interlinking data from disparate sources [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Linked data fosters a
decentralised approach to data integration, wherein each data source is assigned a unique identifier
and linked to related information using standardised ontologies and vocabularies [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This
approach enables seamless data interoperability and facilitates discovering hidden relationships
and patterns that might remain obscured when data sources are treated in isolation. While
Linked data focuses on using the Web standards and standardised ontologies to make data
available, a dataspace aims to develop the necessary set of services to enable a custom view, for
data consumers, of heterogeneous data sources which diferent data controllers manage [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>This paper presents a novel approach in environmental monitoring, which is to create a
dataspace using the Linked data approach. Our proposed dataspace acts as a linked platform
where heterogeneous environmental data sources are integrated, interconnected, and made
accessible in an iterative yet coherent manner. By utilising semantic technologies and ontological
modelling, the dataspace allows for the easy integration of data originating from sources as
diverse as mobile weather platforms, remote sensing platforms, meteorological websites, and
governmental statistical ofices. Towards this end, the primary objective of this paper is to
showcase the feasibility and advantages of employing Linked data principles in environmental
monitoring. We will discuss the intricacies of data source integration, ontology development
and data inter-linkage within the context of the environmental monitoring. Furthermore, we
will demonstrate how the dataspace fosters enhanced data exploration and decision support by
harnessing the power of semantic relationships and contextual enrichment using appropriate
visualisations and forecasting models. A vital aspect of this dataspace is the use of
spatiotemporal data features as the anchoring point of all information integration and utilisation
processes.</p>
      <p>In the subsequent sections, we will delve into the technical aspects of the Linked data approach,
elucidating the methods used for data extraction, transformation, and integration. We will also
highlight the challenges encountered during the process and the strategies to overcome them.
Additionally, we will provide case studies and real-world examples to underscore the practical
implications of the proposed dataspace in aiding environmental research, policy-making, and
public awareness.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Background</title>
      <sec id="sec-3-1">
        <title>2.1. Environment and Climate Data Sources</title>
        <p>
          The internet hosts diverse open data sources for environmental monitoring, making it a valuable
resource for researchers, policymakers, and the public [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. These open datasets cover various
environmental factors, including air quality, climate, water quality, biodiversity, etc. One
prominent example is the European Environmental Agency (EEA), which provides extensive
environmental data, including weather observations, climate records, and satellite imagery, all
accessible to the public for various applications, from climate research to weather forecasting [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
Besides the EEA, each European country has national organisations, such as the Environmental
Protection Agency (EPA) and the Irish Meteorological Service (Met Éireann). In addition to open
data, private companies also contribute valuable information for environmental monitoring.
For instance, technology companies like Google capture real-time data from their mapping
and navigation services, which can be harnessed to create custom views of environmental
data. By analysing people’s movement and trafic patterns, this data can be used to assess
air pollution, identify congested areas, or monitor urban heat islands. Such private data can
enhance our understanding of urban environments and enable more customised approaches to
address environmental challenges. It is a valuable complement to open datasets in pursuing
sustainable and eco-friendly solutions [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Common European Dataspaces</title>
        <p>
          Since 2020, the European Commission has proposed establishing several dataspaces for
sectorspecific data exchange, sharing and pooling [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Such dataspaces can allow organizations to
create dynamic and on-demand custom views over heterogeneous and distributed data sources,
including the organization’s protected data, private data from its partners, and publicly available
open data [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Concerning environmental monitoring, two dataspaces are particularly relevant:
the Mobility dataspace and the Green Deal dataspace. The Mobility dataspace focuses on the
transportation sector where data from transport systems, trafic monitoring systems, transport
companies, etc. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] is collected, and insights are provided. When combined with the data from
EEA, such data can provide a more granular analysis of the impact of transport on climate.
More importantly, the Green Deal dataspace aims to provide a set of common infrastructure
and services that will facilitate easy access to interoperable data related to climate, environment
and sustainability across Europe [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. While both these dataspaces aim to provide technical
and legal guidance regarding sharing and exchanging data, the core challenge organisations
still face is building their local services and interfaces over such data sources.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Dataspace for Environmental Monitoring</title>
      <p>
        A pay-as-you-go approach to creating a custom view of environmental data leverages the variety
of available open and private data in an incremental, flexible, and cost-efective manner [
        <xref ref-type="bibr" rid="ref13 ref4">4, 13</xref>
        ].
This approach is particularly relevant in environmental monitoring, where the diverse data
sources are often voluminous and specialised. Rather than maintaining extensive infrastructure
and data repositories, organisations and individuals can tap into open data resources like the
EEA’s climate data or Google’s real-time location data as needed. This approach allows users to
access the precise data they require, paying only for the specific resources and processing power
necessary to create custom views. For instance, if a researcher needs real-time information on
air quality in an urban area to study the impact of trafic patterns, they can harness private data
sources like Google’s mobility data alongside open datasets. By doing so, they can tailor their
data processing and analysis to their project’s scope, optimising costs and resource utilisation.
This flexibility empowers many stakeholders to engage in environmental monitoring, ultimately
fostering sustainable practices and decision-making.
      </p>
      <p>
        To establish a dataspace, a set of services plays a pivotal role in creating custom views of
environmental data when following a pay-as-you-go approach [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. As illustrated in Figure 1,
Visualizations and Dashboards
      </p>
      <p>results</p>
      <p>Semantic
Integration &amp;</p>
      <p>Querying
Data Catalogue query
query</p>
      <p>Machine</p>
      <p>
        Learning
metadata
query
query model
model
Private Data
Public Data
these services encompass a range of capabilities and tools that facilitate data access, integration,
analysis, and visualisation. First and foremost, data discovery services are essential, as they help
users identify relevant datasets from the vast pool of public and private sources. These services
provide metadata and cataloguing information, making finding the data that suits specific
monitoring needs easier. In the absence of automated discovery services, a more practical
approach is to create a data catalogue of known datasets and data sources. This cataloguing
service serves as a canonical sources of metadata about data sources, and it is improved and
updated overtime to facilitate current and future requirements for data discovery [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>Once the correct set of data sources has been identified, a set of services for querying and
semantic integration come into play. These services enable the harmonisation and blending
of diverse data sources, ensuring compatibility and consistency across heterogeneous sources
when accessing data. For instance, if one is creating a custom view of air quality data by
combining open data from the EEA and EPA and private location data, data integration services
help reconcile diferent data formats and units of measurement. Besides semantic integration,
data querying services are equally vital to facilitate further processing and analysis. They allow
users to apply various algorithms and statistical methods to extract valuable insights from the
integrated datasets. In the context of environmental monitoring, this may involve calculating
pollution trends, identifying hot-spots, or predicting future environmental conditions. Finally,
dataspace services provide data visualisation and presentation tools to visualise and share the
results efectively. These non-core services enable the creation of custom dashboards, maps, and
reports to communicate findings to stakeholders, researchers, or the public in a user-friendly
and actionable manner.</p>
      <p>Overall, dataspace services bridge the wealth of available data and create custom views,
facilitating the pay-as-you-go approach that allows users to tailor their data usage to their
specific requirements while optimising eficiency, accuracy, and cost-efectiveness. To match
the information requirements of custom views with the syntax and semantics of underlying
data sources, applying some form of standardisation and semantic mapping across schemas and
entities of data sources becomes imperative. Furthermore, any required statistical analytics will
be used to present the data and its analysis better.</p>
    </sec>
    <sec id="sec-5">
      <title>4. A Case Study of Air Pollutant Monitoring</title>
      <p>This section presents a case study employing the dataspace approach and federated learning
for air quality monitoring and forecasting. Table 1 shows a list of data sources used to create
the environmental dataspace for pollutant monitoring. In addition, Pollutrack has emerged
as a trailblazer in environmental data collection by implementing a sophisticated approach to
monitoring air quality. Recognizing the critical importance of understanding and mitigating air
pollution, Pollutrack has partnered with DPD to strategically deploy a combination of fixed and
mobile platforms for comprehensive data collection in the city of Dublin. The fixed platforms,
strategically positioned in urban centres and industrial zones, serve as constant monitoring hubs,
capturing baseline air quality metrics over extended periods. Complementing this platform
network, DPD introduced a fleet of mobile platforms equipped with the latest sensors on their
delivery vehicles. These vehicles traverse diverse regions within Dublin, providing real-time
data on the go. This dynamic approach ofers a nuanced understanding of pollution dynamics
influenced by various environmental factors.</p>
      <p>In the following subsections, we will describe the proposed ontology and its integration, data
quality and pre-processing, an overview of the machine learning-based pollutant forecasting
module, and finally, our spatiotemporal visualisation.</p>
      <sec id="sec-5-1">
        <title>4.1. Ontology Development and Integration</title>
        <p>
          To enable linkages between data across multiple sources, the first step involved the definition
of an ontology to define entities and their relationships [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The objective of the Mobility in
Cities ontology, as illustrated in Figure 2, is to establish a standard that represents how
information about environmental monitoring can be expressed via a set of classes and relationships.
Following the Semantic Web standards, this ontology is designed to support various concepts
associated with environmental monitoring and mobility across time and space in an urban
environment.
        </p>
        <p>Core concepts: At the core of the ontology, we have an observation that can be described,
measured and observed. Each observation can have a location and time of when and where the
observation occurred. We express a measurement of an observation as a phenomenon which
needs to be observed. These phenomena can have multiple types of measurements, such as
  2.5 (defined as particles that are 2.5 microns or less in diameter), trafic, weather or airport
activity. Similarly, an observation can be observed by a sensor where a sensor is a device e.g. an
IoT or edge sensor with a particular measurement capability with defined features. The sensor
entity can implement sensing, which has a specific sensor output that can be considered an
observation of a defined phenomenon.</p>
        <p>
          Linkage with other ontologies: The ontology is designed to describe issues related to urban
environments, such as air quality, water quality, trafic volume, etc. Following the principles of
Linked data, some concepts within this ontology are defined using existing and well-known
ontologies in the Semantic Web. For instance, the HealthRisk Ontology [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] is used to describe
an observation of air quality. Similarly, the concepts related to time and location have been
defined using existing Time and GEO ontologies, respectively. A key benefit of such linkages
with existing ontologies is supporting the semantic integration of data fetched from multiple
sources.
        </p>
        <p>The ontology design was carefully developed to allow flexibility for future extension to cover
diferent urban mobility observational phenomena. For instance, we detail air quality
measurements from sensors capturing trafic volume at trafic lights. Adding sub-classes covering data
sources from, for example, public transport or cycling infrastructure can significantly enhance
the ontology’s impact. With the development of this dataspace, we anticipate it will support
several future implementations using public data sources, which will validate our hypothesis
and support continual extensions to the ontology.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Data Quality &amp; Pre-processing</title>
        <p>The raw data about the air quality was not usable straightaway due to (i) the wide area covered
by the mobile platforms, (ii) some readings recorded outside the normal times of business, (iii)
issues with the sensor identification scheme, and (iv) diferences of data recording intervals
between the diferent platforms (v) missing data due to breakdown of sensors or other reasons.
To address these issues and improve data quality, a set of pre-processing steps were applied
before any machine learning tasks. The air pollutant data is filtered through to keep only an
area of 1515 kilometres that encompasses the centre of Dublin (latitude ranging from 53.2821
to 53.417, longitude ranging from -6.377 to -6.15065) for both fixed and mobile platforms data
and sorted into three diferent spatial granularity. Each of the following datasets was used in a
series of tests to allow the predictions of particulate matter at both   2.5 and   10 sizes for all
spatial subdivisions.</p>
        <p>• Global (33 ): 9 squares with 5 kilometres sides
• Local (55 ): 15 squares with 3 kilometres sides
• Hyper-local (1515 ): 175 squares with 1-kilometre sides
Data standardisation: In the first step, new identifiers were assigned to the diferent
platforms, which made it more easily human readable and helped for the display in applications
and debugging. All data was homogenised with a window size of one hour. For the fixed data,
the values of the readings were added up and divided by the number of records to give the
ifnal results per hour for each platform concerned. For the mobile data, each reading was kept
untouched concerning the values and given the timestamp associated with the original reading
time, which means that the same timestamps can appear several times for the platforms. Still,
the uniqueness is ensured by the coordinates where the reading has been taken. It was found
that the mobile platforms were used far outside the area where the fixed platforms are located,
and those data points were removed. Also, only the readings were taken between 8am and 8pm,
as these are the hours when mobile platforms should be out for data collection.
Outlier detection: The Local Outlier Factor (LOF) algorithm is an unsupervised anomaly
detection method which computes the local density deviation of a given data point concerning
its neighbours. It considers the samples with a substantially lower density than their neighbours
as outliers. The LOF method was used by looking into the relation between   2.5 and   10
values using 1000 neighbours; between 1% to 3% of the data points were deemed outliers and
were removed from the final datasets.</p>
        <p>
          Data imputation: Dealing with missing values is one of the most common data quality
requirements in environmental data processing, and previously, several data imputation methods
for missing air quality are hyper-local level [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. When the particulate matter data was mapped
according to three levels of granularity, not all the squares created were necessarily populated
(i.e., the square either had no fixed platform around it or was located outside the urban area).
To overcome this problem, the values of the neighbouring squares were used to extrapolate
the missing data by adding the values of all the neighbour-populated squares for a given
timestamp. The result was divided by the number of neighbour-populated squares. After this
data imputation, both fixed and mobile data for datasets of both types of particulate matter was
integrated with weather data (e.g., dry and wet bulb temperatures, dew point, and atmospheric
water content) for pollutant forecasting and visualization.
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>4.3. Pollutant Forecasting with Machine Learning</title>
        <p>
          Machine learning, specifically its deep learning sub-branch, has shown promising results for
various applications, e.g., autonomous vehicles and the medical domain. Similarly, it is used
to predict and forecast air pollution. In this work, a hybrid model [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] resulting from the
combination of a conv layer, LSTM layer, and an attention-based layer is adapted to be used in
a federated learning approach [
          <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
          ] to forecast pollutants (  2.5 and   10) in the air. The
model is trained on data from Dublin City. It consists of both private pollutant data collected by
DPD and public data collected by the Irish Meteorological Service (Met Éireann).
        </p>
        <p>In time series data, the sliding window approach is normally adopted to select the input for
predicting the next value/values. In this scenario, the window can comprise diferent time lags,
i.e., the number of previous hours of data necessary to generate the predictions for the next
second/hour/day. We chose 6-time lags, i.e., 2, 4, 6, 8, 10, and 12 hours, to experiment at diferent
granularities. The RMSE and MAE are used to assess the model’s performance. In addition, to
divide and select only specific region data, we divided the regions with in three regions called
global 33 , local 55 , and hyper local 1515 as being defined before in section 4.2.</p>
        <p>Table 2 shows that a decrease in the values of MAE and RMSE occurs as the number of
divisions increases. When the data is grouped within a 33 granularity, the best time lag found
is 12 hrs for both   2.5 and   10. The 55 granularity shows diferent optimum results for
  2.5 (12hrs) and   10 (8hrs) but given the small increase for RMSE and MAE in the 12hrs
results compared the 8hrs result for   10. The 1515 shows diferent optimum results for   2.5
(8hrs) and   10 (4hrs), but given the small increase for RMSE and MAE in the 8hrs results
compared to the 4hrs result for   10. It was observed that for all time lags, as the granularity
increases, both RMSE and MAE values decrease for both   2.5 and   10.</p>
      </sec>
      <sec id="sec-5-4">
        <title>4.4. Spatio-temporal Visualisation</title>
        <p>Besides forecasting, the objective is to create custom views so that any user would be able
to visualise the clean and integrated raw data with the help of appropriate visualizations and
graphing tools. For this purpose, a web-based visualization tool was created that allowed a user
to select a geo-spatial area of interest and time frame for viewing data on a graph. To help with
the selection of an area time frame, the location of each fixed platform is shown on a map (see
Figure 3) and when the mouse hovers on one of the locations, the coordinates, first and last
reading timestamps are shown.</p>
        <p>To query the relevant Linked data, users are asked to provide the maximum and minimum
latitude, longitude and dates as well as the type of pollutant (see Figure 4). The dates are used to
ifrst filter the mobile data, it is then further refined using the latitude and longitude coordinates
and the values are averaged for each each hour of the designed time frame and over the selected
area. The fixed platform identifiers within the search area’s compound are determined using the
metadata of fixed platforms. Once the identifiers are found, they are used in concert with the
start and end dates to query data. These processes are used for any of the types of pollutant(s)
selected (  2.5 and/or   10). Once this is over, the data of each fixed platform within the
designated area and the associated averaged data are plotted (see Figure 5) and their associated
mean and standard deviation are displayed in a table below the graph for each pollutant type
requested.</p>
        <p>As shown in Figure 5, the machine learning forecast can be added to this time-series
visualisation. It can allow the user to see future forecasts of pollutants in a specific area. Similar
visualizations for descriptive and predictive analytics will be added as part of the future work.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusion and Future Research</title>
      <p>This paper presents a novel approach combining public and private data to create an on-demand
dataspace for the environment. Our proposed dataspace acts as a linked platform where
heterogeneous environmental data sources are integrated using semantic technologies and
ontological modelling. The utility of the proposed data is demonstrated with the help of a case
study that combines pollutant data from a company with publicly available weather data to
create interactive visualizations and a pollutant forecasting model. Both public and private data
are relevant to pollution monitoring in Dublin city, and the outputs of this research work can
help in achieving the Net Zero 20501 target of the Government of Ireland. Furthermore, this</p>
      <sec id="sec-6-1">
        <title>1https://www.gov.ie/en/press-release/16421-climate-action-plan-2021-securing-our-future/</title>
        <p>work is aligned with UN’s SDG 132 on climate action.</p>
        <p>
          To extend the work presented in this paper, several challenges still require further
investigation. One of those challenges is the data quality. For instance, if there are significant time gaps
in the data, how can we address these issues? Various imputation techniques can be adopted,
including but not limited to averaging or EM algorithm [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Such techniques can be further
extended with more intelligent mechanisms to fill the gap closer to the original data. In this
regard, a proposed work could utilise public data to pre-train a base model and then use it to
ifll the gaps in private data. Another challenge is the integration of public and private data
into data provenance. Providing details of the sources of data and its lineage can help improve
the opacity of both the visualization and machine learning models. For this purpose, existing
ontologies such as PROV-DM3 and Dublin Core4 can be used to generate provenance metadata
[
          <xref ref-type="bibr" rid="ref22">22</xref>
          ].
        </p>
        <p>
          Creating complex artificial intelligence services over combined public and private brings its
own set of challenges. One of those challenges is to deal with diferent scopes of data privacy
and protection applicable to each source dataset [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. One approach to address this issue is
to follow a federated learning approach and train models individually on each data source
[
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. This approach allows the building of aggregate models without explicit integration of
heterogeneous datasets.
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This publication has emanated from research conducted with the financial support of Science
Foundation Ireland under Grant Number 12/RC/2289_P2 - Insight SFI Centre at the University
of Galway. For the purpose of Open Access, the authors have applied a CC BY public copyright
licence to any Author Accepted Manuscript version arising from this submission.</p>
      <sec id="sec-7-1">
        <title>2https://sdgs.un.org/goals/goal13 3https://www.w3.org/TR/prov-dm/ 4https://www.dublincore.org/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Messer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zinevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Alpert</surname>
          </string-name>
          ,
          <article-title>Environmental monitoring by wireless communication networks</article-title>
          ,
          <source>Science</source>
          <volume>312</volume>
          (
          <year>2006</year>
          )
          <fpage>713</fpage>
          -
          <lpage>713</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N. P.</given-names>
            <surname>Melville</surname>
          </string-name>
          ,
          <article-title>Information systems innovation for environmental sustainability</article-title>
          ,
          <source>MIS quarterly</source>
          (
          <year>2010</year>
          )
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I.-I.</given-names>
            <surname>Prado-Rujas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Serrano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García-Dopico</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Córdoba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <article-title>Combining heterogeneous data sources for spatio-temporal mobility demand forecasting</article-title>
          ,
          <source>Information Fusion</source>
          <volume>91</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E.</given-names>
            <surname>Curry</surname>
          </string-name>
          ,
          <article-title>Real-time linked dataspaces: Enabling data ecosystems for intelligent systems</article-title>
          , Springer Nature,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Curry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Derguech</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kouroupetroglou</surname>
          </string-name>
          , U. ul
          <string-name>
            <surname>Hassan</surname>
          </string-name>
          ,
          <article-title>A real-time linked dataspace for the internet of things: enabling “pay-as-you-go” data management in smart environments</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>90</volume>
          (
          <year>2019</year>
          )
          <fpage>405</fpage>
          -
          <lpage>422</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <article-title>Linked data: Evolving the web into a global data space</article-title>
          , Springer Nature,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Franklin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Halevy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maier</surname>
          </string-name>
          ,
          <article-title>From databases to dataspaces: a new abstraction for information management</article-title>
          ,
          <source>ACM Sigmod Record</source>
          <volume>34</volume>
          (
          <year>2005</year>
          )
          <fpage>27</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gemeinholzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Treloar</surname>
          </string-name>
          ,
          <article-title>Open data in global environmental research: The belmont forum's open data survey</article-title>
          ,
          <source>PloS one 11</source>
          (
          <year>2016</year>
          )
          <article-title>e0146695</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zuiderwijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Janssen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Poulis</surname>
          </string-name>
          , G. van de Kaa,
          <article-title>Open data for competitive advantage: insights from open data use by companies</article-title>
          ,
          <source>in: Proceedings of the 16th annual international conference on digital government research</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>European</given-names>
            <surname>Comission</surname>
          </string-name>
          ,
          <article-title>A european strategy for data</article-title>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pretzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Drees</surname>
          </string-name>
          , L. Rittershaus,
          <article-title>Mobility data space</article-title>
          , in: B.
          <string-name>
            <surname>Otto</surname>
          </string-name>
          , M. ten
          <string-name>
            <surname>Hompel</surname>
          </string-name>
          , S. Wrobel (Eds.),
          <article-title>Designing Data Spaces: The Ecosystem Approach</article-title>
          to Competitive Advantage, Springer International Publishing,
          <year>2022</year>
          , pp.
          <fpage>343</fpage>
          -
          <lpage>361</lpage>
          . doi:
          <volume>10</volume>
          .1007/978- 3-
          <fpage>030</fpage>
          - 93975- 5_
          <fpage>21</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M. Gutierrez</given-names>
            <surname>David</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dietrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Raczko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Denvil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Santoro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chatzikyriakou</surname>
          </string-name>
          , W. Borejko,
          <article-title>Towards the european green deal data space</article-title>
          ,
          <source>in: EGU General Assembly Conference Abstracts</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>EGU</fpage>
          -
          <volume>8788</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>U.</given-names>
            <surname>ul Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Derguech</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hannon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cliford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kouroupetroglou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Smit</surname>
          </string-name>
          , E. Curry,
          <article-title>Water analytics and management with real-time linked dataspaces</article-title>
          ,
          <source>in: Government 3</source>
          .
          <fpage>0</fpage>
          -
          <string-name>
            <given-names>Next</given-names>
            <surname>Generation Government Technology Infrastructure</surname>
          </string-name>
          and Services: Roadmaps,
          <string-name>
            <given-names>Enabling</given-names>
            <surname>Technologies</surname>
          </string-name>
          &amp; Challenges, Springer,
          <year>2017</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>196</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Halevy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Franklin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maier</surname>
          </string-name>
          ,
          <article-title>Principles of dataspace systems</article-title>
          ,
          <source>in: Proceedings of the twenty-fith ACM SIGMOD-SIGACT-SIGART symposium on Principles of database systems</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>U.</given-names>
            <surname>ul Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ojo</surname>
          </string-name>
          , E. Curry,
          <article-title>Catalog and entity management service for internet of things-based smart environments, in: Real-time Linked Dataspaces: Enabling Data Ecosystems for Intelligent Systems</article-title>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>X.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          , G. Song, S. Ma, S. Hu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>An ontology-driven approach for integrating intelligence to manage human and ecological health risks in the geospatial sensor web</article-title>
          ,
          <source>Sensors</source>
          <volume>18</volume>
          (
          <year>2018</year>
          )
          <fpage>3619</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Junninen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Niska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tuppurainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ruuskanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kolehmainen</surname>
          </string-name>
          ,
          <article-title>Methods for imputation of missing values in air quality data sets</article-title>
          ,
          <source>Atmospheric environment 38</source>
          (
          <year>2004</year>
          )
          <fpage>2895</fpage>
          -
          <lpage>2907</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Urban</surname>
          </string-name>
          <year>pm2</year>
          .
          <article-title>5 concentration prediction via attention-based cnn-lstm</article-title>
          ,
          <source>Applied Sciences</source>
          <volume>10</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .3390/app10061953.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>G.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <article-title>Federated learning-based multi-energy load forecasting method using cnn-attention-lstm model</article-title>
          ,
          <source>Sustainability</source>
          <volume>14</volume>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .3390/su141912843.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>I.</given-names>
            <surname>Ullah</surname>
          </string-name>
          , U. ul Hassan,
          <string-name>
            <surname>M. I.</surname>
          </string-name>
          <article-title>Ali, Multi-level federated learning for industry 4.0 - a crowdsourcing approach</article-title>
          ,
          <source>Procedia Computer Science</source>
          <volume>217</volume>
          (
          <year>2023</year>
          )
          <fpage>423</fpage>
          -
          <lpage>435</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.procs.
          <year>2022</year>
          .
          <volume>12</volume>
          .238.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>W.</given-names>
            <surname>Junger</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. P. De Leon</surname>
          </string-name>
          ,
          <article-title>Imputation of missing data in time series for air pollutants</article-title>
          ,
          <source>Atmospheric Environment</source>
          <volume>102</volume>
          (
          <year>2015</year>
          )
          <fpage>96</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>D. L.</given-names>
            da
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Batista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. L.</given-names>
            <surname>Correa</surname>
          </string-name>
          ,
          <article-title>Data provenance in environmental monitoring</article-title>
          ,
          <source>in: 2016 IEEE 13th International Conference on Mobile Ad Hoc and Sensor Systems (MASS)</source>
          , IEEE,
          <year>2016</year>
          , pp.
          <fpage>337</fpage>
          -
          <lpage>342</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ré</surname>
          </string-name>
          ,
          <article-title>Reasoning over public and private data in retrieval-based systems</article-title>
          ,
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>11</volume>
          (
          <year>2023</year>
          )
          <fpage>902</fpage>
          -
          <lpage>921</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>