<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bootstrapping the Publication of Linked Data Streams</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Riccardo Tommasini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamed Ragab</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Falcetta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Della Valle</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sherif Sakr</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano, DEIB</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Tartu, DataSystem Group</institution>
          ,
          <country country="EE">Estonia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data Velocity reached the Web. New protocols and APIs (e.g. WebSockets, and EventSource) are emerging, and the Web of Data is also evolving to tame Velocity without neglecting Variety. The RDF Stream Processing (RSP) community is actively addressing these challenges by proposing continuous query languages and working prototypes. Nevertheless, the problem of Streaming Linked Data publication is still an open challenge. In this paper, we present the rst attempt to tackle this challenge by introducing a set of guidelines to publish streaming linked data by reusing existing resources such as TripleWave, R2RML/RML, VoCaLS, and RSP-QL. We design a publication life-cycle that follows the W3C best practices. Besides, we present an example for publishing the Global Database of Events, Language and Tone (GDELT) as a Streaming Linked Data resource. We open-sourced the code of our resource and made it available for public use.</p>
      </abstract>
      <kwd-group>
        <kwd>RDF Streams</kwd>
        <kwd>Streaming Linked Data</kwd>
        <kwd>RDF Stream Processing</kwd>
        <kwd>Stream Reasoning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1 A new generation of Web applications shows the need to process data reactively
and continuously [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This challenge, known as Data Velocity in the Big Data
domain, is now critical for the Web too [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The RDF Stream Processing (RSP)
tries to process heterogeneous data streams that come from complex domains
on-the- y. Indeed, the RSP community2 proposed the a data model, (i.e., RDF
Streams) a query model (i.e., RSP-QL) and several resources for processing, and
publishing Web streams [
        <xref ref-type="bibr" rid="ref2 ref4 ref9">2, 4, 9</xref>
        ].
      </p>
      <p>The Global Database of Events, Language &amp; Tone (GDELT) project3 is a
family of vast and heterogeneous Web Streams that is considered as the largest
open-access Spatio-temporal archive for human society. Its Global Knowledge
Graph spans more than 215 years, and connects people, organizations, locations,
all over the world. In particular, GDELT complex from a complex domain. It
1 Copyright 2019 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
2 https://www.w3.org/community/rsp/
3 www.gdeltproject.org
captures themes, images, and emotions into a single holistic global network.
GDELT data can be accessed via Google Big Query or via a number of APIs that
run pre-de ned analyses. Therefore, interested researches are forced to either run
their analysis via the pay-per-use service or stick with the provided APIs.</p>
      <p>GDELT Project provisions, every 15 minutes, many TSV Web Streams: the
Event Stream makes use of the dyadic CAMEO format, describing the two
actors participating in each event, and the action they perform; the Mention
Stream records every mention of an event in the Event Stream over time, along
with the timestamp the article was published; the GKG Stream connects
people, organizations, locations, events, and more across the planet. In this paper,
we present our approach for consuming and analyzing the shared GDELT TSV
Streams. In particular, we provide a set of methodological guidelines that explain
how to employ existing RSP resources { RML, TripleWave, VoCaLS { to publish
and consume Streaming Linked Data. The guidelines and an extended version
of the documents included in this paper are available at http://gdelt.stream.
2</p>
      <p>
        Publishing GDELT Streams as Streaming Linked Data
In this section, we present our methodological guidelines to publish GDELT,
an example of RDF streams, as Streaming Linked Data. Figure 1 depicts the
steps of our publication life-cycle that has been designed following the W3C Best
practices4 and the guidelines presented by Hyland et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        The Step (0) Name Things with (HTTP) URIs aims of designing
(HTTP) URIs that identify the relevant resources according to the Linked Data
principles [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and W3C best practices for URI design5. We designed URIs with
the following base http://gdelt.stream/, distinguishing between base
vocabulary (vocab), CAMEO (onto/cameo), instances (ist ) and time instants (time).
      </p>
      <p>
        The Step (1) Model the Streams Domain aims of (i) Understanding
and capturing the domain knowledge into an ontological model, and reuse
existing authoritative vocabularies. (ii) Identifying related resources, and (iii)
Formulating canonical information needs [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This process requires collecting and
reviewing applications, documentations and analyzing sample data. The data of
GDELT is streamed from a multitude of news sources. The extraction process
follows Natural Language Processing techniques and makes use of the Con ict
and Mediation Event Observations (CAMEO) Ontology [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>We collected all the information regarding the schemas of the streams, and
we studied the CAMEO documentations. Then we designed a set of OWL2
ontologies describing the GDELT domain. The CAMEO ontology is a coding scheme</p>
    </sec>
    <sec id="sec-2">
      <title>4 https://www.w3.org/TR/ld-bp/ 5 https://www.w3.org/TR/cooluris/#cooluris</title>
      <p>Bootstrapping the Publication of Linked Data Streams
&lt;&gt; a vocals:StreamDescriptor ; dcat:dataset :eventstream 1</p>
      <p>.
:eventstream a vocals:Stream ;
dcat:title "GDELT Event Stream"^^xsd:string ;
dcat:publisher &lt;http://www.streamreasoning.org&gt; ;
dcat:description "GDELT Events Stream"^^xsd:string ;
vocals:windowType vocals:logicalTumbling ;
vocals:windowSize "PT15M"^^xsd:duration ;
vocals:hasEndpoint [
a vocals:StreamEndpoint ;
dcat:license &lt;https://cc.org/licenses/by-nc/4.0/&gt; ;
dcat:format frmt:JSON-LD;
dcat:accessURL "ws://examples:8080/events" ] .
which is designed for the study of third-party mediation in international
disputes. It contains a hierarchical coding-scheme for dealing with sub-state actors,
event types, and an extensive taxonomy for religious groups and ethnic groups.
For CAMEO, we model event and actors types, into a comprehensive hierarchy.</p>
      <p>
        The Step (2) Describe the Stream aims for providing human-readable
and machine-readable representations of the streams [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] as commonly done for
datasets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To this extent, Tommasini et al. proposed a vocabulary based on
DCAT6 named VoCaLS [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>We used VoCaLS to
describe GDELT streams.
Listing 1.1 shows an
example of description for the
GDELT Event Stream. In
VoCaLS, a Web Stream is
represented using vocals:Stream,
i.e., an unbounded sequence
of data items that might be
accessible on the Web. A Listing 1.1: VoCaLS GDELT Stream
vocals:StreamDescriptor, i.e, description.
a HTTP-accessible document
that contains stream metadata. Finally, the stream content can be consumed
via a vocals:StreamEndpoint, that refer to actual sources using dcat:accessURL
property. GDELT does not use a license format, thus we include a license that
is compliant with the terms of use. We also linked ontologies and mappings
using rdfs:seeAlso. Since the stream is published regularly as 15 minutes batch,
we include metadata about the rate, i.e., at lines 6-7 vocals:windowType and
vocals:windowSize.</p>
      <p>
        The Step (3) Convert
to RDF Stream aims of
enabling RDF data
provisioning so that data could be
enriched with domain
knowledge [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] from Step (1). The
conversion to RDF can occur
either automatically or via
expert modeling [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Approaches
based on mapping languages Listing 1.2: Subset of GDELT RML Mapping.
decouple conversion and
modelling, and support the conversion process using a formal language like R2RML
7 for relational data, and RML for logical sources [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] sharing document-based
formats (e.g., JSON), or semi-structured formats (e.g., CSV). Similarly, Web
streams are often not shared as RDF Streams and needs to be converted. In
practice, GDELT content is not natively RDF. Therefore, we need to set up
a conversion mechanism. For data conversion, we followed the RML approach
where CSV data is converted using its mappings. Listing 1.2 shows a sample of
&lt;GEM&gt; a rr:TriplesMap ; rml:logicalSource &lt;source&gt; ; 1
rr:subjectMap [ 2
rr:template "http://gdelt.stream/ist/{GLOBALEVENTID}"; 3
rr:class gdelt:Event; 4
rr:graphMap [ 5
rr:template "http://gdelt.stream/time/{DATEADDED}"]]; 6
rr:predicateObjectMap [ rr:predicate rdf:type; 7
rr:objectMap [ 8
rr:template "http://gdelt.stream/cameo/{EventCode}" 9
      </p>
      <p>;]];
rr:predicateObjectMap [
rr:predicate gdelt:actor;
rr:objectMap [ rr:parentTriplesMap &lt;Actor1TM&gt; ]];...</p>
    </sec>
    <sec id="sec-3">
      <title>6 https://www.w3.org/TR/vocab-dcat/ 7 https://www.w3.org/2001/sw/rdb2rdf/r2rml/</title>
      <p>Tommasini et al.
such R2RML mapping. Notably, we used rr:class to assign the gdelt:Event
type to the data at line 4. Moreover, we assign the CAMEO type using rdf:type.</p>
      <p>
        The Step (4) Publish The RDF Stream aims for serving the data to
the audience of interest. Datasets { opportunely described with contextual
vocabularies [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] { are either added to the Linked Open Data (LOD) Cloud, shared
using REST APIs, or exposed via SPARQL endpoints. Critical aspects of this
step are licensing, audit, and access control [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        The unbounded nature of Streaming Linked Data makes it impossible to
directly add a data stream to the LOD Cloud [
        <xref ref-type="bibr" rid="ref10 ref9">10, 9</xref>
        ]. However, using an RSP
Engine that focuses on RDF Stream provisioning like TripleWave [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], one can
publish a stream describing using VoCaLS. TripleWave relies on Barbieri et al's
vision [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] of serving a "static" named graph called S-Graph via REST API, to
identify and describe the streams. Moreover, they envisioned that the stream
elements of an RDF Graph are called I-Graphs. We shared the VoCALS
description les like the one in Listing 1.1 as S-GRAPH via REST APIs. Since the
concept of event is rst-class, we opted for a graph-based stream data model.
      </p>
      <p>
        Conclusion
This paper presents a rst attempt to publish streaming Linked Data following
W3C's best practices [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Some real-world Web streams from the GDELT project
are available as RDF Stream, described using VoCaLS, and accessible at http:
//gdelt.stream. Future work include publishing other Web streams towards a
catalog of Linked Streams.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hausenblas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Describing linked datasets</article-title>
          .
          <source>In: Proceedings of the WWW2009 Workshop on Linked Data on the Web, LDOW</source>
          <year>2009</year>
          , Madrid, Spain, April
          <volume>20</volume>
          ,
          <year>2009</year>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barbieri</surname>
            ,
            <given-names>D.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Della Valle</surname>
          </string-name>
          , E.:
          <article-title>A proposal for publishing data streams as linked data - A position paper</article-title>
          .
          <source>In: Proceedings of the WWW2010 Workshop on Linked Data on the Web, LDOW</source>
          <year>2010</year>
          ,
          <article-title>Raleigh</article-title>
          , USA, April
          <volume>27</volume>
          ,
          <year>2010</year>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Della</given-names>
            <surname>Valle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Dell'Aglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Margara</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Taming velocity and variety simultaneously in big data with stream reasoning: tutorial</article-title>
          . In: DEBS (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>DellAglio</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Della</given-names>
            <surname>Valle</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>van Harmelen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Stream reasoning: A survey and outlook</article-title>
          .
          <source>Data Science</source>
          <volume>1</volume>
          (
          <issue>1-2</issue>
          ),
          <volume>59</volume>
          {
          <fpage>83</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dimou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sande</surname>
            ,
            <given-names>M.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Slepicka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szekely</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mannens</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoblock</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de</surname>
            <given-names>Walle</given-names>
          </string-name>
          , R.V.:
          <article-title>Mapping hierarchical sources into RDF using the RML mapping language</article-title>
          .
          <source>In: International Conference on Semantic Computing</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gerner</surname>
            ,
            <given-names>D.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schrodt</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yilmaz</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abu-Jabr</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Con ict and mediation event observations (cameo)</article-title>
          .
          <source>International Studies Association</source>
          , New Orleans (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hyland</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>The joy of data-a cookbook for publishing linked government data on the web</article-title>
          .
          <source>In: Linking government data</source>
          , pp.
          <volume>3</volume>
          {
          <fpage>26</fpage>
          . Springer (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Margara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urbani</surname>
          </string-name>
          , J., van
          <string-name>
            <surname>Harmelen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bal</surname>
            ,
            <given-names>H.E.</given-names>
          </string-name>
          :
          <article-title>Streaming the web: Reasoning over dynamic data</article-title>
          .
          <source>J. Web Sem</source>
          .
          <volume>25</volume>
          ,
          <issue>24</issue>
          {
          <fpage>44</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mauri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calbimonte</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dell'Aglio</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balduini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brambilla</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Della</given-names>
            <surname>Valle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Aberer</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Triplewave: Spreading RDF streams on the web</article-title>
          .
          <source>In: ISWC</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tommasini</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sedira</surname>
            ,
            <given-names>Y.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dell'Aglio</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balduini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ali</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Phuoc</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Della Valle</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calbimonte</surname>
          </string-name>
          , J.:
          <article-title>Vocals: Vocabulary and catalog of linked streams</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>