<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Roomba: An Extensible Framework to Validate and Build Dataset Pro les</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ahmad Assaf</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raphael Troncy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aline Senart</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>EURECOM</institution>
          ,
          <addr-line>Sophia Antipolis</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <country>SAP Labs France</country>
        </aff>
      </contrib-group>
      <fpage>32</fpage>
      <lpage>46</lpage>
      <abstract>
        <p>Linked Open Data (LOD) has emerged as one of the largest collections of interlinked datasets on the web. In order to bene t from this mine of data, one needs to access to descriptive information about each dataset (or metadata). This information can be used to delay data entropy, enhance dataset discovery, exploration and reuse as well as helping data portal administrators in detecting and eliminating spam. However, such metadata information is currently very limited to a few data portals where they are usually provided manually, thus being often incomplete and inconsistent in terms of quality. To address these issues, we propose a scalable automatic approach for extracting, validating, correcting and generating descriptive linked dataset pro les. This approach applies several techniques in order to check the validity of the metadata provided and to generate descriptive and statistical information for a particular dataset or for an entire data portal.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>Dataset Pro le</kwd>
        <kwd>Metadata</kwd>
        <kwd>Data Quality</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        From 12 datasets cataloged in 2007, the Linked Open Data cloud has grown to
nearly 1000 datasets containing more than 82 billion triples3 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Data is being
published by both the public and private sectors and covers a diverse set of
domains from life sciences to media or government data. The Linked Open Data
cloud is potentially a gold mine for organizations and individuals who are trying
to leverage external data sources in order to produce more informed business
decisions [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Dataset discovery can be done through public data portals like Datahub.io
and publicdata.eu or private ones like quandl.com and enigma.io. Private
portals harness manually curated data from various sources and expose them
to users either freely or through paid plans. Similarly, in some public data
portals, administrators manually review datasets information, validate, correct and
attach suitable metadata information. This information is mainly in the form
of prede ned tags such as media, geography, life sciences for organization and
clustering purposes. However, the diversity of those datasets makes it harder
to classify them in a xed number of prede ned tags that can be subjectively
3 http://datahub.io/dataset?tags=lod
assigned without capturing the essence and breadth of the dataset [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
Furthermore, the increasing number of datasets available makes the metadata review
and curation process unsustainable even when outsourced to communities.
      </p>
      <p>There are several Data Management Systems (DMS) that power public data
portals. CKAN4 is the world's leading open-source data portal platform
powering web sites like DataHub, Europe's Public Data and the U.S Government's
open data. Modeled on CKAN, DKAN5 is a standalone Drupal distribution
that is used in various public data portals as well. Socrata6 helps public sector
organizations improve data-driven decision making by providing a set of
solutions including an open data portal. In addition to these tradition data portals,
there is a set of tools that allow exposing data directly as RESTful APIs like
thedatatank.com.</p>
      <p>
        Metadata provisioning is one of the Linked Data publishing best practices
mentioned in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Datasets should contain the metadata needed to e ectively
understand and use them. This information includes the dataset's license,
provenance, context, structure and accessibility. The ability to automatically check
this metadata helps in:
{ Delaying data entropy: Information entropy refers to the degradation or
loss limiting the information content in raw or metadata. As a consequence
of information entropy, data complexity and dynamicity, the life span of
data can be very short. Even when the raw data is properly maintained, it
is often rendered useless when the attached metadata is missing, incomplete
or unavailable. Comprehensive high quality metadata can counteract these
factors and increase dataset longevity [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
{ Enhancing data discovery, exploration and reuse: Users who are
unfamiliar with a dataset require detailed metadata to interpret and analyze
accurately unfamiliar datasets. A study conducted by the European Union
commission [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] found that both business and users are facing di culties in
discovering, exploring and reusing public data due to missing or inconsistent
metadata information.
{ Enhancing spam detection: Portals hosting public open data like Datahub
allow anyone to freely publish datasets. Even with security measures like
captchas and anti-spam devices, detecting spam is increasingly di cult. In
addition to that, the increasing number of datasets hinders the scalability of
this process, a ecting the correct and e cient spotting of datasets spam.
      </p>
      <p>
        Data pro ling is the process of creating descriptive information and collect
statistics about that data. It is a cardinal activity when facing an unfamiliar
dataset [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Data pro les re ect the importance of datasets without the need
for detailed inspection of the raw data. It also helps in assessing the importance
of the dataset, improving users' ability to search and reuse part of the dataset
and in detecting irregularities to improve its quality. Data pro ling includes
typically several tasks:
      </p>
    </sec>
    <sec id="sec-2">
      <title>4 http://ckan.org 5 http://nucivic.com/dkan/ 6 http://www.socrata.com</title>
      <p>{ Metadata pro ling: Provides general information on the dataset (dataset
description, release and update dates), legal information (license information,
openness), practical information (access points, data dumps), etc.
{ Statistical pro ling: Provides statistical information about data types and
patterns in the dataset (e.g. properties distribution, number of entities and
RDF triples).
{ Topical pro ling: Provides descriptive knowledge on the dataset content
and structure. This can be in form of tags and categories used to facilitate
search and reuse.</p>
      <p>In this work, we address the challenges of automatic validation and
generation of descriptive dataset pro le. This paper proposes Roomba, an extensible
framework consisting of a processing pipeline that combines techniques for data
portals identi cation, datasets crawling and a set of pluggable modules
combining several pro ling tasks. The framework validates the provided dataset
metadata against an aggregated standard set of information. Metadata elds
are automatically corrected when possible (e.g. adding a missing license URL
reference). Moreover, a report describing all the issues highlighting those that
cannot be automatically xed is created to be sent by email to the dataset's
maintainer. There exist various statistical and topical pro ling tools for both
relational and Linked Data. The architecture of the framework allows to easily
add them as additional pro ling tasks. However, in this paper, we focus on the
task of dataset metadata pro ling. We validate our framework against a
manually created set of pro les and manually check its accuracy by examining the
results of running it on various CKAN-based data portals.</p>
      <p>The remainder of the paper is structured as follows. In Section 2, we review
relevant related work. In Section 3, we describe our proposed framework's
architecture and components that validate and generate dataset pro les. In Section 4,
we evaluate the framework and we nally conclude and outline some future work
in Section 5.
2</p>
      <sec id="sec-2-1">
        <title>Related Work</title>
        <p>
          Data Catalog Vocabulary (DCAT) [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and the Vocabulary of Interlinked Datasets
(VoID) [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] are concerned with metadata about RDF datasets. There exist
several tools aiming at exposing dataset metadata using these vocabularies. In [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ],
the authors generate VoID descriptions limited to a subset of properties that
can be automatically deduced from resources within the dataset. However, it
still provides data consumers with interesting insights. Flemming's Data
Quality Assessment Tool7 provides basic metadata assessment as it computes data
quality scores based on manual user input. The user assigns weights to the
prede ned quality metrics and answers a series of questions regarding the dataset.
These include, for example, the use of obsolete classes and properties by de ning
the number of described entities that are assigned disjoint classes, the usage of
7 http://linkeddata.informatik.hu-berlin.de/LDSrcAss/datenquelle.php
stable URIs and whether the publisher provides a mailing list for the dataset.
The ODI certi cate8, on the other hand, provides a description of the published
data quality in plain English. It aspires to act as a mark of approval that helps
publishers understand how to publish good open data and users how to use it.
It gives publishers the ability to provide assurance and support on their data
while encouraging further improvements through an ascending scale. ODI comes
as an online and free questionnaire for data publishers focusing on certain
characteristics about their data.
        </p>
        <p>Metadata pro ling: The Project Open Data Dashboard9 tracks and
measures how US government web sites implement the Open Data principles to
understand the progress and current status of their public data listings. A
validator analyzes machine readable les: e.g. JSON les for automated metrics
like the resolved URLs, HTTP status and content-type. However, deep schema
information about the metadata is missing like description, license information
or tags. Similarly on the LOD cloud, the Datahub LOD Validator10 gives an
overview of Linked Data sources cataloged on the Datahub. It o ers a
step-bystep validator guidance to check a dataset completeness level for inclusion in the
LOD cloud. The results are divided into four di erent compliance levels from
basic to reviewed and included in the LOD cloud. Although it is an excellent
tool to monitor LOD compliance, it still lacks the ability to give detailed
insights about the completeness of the metadata and overview on the state of the
entire LOD cloud group and it is very speci c to the LOD cloud group rules and
regulations.</p>
        <p>
          Statistical pro ling: Calculating statistical information on datasets is vital
to applications dealing with query optimization and answering, data cleansing,
schema induction and data mining [
          <xref ref-type="bibr" rid="ref15 ref18 ref22">18, 15, 22</xref>
          ]. Semantic sitemaps [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and
RDFStats [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] are one of the rst to deal with RDF data statistics and summaries.
ExpLOD [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] creates statistics on the interlinking between datasets based on
owl:sameAs links. In [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], the author introduces a tool that induces the actual
schema of the data and gather corresponding statistics accordingly. LODStats
[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is a stream-based approach that calculates more general dataset statistics.
ProLOD++ [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] is a Web-based tool that allows LOD analysis via automatically
computed hierarchical clustering [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Aether [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] generates VoID statistical
descriptions of RDF datasets. It also provides a Web interface to view and compare
VoID descriptions. LODOP [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] is a MapReduce framework to compute,
optimize and benchmark dataset pro les. The main target for this framework is to
optimize the runtime costs for Linked Data pro ling. In [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] authors calculate
certain statistical information for the purpose of observing the dynamic changes
in datasets.
        </p>
        <p>
          Topical Pro ling: Topical and categorical information facilitates dataset
search and reuse. Topical pro ling focuses on content-wise analysis at the
instances and ontological levels. GERBIL [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ] is a general entity annotation
frame
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>8 https://certificates.theodi.org/ 9 http://labs.data.gov/dashboard/ 10 http://validator.lod-cloud.net/</title>
      <p>
        work that provides machine processable output allowing e cient querying. In
addition, there exist several entity annotation tools and frameworks [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] but none
of those systems are designed speci cally for dataset annotation. In [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], the
authors created a semantic portal to manually annotate and publish metadata
about both LOD and non-RDF datasets. In [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], the authors automatically
assigned Freebase domains to extracted instance labels of some of the LOD Cloud
datasets. The goal was to provide automatic domain identi cation, thus enabling
improving datasets clustering and categorization. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the authors extracted
dataset topics by exploiting the graph structure and ontological information,
thus removing the dependency on textual labels. In [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the authors generate
VoID and VoL descriptions via a processing pipeline that extracts dataset topic
models ranked on graphical models of selected DBpedia categories.
      </p>
      <p>Although the above mentioned tools are able to provide various types of
information about a dataset, there exists no approach that aggregates this
information and is extensible to combine additional pro ling tasks. To the best
of our knowledge, this is the rst e ort towards extensible automatic validation
and generation of descriptive dataset pro les.
3</p>
      <sec id="sec-3-1">
        <title>Pro ling Data Portals</title>
        <p>In this section, we provide an overview of Roomba's architecture and the
processing steps for validating and generating dataset pro les. Figure 1 shows the
main steps which are the following: (i) data portal identi cation; (ii) metadata
extraction; (iii) instance and resource extraction; (iv) pro le validation (v) pro le
and report generation.</p>
        <p>Roomba is built as a Command Line Interface (CLI) application using Node.js.
Instructions on installing and running the framework are available on its public
Github repository11. The various steps are explained in detail below.
3.1</p>
        <sec id="sec-3-1-1">
          <title>Data Portal Identi cation</title>
          <p>Roomba should be extensible to any data portal that exposes its functionalities
via an external accessible API. Since every portal ca have its own data model,
identifying the software powering data portals is a vital rst step. We rely on
several Web scraping techniques in the identi cation process which includes a
combination of the following:
{ URL inspection: Various CKAN based portals are hosted on subdomains
of the http://ckan.net. For example, CKAN Brazil (http://br.ckan.
net). Checking the existence of certain URL patterns can detect such cases.
{ Meta tags inspection: The &lt;meta&gt; tag provides metadata about the HTML
document. They are used to specify page description, keywords, author, etc.
Inspecting the content attribute can indicate the type of the data portal.
We use CSS selectors to check the existence of these meta tags. An
example of a query selector is meta[content*=``ckan''] (all meta tags with
11 https://github.com/ahmadassaf/opendata-checker
the attribute content containing the string CKAN ). This selector can
identify CKAN portals whereas the meta[content*=``Drupal''] can identify
DKAN portals.
{ Document Object Model (DOM) inspection: Similar to the meta tags
inspection, we check the existence of certain DOM elements or properties. For
example, CKAN powered portals will have DOM elements with class names
like ckan-icon or ckan-footer-logo. A CSS selector like .ckan-icon will
be able to check if a DOM element with the class name ckan-icon exists. The
list of elements and properties to inspect is stored in a separate con gurable
object for each portal. This allows the addition and removal of elements as
deemed necessary.</p>
          <p>The identi cation process for each portal can be easily customized by overriding
the default function. Moreover, adding or removing steps from the identi cation
process can be easily con gured.</p>
          <p>After those preliminary checks, we query one of the portal's API endpoints.
For example, DataHub is identi ed as CKAN, so we will query the API endpoint
on http://datahub.io/api/action/package\_list. A successful request will
list the names of the site's datasets, whereas a failing request will signal a possible
failure of the identi cation process.
3.2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Metadata Extraction</title>
          <p>Data portals expose a set of information about each dataset as metadata. The
model used varies across portals. However, a standard model should contain
information about the dataset's title, description, maintainer email, update and
creation date, etc. We divided the metadata information into the following types:</p>
          <p>General information: General information about the dataset. e.g., title,
description, ID, etc. This general information is manually lled by the dataset
owner. In addition to that, tags and group information is required for classi
cation and enhancing dataset discoverability. This information can be entered
manually or inferred modules plugged into the topical pro ler.</p>
          <p>Access information: Information about accessing and using the dataset.
This includes the dataset URL, license information i.e., license title and URL
and information about the dataset's resources. Each resource has as well a set
of attached metadata e.g., resource name, URL, format, size.</p>
          <p>Ownership information: Information about the ownership of the dataset.
e.g., organization details, maintainer details, author. The existence of this
information is important to identify the authority on which the generated report and
the newly corrected pro le will be sent to.</p>
          <p>Provenance information: Temporal and historical information on the dataset
and its resources. For example, creation and update dates, version information,
version, etc. Most of this information can be automatically lled and tracked.</p>
          <p>Building a standard metadata model is not the scope of this paper, and since
we focus on CKAN-based portals, we validate the extracted metadata against
the CKAN standard model12.</p>
          <p>After identifying the underlying portal software, we perform iterative queries
to the API in order to fetch datasets metadata and persist them in a le-based
cache system. Depending on the portal software, we can issue speci c extraction
jobs. For example, in CKAN-based portals, we are able to crawl and extract the
metadata of a speci c dataset, all the datasets in a speci c group (e.g. LOD
cloud) or all the datasets in the portal.
3.3</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Instance and Resource Extraction</title>
          <p>From the extracted metadata we are able to identify all the resources associated
with that dataset. They can have various types like a SPARQL endpoint, API,
le, visualization, etc. However, before extracting the resource instance(s) we
perform the following steps:
{ Resource metadata validation and enrichment: Check the resource
attached metadata values. Similar to the dataset metadata, each resource
should include information about its mimetype, name, description, format,
valid de-referenceable URL, size, type and provenance. The validation
process issues an HTTP request to the resource and automatically lls up
various missing information when possible, like the mimetype and size by
extracting them from the HTTP response header. However, missing elds like
name and description that needs manual input are marked as missing and
will appear in the generated summary report.
{ Format validation: Validate speci c resource formats against a linter or
a validator. For example, node-csv13 for CSV les and n314 to validate N3
and Turtle RDF serializations.
12 http://demo.ckan.org/api/3/action/package\_show?id=adur\_district\
_spending
13 https://github.com/wdavidw/node-csv
14 https://github.com/RubenVerborgh/N3.js</p>
          <p>
            Considering that certain datasets contain large amounts of resources and the
limited computation power of some machines on which the framework might
run on, a sampler module can be introduced to execute various sample-based
strategies detailed as they were found to generate accurate results even with
comparably small sample size of 10%. These strategies introduced in [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] are:
{ Random Sampling: Randomly selects resource instances.
{ Weighted Sampling: Weighs each resources as the ratio of the number of
datatype properties used to de ne a resource over the maximum number of
datatype properties over all the datasets resources.
{ Resource Centrality Sampling: Weighs each resource as the ration of the
number of resource types used to describe a particular resource divided by
the total number of resource types in the dataset. This is speci c and
important to RDF datasets where important concepts tend to be more structured
and linked to other concepts.
          </p>
          <p>
            However, the sampler is not restricted only to these strategies. Strategies
like those introduced in [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ] can be con gured and plugged in the processing
pipeline.
3.4
          </p>
        </sec>
        <sec id="sec-3-1-4">
          <title>Pro le Validation</title>
          <p>A dataset pro le should include descriptive information about the data
examined. In our framework, we have identi ed three main categories of pro ling
information. However, the extensibility of our framework allows for additional
pro ling techniques to be plugged in easily (i.e. a quality pro ling module
reecting the dataset quality). In this paper, we focus on the task of metadata
pro ling.</p>
          <p>Metadata validation process identi es missing information and the ability to
automatically correct them. Each set of metadata (general, access, ownership
and provenance) is validated and corrected automatically when possible. Each
pro ler task has a set of metadata elds to check against. The validation process
check if each eld is de ned and if the value assigned is valid.</p>
          <p>There exist many special validation steps for various elds. For example, the
email addresses and urls should be validated to ensure that the value entered
is syntactically correct. In addition to that, for urls, we issue an HTTP HEAD
request in order to check if that URL is reachable. We also use the information
contained in a valid content-header response to extract, compare and correct
some resources metadata values like mimetype and size.</p>
          <p>From our experiments, we found out that datasets' license information is
noisy. The license names if found are not standardized. For example, Creative
Commons CCZero can be also CC0 or CCZero. Moreover,the license URI if
found and if de-referenceable can point to di erent reference knowledge bases
e.g., http://opendefinition.org. To overcome this issue, we have manually
created a mapping le standardizing the set of possible license names and the
reference knowledge base15. In addition, we have also used the open source and
15 https://github.com/ahmadassaf/opendata-checker/blob/master/util/
licenseMappings.json
knowledge license information16 to normalize the license information and add
extra metadata like the domain, maintainer and open data conformance.
f
g ,
f
g
" l i c e n s e i d " : [ "ODC PDDL 1.0" ] ,
" disambiguations " : [ "Open Data Commons Public Domain Dedication and License
(PDDL) " ]
" l i c e n s e i d " : [ "CC BY SA 4.0" , "CC BY SA 3.0" ] ,
" disambiguations " : [ " cc by sa " , "CC BY SA" , " Creative Commons A t t r i b u t i o n
Share Alike " ]</p>
          <p>Listing 1.1. License mapping le sample
3.5</p>
        </sec>
        <sec id="sec-3-1-5">
          <title>Pro le and Report Generation</title>
          <p>The validation process highlights the missing information and presents them in
a human readable report. The report can be automatically sent to the dataset
maintainer email if exists in the metadata. In addition to the generated report,
the enhanced pro les are represented in JSON using the CKAN data model and
are publicly available17.</p>
          <p>Data portal administrators need an overall knowledge of the portal datasets
and their properties. Our framework has the ability to generate numerous reports
of all the datasets by passing formatted queries. There are two main sets of
aggregation tasks that can be run:
{ Aggregating meta- eld values: Passing a string that corresponds to a
valid eld in the metadata. The eld can be at like license title
(aggregates all the license titles used in the portal or in a speci c group) or nested
like resource&gt;resource type (aggregates all the resources types for all the
datasets). Such reports are important to have an overview of the possible
values used for each metadata eld.
{ Aggregating key:object meta- eld values: Passing two meta- eld
values separated by a colon : e.g., resources&gt;resource type:resources&gt;name.
These reports are important as you can aggregate the information needed
when also having the set of values associated to it printed.</p>
          <p>For example, the meta- eld value query resource&gt;resource type run against
the LODCloud group will result in an array containing [f ile; api; documentation:::]
values. These are all the resource types used to describe all the datasets of
the group. However, to be able to know also what are the datasets containing
resources corresponding to each type, we issue a key:object meta- eld query
resource&gt;resource type:name. The result will be a JSON object having the
resource type as the key and an array of corresponding datasets titles that has
a resource of that type.
16 https://github.com/okfn/licenses
17 https://github.com/ahmadassaf/opendata-checker/tree/master/results
=======================================================================</p>
          <p>Metadata Report
=======================================================================
group i n f o r m a t i o n i s m i s s i n g . Check o r g a n i z a t i o n i n f o r m a t i o n as they
can be mixed sometimes
o r g a n i z a t i o n i m a g e u r l f i e l d e x i s t s but t h e r e i s no value d e f i n e d
=======================================================================</p>
          <p>Tag S t a t i s t i c s
=======================================================================</p>
          <p>There i s a t o t a l o f : 21 [ undefined ] v o c a b u l a r y i d f i e l d s 100.00%
=======================================================================</p>
          <p>L i c e n s e Report
=======================================================================</p>
          <p>L i c e n s e i n f o r m a t i o n has been normalized !
=======================================================================</p>
          <p>Resource S t a t i s t i c s
=======================================================================
There i s a t o t a l o f : 10 [ m i s s i n g ] url type f i e l d s 100.00%
There i s a t o t a l o f : 9 [ m i s s i n g ] c r e a t e d f i e l d s 90.00%
There i s a t o t a l o f : 10 [ undefined ] c a c h e l a s t u p d a t e d f i e l d s 100.00%
There i s a t o t a l o f : 10 [ undefined ] s i z e f i e l d s 100.00%
There i s a t o t a l o f : 10 [ undefined ] hash f i e l d s 100.00%
There i s a t o t a l o f : 10 [ undefined ] mimetype inner f i e l d s 100.00%
There i s a t o t a l o f : 7 [ undefined ] mimetype f i e l d s 70.00%
There i s a t o t a l o f : 10 [ undefined ] c a c h e u r l f i e l d s 100.00%
There i s a t o t a l o f : 6 [ undefined ] name f i e l d s 60.00%
There i s a t o t a l o f : 9 [ undefined ] w e b s t o r e u r l f i e l d s 90.00%
There i s a t o t a l o f : 9 [ undefined ] l a s t m o d i f i e d f i e l d s 90.00%
There i s one [ undefined ] format f i e l d 10.00%
=======================================================================</p>
          <p>Resource C o n n e c t i v i t y I s s u e s
=======================================================================
There are 2 c o n n e c t i v i t y i s s u e s with the f o l l o w i n g URLs :</p>
          <p>n u r l f http : / / dbpedia . org / void / Dataset g
=======================================================================</p>
          <p>Un Reachable URLs Types
=======================================================================
There are : 1 unreachable URLs o f type [ f i l e ]</p>
          <p>Listing 1.2. Excerpt of the DBpedia validation report</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>4 Experiments and Evaluation</title>
        <p>In this section, we provide the experiments and evaluation of the proposed
framework. All the experiments are reproducible by our tool and their results are
available in its Github repository. A CKAN dataset metadata describes four main
sections in addition to the core dataset's properties. These sections are:
{ Resources: The distributable parts containing the actual raw data. They
can come in various formats (JSON, XML, RDF, etc.) and can be
downloaded or accessed directly (REST API, SPARQL endpoint).
{ Tags: Provide descriptive knowledge on the dataset content and structure.</p>
        <p>They are used mainly to facilitate search and reuse.
{ Groups: A dataset can belong to one or more group that share common
semantics. A group can be seen as a cluster or a curation of datasets based
on shared categories or themes.
{ Organizations: A dataset can belong to one or more organization controlled
by a set of users. Organizations are di erent from groups as they are not
constructed by shared semantics or properties, but solely on their association
to a speci c administration party.</p>
        <p>Each of these sections contains a set of metadata corresponding to one or
more type (general, access, ownership and provenance). For example, a dataset
resource will have general information such as the resource name, access
information such as the resource url and provenance information such as creation
date. The framework generates a report aggregating all the problems in all these
sections, xing eld values when possible. Errors can be the result of missing
metadata elds, unde ned eld values or eld value errors (e.g. unreachable URL
or incorrect email addresses).
4.1</p>
        <sec id="sec-3-2-1">
          <title>Experimental Setup</title>
          <p>
            We ran our tool on two CKAN-based data portals. The rst one is datahub.io
targeting speci cally the LOD cloud group. The current state of the LOD cloud
report [
            <xref ref-type="bibr" rid="ref27">27</xref>
            ] indicates that the LOD cloud contains 1014 datasets. They were
harvested via a LDSpider crawler [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] seeded with 560 thousands URIs. Roomba, on
the other hand, fetches datasets hosted in data portals where datasets have
attached relevant metadata. As a result, we relied on the information provided by
the Datahub CKAN API. Examining the tags available, we found two candidate
groups. The rst one tagged with \lodcloud" returned 259 datasets, while the
second one tagged with \lod" returned only 75 datasets. After manually
examining the two lists, we found out the datasets grouped with the tag \lodcloud"
are the correct ones. To qualify other CKAN-based portals for the experiments,
we use http://dataportals.org/ which contains a comprehensive list of Open
Data portals from around the world. In the end, we chose the Amsterdam data
portal18. The portal was commissioned in 2012 by the Amsterdam Economic
Board Open Data Exchange (ODE) and covers a wide range of information
domains (energy, economy, education, urban development, etc.) about Amsterdam
metropolitan region.
          </p>
          <p>We ran the Roomba instance and resource extractors in order to cache the
metadata les for these datasets locally and ran the validation process. The
experiments were executed on a 2.6 Ghz Intel Core i7 processor with 16GB of
DDR3 memory machine. The approximate execution time alongside the
summary of the datasets' properties are presented in table 1.</p>
          <p>In our evaluation, we focused on two aspects: i)pro ling correctness which
manually assesses the validity of the errors generated in the report, and ii)pro ling
18 http://data.amsterdamopendata.nl/
completeness which assesses if the pro lers cover all the errors in the datasets
metadata.
4.2</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Pro ling Correctness</title>
          <p>To measure pro le correctness, we need to make sure that the issues reported
by Roomba are valid on the dataset, group and portal levels.</p>
          <p>On the dataset level, we choose three datasets from both the LOD Cloud
and the Amsterdam data portal. The datasets details are shown in table 2.</p>
          <p>To measure the pro ling correctness on the groups level, we selected four
groups from the Amsterdam data portal containing a total of 25 datasets. The
choice was made to cover groups in various domains that contain a moderate
number of datasets that can be checked manually (between 3-9 datasets). Table
3 summarizes the groups chosen for the evaluation.</p>
          <p>After running Roomba and examining the results on the selected datasets
and groups, we found out that our framework provides 100% correct results
on the individual dataset level and on the aggregation level over groups. Since
our portal level aggregation is extended from the group aggregation, we can
infer that the portal level aggregation also produces complete correct pro les.
However, the lack of a standard way to create and manage collections of datasets
was the source of some errors when comparing the results from these two portals.
For example, in Datahub, we noticed that all the datasets groups information
were missing, while in the Amsterdam Open Data portal, all the organisation
information was missing. Although the error detection is correct, the overlap
in the usage of group and organization can give a false indication about the
metadata quality.
4.3</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>Pro ling Completeness</title>
          <p>We analyzed the completeness of our framework by manually constructing a
set of pro les that act as a golden standard. These pro les cover the range of
uncommon problems that can occur in a certain dataset19. These errors are:
{ Incorrect mimetype or size for resources;
{ Invalid number of tags or resources de ned;
{ Check if the license information can be normalized via the license id or
the license title as well as the normalization result;
{ Syntactically invalid author email or maintainer email.</p>
          <p>After running our framework at each of these pro les, we measured the
completeness and correctness of the results. We found out that our framework covers
indeed all the metadata problems that can be found in a CKAN standard model
correctly.
5</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Conclusion and Future Work</title>
        <p>In this paper, we proposed a scalable automatic approach for extracting,
validating, correcting and generating descriptive linked dataset pro les. This approach
applies several techniques in order to check the validity of the metadata provided
and to generate descriptive and statistical information for a particular dataset
or for an entire data portal. Based on our experiments running the tool on the
LOD cloud, we discovered that the general state of the datasets needs attention
as most of them lack informative access information and their resources su er
low availability. These two metrics are of high importance for enterprises looking
to integrate and use external linked data.</p>
        <p>It has been noticed that the issues surrounding metadata quality a ect
directly dataset search as data portals rely on such information to power their
search index. We noted the need for tools that are able to identify various issues
in this metadata and correct them automatically. We evaluated our framework
manually against two prominent data portals and proved that we can
automatically scale the validation of datasets metadata pro les completely and correctly.</p>
        <p>As part of our future work, we plan to introduce work ows that will be
able to correct the rest of the metadata either automatically or through
intuitive manually-driven interfaces. We also plan to integrate statistical and topical
pro lers to be able to generate full comprehensive pro les. We also intend to
suggest a ranked standard metadata model that will help generate more accurate
and scored metadata quality pro les. We also plan to run this tool on various
CKAN-based data portals, schedule periodic reports to monitor the evolvement
of datasets metadata. Finally, at some stage, we plan to extend this tool for
other data portal types like DKAN and Socrata.
19 https://github.com/ahmadassaf/opendata-checker/tree/master/test</p>
      </sec>
      <sec id="sec-3-4">
        <title>Acknowledgments</title>
        <p>This research has been partially funded by the European Union's 7th Framework
Programme via the project Apps4EU (GA No. 325090).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gruetze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jentzsch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          .
          <article-title>Pro ling and mining RDF data with ProLOD++</article-title>
          .
          <source>In 30th IEEE International Conference on Data Engineering (ICDE)</source>
          , pages
          <fpage>1198</fpage>
          {
          <fpage>1201</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Demter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <surname>and J. Lehmann.</surname>
          </string-name>
          <article-title>LODStats - an Extensible Framework for High-performance Dataset Analytics</article-title>
          .
          <source>In 18th International Conference on Knowledge Engineering and Knowledge Management (EKAW)</source>
          , pages
          <fpage>353</fpage>
          {
          <fpage>362</fpage>
          ,
          <string-name>
            <surname>Galway</surname>
          </string-name>
          , Ireland,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. C. Bohm, G. Kasneci, and
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          .
          <article-title>Latent Topics in Graph-structured Data</article-title>
          .
          <source>In 21st ACM International Conference on Information and Knowledge Management (CIKM)</source>
          , pages
          <fpage>2663</fpage>
          {
          <fpage>2666</fpage>
          ,
          <string-name>
            <surname>Maui</surname>
          </string-name>
          , Hawaii, USA,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Bohm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abedjan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fenz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Grutze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hefenbrock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pohl</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Sonnabend</surname>
          </string-name>
          .
          <article-title>Pro ling linked open data with ProLOD</article-title>
          .
          <source>In 26th International Conference on Data Engineering Workshops (ICDEW)</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>D.</given-names>
            <surname>Boyd</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Crawford</surname>
          </string-name>
          .
          <article-title>Six provocations for big data</article-title>
          .
          <source>A Decade in Internet Time: Symposium on the Dynamics of the Internet and Society</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>B.</given-names>
            <surname>Christian</surname>
          </string-name>
          .
          <article-title>Evolving the Web into a Global Data Space</article-title>
          .
          <source>In 28th British National Conference on Advances in Databases</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>B.</given-names>
            <surname>Christian</surname>
          </string-name>
          , H. T, and B.
          <string-name>
            <surname>-L. T. Linked Data - The Story</surname>
          </string-name>
          So Far.
          <source>International Journal on Semantic Web and Information Systems (IJSWIS)</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>B.</given-names>
            <surname>Christoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Johannes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Felix</surname>
          </string-name>
          .
          <article-title>Creating voiD Descriptions for Web-scale Data</article-title>
          .
          <source>Journal of Web Semantics</source>
          ,
          <volume>9</volume>
          (
          <issue>3</issue>
          ):
          <volume>339</volume>
          {
          <fpage>345</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornolti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ferragina</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ciaramita</surname>
          </string-name>
          .
          <article-title>A Framework for Benchmarking Entity-annotation Systems</article-title>
          .
          <source>In 22nd World Wide Web Conference (WWW)</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Stenzhorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Delbru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Tummarello. Semantic Sitemaps</surname>
          </string-name>
          :
          <article-title>E cient and Flexible Access to Datasets on the Semantic Web</article-title>
          .
          <source>In 5th European Semantic Web Conference (ESWC)</source>
          , pages
          <fpage>690</fpage>
          {
          <fpage>704</fpage>
          ,
          <string-name>
            <surname>Tenerife</surname>
          </string-name>
          , Spain,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hausenblas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Alexander</surname>
          </string-name>
          .
          <article-title>Describing Linked Datasets with the VoID Vocabulary</article-title>
          .
          <source>W3C Note</source>
          ,
          <year>2011</year>
          . http://www.w3.org/TR/ void/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>M.</given-names>
            <surname>Fadi</surname>
          </string-name>
          and
          <string-name>
            <surname>E. John. Data Catalog</surname>
          </string-name>
          <article-title>Vocabulary (DCAT)</article-title>
          .
          <source>W3C Recommendation</source>
          ,
          <year>2014</year>
          . http://www.w3.org/TR/vocab-dcat/.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>B.</given-names>
            <surname>Fetahu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dietze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Pereira</given-names>
            <surname>Nunes</surname>
          </string-name>
          , M. Antonio Casanova,
          <string-name>
            <given-names>D.</given-names>
            <surname>Taibi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Nejdl</surname>
          </string-name>
          .
          <article-title>A Scalable Approach for E ciently Generating Structured Dataset Topic Proles</article-title>
          .
          <source>In 11th European Semantic Web Conference (ESWC)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>B.</given-names>
            <surname>Forchhammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jentzsch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann. LODOP</surname>
          </string-name>
          <article-title>- Multi-Query Optimization for Linked Data Pro ling Queries</article-title>
          . In International Workshop on Dataset PROFIling and
          <article-title>fEderated Search for Linked Data (PROFILES), Heraklion</article-title>
          , Greece,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>M. Frosterus</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <article-title>Hyvonen, and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Laitio</surname>
          </string-name>
          .
          <article-title>Creating and Publishing Semantic Metadata about Linked and Open Datasets</article-title>
          .
          <source>In Linking Government Data</source>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>M. Frosterus</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <article-title>Hyvonen, and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Laitio. DataFinland - A Semantic</surname>
          </string-name>
          <article-title>Portal for Open and Linked Datasets</article-title>
          .
          <source>In 8th Extended Semantic Web Conference (ESWC)</source>
          , pages
          <fpage>243</fpage>
          {
          <fpage>254</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>R.</given-names>
            <surname>Isele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Umbrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. Harth.</surname>
          </string-name>
          <article-title>LDspider: An Open-source Crawling Framework for the Web of Linked Data</article-title>
          .
          <source>In 9th International Semantic Web Conference (ISWC)</source>
          ,
          <source>Posters &amp; Demos Track</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>A.</given-names>
            <surname>Jentzsch</surname>
          </string-name>
          .
          <article-title>Pro ling the Web of Data</article-title>
          .
          <source>In 13th International Semantic Web Conference (ISWC)</source>
          , Doctoral Consortium, Trentino, Italy,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. T. Kafer,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdelrahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Umbrich</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <article-title>O'Byrne, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          .
          <article-title>Observing Linked Data Dynamics</article-title>
          .
          <source>In 10th European Semantic Web Conference (ESWC)</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>S.</given-names>
            <surname>Khatchadourian</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Consens</surname>
          </string-name>
          . ExpLOD:
          <article-title>Summary-based Exploration of Interlinking and RDF Usage in the Linked Open Data Cloud</article-title>
          .
          <source>In 7th Extended Semantic Web Conference (ESWC)</source>
          , pages
          <fpage>272</fpage>
          {
          <fpage>287</fpage>
          ,
          <string-name>
            <surname>Heraklion</surname>
          </string-name>
          , Greece,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Kovacs-Lang</surname>
          </string-name>
          .
          <article-title>Global Terrestrial Observing System</article-title>
          .
          <source>Technical report, GTOS Central and Eastern European Terrestrial Data Management and Accessibility Workshop</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. S. Lalithsena,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hitzler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sheth</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Jain</surname>
          </string-name>
          .
          <article-title>Automatic Domain Identi cation for Linked Open Data</article-title>
          .
          <source>In IEEE/WIC/ACM International Joint Conferences on Web Intelligence (WI) and Intelligent Agent Technologies (IAT)</source>
          , pages
          <fpage>205</fpage>
          {
          <fpage>212</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>A.</given-names>
            <surname>Langegger</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Woss. RDFStats - An Extensible RDF Statistics</surname>
          </string-name>
          <article-title>Generator and Library</article-title>
          .
          <source>In 20th International Workshop on Database and Expert Systems Application (DEXA)</source>
          , pages
          <fpage>79</fpage>
          {
          <fpage>83</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Faloutsos</surname>
          </string-name>
          .
          <article-title>Sampling from Large Graphs</article-title>
          .
          <source>In 12thth ACM International Conference on Knowledge Discovery and Data Mining (KDD'12)</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Data Pro ling for Semantic Web Data</article-title>
          .
          <source>In International Conference on Web Information Systems and Mining (WISM)</source>
          , pages
          <fpage>472</fpage>
          {
          <fpage>479</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26. E. Makela.
          <article-title>Aether - Generating and Viewing Extended VoID Statistical Descriptions of RDF Datasets</article-title>
          .
          <source>In 11th European Semantic Web Conference (ESWC)</source>
          , Demo Track, Heraklion, Greece,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>S.</given-names>
            <surname>Max</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Christian</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Heiko</surname>
          </string-name>
          .
          <article-title>Adoption of the Linked Data Best Practices in Di erent Topical Domains</article-title>
          .
          <source>In 13th International Semantic Web Conference (ISWC)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <given-names>R.</given-names>
            <surname>Usbeck</surname>
          </string-name>
          , M. Roder, A.
          <string-name>
            <surname>-C.</surname>
            Ngonga-Ngomo,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Baron</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Both</surname>
            , M. Brummer,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ceccarelli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Cornolti</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Cherix</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Eickmann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Ferragina</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lemke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Moro</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Piccinno</surname>
            , G. Rizzo,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Sack</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Speck</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Troncy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Waitelonis</surname>
            , and
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Wesemann. GERBIL - General Entity Annotation Benchmark</surname>
          </string-name>
          <article-title>Framework</article-title>
          .
          <source>In 24th World Wide Web Conference (WWW)</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29. G. Vickery.
          <article-title>Review of Recent Studies on PSI-use and Related Market Developments</article-title>
          .
          <source>Technical report, EC DG Information Society</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>