<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Thematic Exploration of Linked Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Silvana Castano</string-name>
          <email>silvana.castano@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alfio Ferrara</string-name>
          <email>o.ferrara@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Montanelli</string-name>
          <email>stefano.montanelli@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universita` degli Studi di Milano</institution>
          ,
          <addr-line>DICo - Via Comelico, 39 20135 Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Now that a huge amount of data is available in the Linked Data Cloud, providing techniques for its e ective exploration is becoming more and more important. In this paper, we propose aggregation and abstraction techniques for thematic exploration of linked data. These techniques transform a basic, at view of a potentially large set of messy linked data for a given search target, into a high-level, thematic view called inCloud. In an inCloud, thematic exploration is guided by few essentials auto-describing their prominence for the search target and by their reciprocal proximity relations. Linked data aggregation, labeling, and exploration</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>The Linked Data paradigm promoted a new way of
exposing, sharing, and connecting pieces of data, information,
and knowledge on the Semantic Web, based on URIs
(Universal Resource Identi er) and RDF (Resource Description
Framework) [1]. Now that a huge amount of data is available
in the Linked Data Cloud, providing techniques for e ective
linked data searching, exploration, and visualization is
becoming crucial [7, 9]. In the recent literature, issues related
to linked data exploration are getting more and more
importance [8, 12]. One of the most challenging questions is
to provide e ective browsing solutions capable to deal with
the inherent at organization of linked data and to manage
the existing huge-sized repositories storing millions of RDF
triples.</p>
      <p>Permission to make digital or hard copies of all or part of this work for
personal or classroom use is granted without fee provided that copies are
not made or distributed for profit or commercial advantage and that copies
bear this notice and the full citation on the first page. To copy otherwise, to
republish, to post on servers or to redistribute to lists, requires prior specific
permission and/or a fee. This article was presented at:
Very Large Data Search (VLDS) 2011.</p>
      <p>Copyright 2011.</p>
      <p>In this context, we propose abstraction and aggregation
techniques to transform a basic, at view of a potentially
large set of messy linked data, into an inCloud, that is, a
high-level, thematic view enabling a more e ective,
themedriven exploration of the same dataset. Through
aggregation techniques, we identify clusters of semantically related
linked data in a (even large) collection representing the
response to a search target. Through abstraction techniques,
we mine suitable essentials capturing the theme dealt with
a linked data cluster and its relevance for the search
target, as well proximity relations re ecting reciprocal degree of
closeness between cluster essentials. We motivate the role of
inClouds through a real example of linked data collection
extracted from the Freebase repository considering Van Gogh as
search target. Moreover, we will describe the construction of
an inCloud through aggregation and abstraction techniques.
Finally, we show how inCloud representation can be used for
thematic browsing and exploration of the underlying linked
data collection.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>MOTIVATING EXAMPLE</title>
      <p>In a common scenario, the user interested in exploring a
linked data repository to satisfy a certain search target
usually has to face a long and loosely-intuitive browsing activity.
This is due to the inherent at organization of linked data
repositories where the URIs of interest for a given target
frequently require the user to follow more than one property
link before being explored. In particular, the user
exploration is typically characterized by the following steps:
Submission to the repository of a search target (t),
namely a keyword (or a list of keywords) that describes
the subject of interest for the search. An example of
search target is the name of the famous painter Vincent
van Gogh.</p>
      <p>Selection of the seed of interest (s), namely an URI
that represents the \point of origin" for the exploration
about the search target. The seed of interest is chosen
from the list of URIs returned by the repository as a
reply to the search target. In the Freebase linked data
repository1, an example of seed for our target is the
URI /en/vincent_van_gogh.</p>
      <p>Exploration of the URIs reachable from the seed with
the aim to get access to more information about the
search target. This requires the user to submit
appropriate queries to the repository to extract the seed</p>
      <sec id="sec-2-1">
        <title>1http://www.freebase.com/.</title>
        <p>prop erties and the URIs directly linked to s through
these prop erties. An example of MQL query for the
Freebase rep ository to extract the artworks directly
connected to the seed s = /en/vincent_van_gogh through
the prop erty /visual_art/visual_artist/artworks
is [f \id": \/en/vincent van gogh", \type": \/visual
art/visual artist", \/visual art/visual artist/artworks": fgg].</p>
        <p>The exploration step can b e recursively applied to the
visited URIs to progressively discover further URIs at higher
distance from the seed according to the user choices and
interests.</p>
        <p>Due to the huge numb er of linked data that is usually
concerned with a search target, a lot of exploration steps are
required to build a (more or less) comprehensive picture of
the available information ab out the target. As an example,
in Figure 1, we show the set of linked data extracted from
Freebase for the target Vincent van Gogh. In this example, we
/fictional_universe/ethnicity_in_fiction
/media_common/lost_work /people/ethnicity
considered the seed s = /en/vincent_van_gogh, we explored
the complete set of directly linked URIs and some selected
URIs at distance d = 2 from s. As it is clear from this simple
example, exploring such a at and huge collection of data
is cumb ersome. First, b ecause the representation is at and
it is imp ossible to immediately understand whether some
URIs are more imp ortant than others. Moreover, p ossible
sets of URIs addressing the same/similar argument ab out
the target are not highlighted nor group ed.</p>
        <p>The solution we prop ose is based on aggregation and
abstraction techniques to transform a basic, at view of linked
data like the one in Figure 1, into an inCloud providing a
high-level, thematic view of the same data. inClouds are
conceived to b e coupled with the conventional query
interfaces of the existing Linked Data rep ositories, in that they
can b e built on top of an extracted dataset to provide a more
e ective presentation of the result.</p>
        <p>An example of inCloud for the seed s = /en/vincent_
van_gogh
is
shown
in</p>
        <p>Fig
ure
2.
an
inCloud
a circle-b ox represents
linked data fo cused on
lated to the considered
that in uenced or have
(cluster C l3 ) or the set
sun owers (Cluster C l4 );
a
a
a cluster, namely a group of
sp eci c argument/topic
reseed (e.g., the set of artists
b een in uenced by van Gogh
of van Gogh artworks ab out
a square-b ox represents an essential, namely a concise
and convenient summary of the content of a cluster at a
glance (e.g., Topic Artwork, Sun ower used to summarize
the content of Cluster C l4 ). Clusters in an inCloud
are also characterized by a prominence value denoting
its level of imp ortance for the target in the framework
of the overall inCloud. Prominence values determine
the size of the cluster circles, thus the most prominent
cluster in the inCloud of Figure 2 is Cluster C l1 ;
an arrow represents a proximity relation b etween
clusters/essentials, namely a closeness relationship b etween
the themes/topics their represent. The arrow
thickness denotes the degree of proximity b etween the two
clusters/essentials connected by the arrow.</p>
        <p>Aggregation techniques are rst employed to enforce a
\thematic" clustering of the initial set of linked data, as
describ ed in Section 3. Abstraction techniques are then
applied to synthesize an inCloud over the thematic clusters, as
describ ed in Section 4.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. LINKED DATA AGGREGATION</title>
      <p>The goal of aggregation techniques is to transform an
initial set of linked data into a numb er of thematic clusters.
The starting p oint is a RDF graph Gs containing the linked
data ab out a certain seed s of interest automatically
extracted from a Linked Data rep ository R. Appropriate
extraction queries are de ned to this end according to the
language (e.g., SPARQL, MQL) supp orted by the rep ository
R. These queries generally enforce the following
extraction/ ltering op erations:</p>
      <p>Extraction of properties and corresponding values within
a distance d from the seed s. We consider that an
URI in the rep ository R is concerned with the seed s if
there is a prop erty path of length d b etween the URI
and s. The distance d can b e dynamically changed and
it has an impact on the numb er of extracted linked
data and thus on the size of the resulting RDF graph.
In usual scenarios, a distance d = 2 is a go o d trade-o
to obtain a su cient numb er of linked data ab out s
and a well-sized RDF graph.</p>
      <p>Extraction of the URI types. For each URI within
distance d from the seed s, we extract the list
typ es (i.e., classes) the URI b elongs to. The appro
ate prop erty of the rep ository R is exploited to
end (e.g., the prop erty type in Freebase).
a
of
prithis
Filtering of non-relevant properties. Lo osely
meaningful prop erties of a rep ository, like the prop erty image
of Freebase, can b e excluded from the resulting RDF
graph since they are p o orly useful in providing
information ab out s.
Written Work Book /en/letters_from_</p>
      <p>
        provence (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
Van Gogh /en/vincent_by_
      </p>
      <p>
        himself (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
/en/the_works_of_vincent_van_gogh (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
/en/color_your_own_van_gogh_paintings (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
/en/vincent_
van_gogh (20)
/en/portrait_of_
camille_roulin (
        <xref ref-type="bibr" rid="ref7">7</xref>
        )
/en/portrait_of_eugene_boch (
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
/en/portrait_of_adeline_ravoux (
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
/en/the_painter_of_sunflowers (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
/en/self_portrait_with_
bandaged_ear (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
      </p>
      <p>...</p>
      <p>Topic Artwork</p>
      <p>
        Sunflower
/en/the_painter_of_sunflowers (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
//en/vase_with_three_sunflowers (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
      </p>
      <p>
        /en/vase_with_twelve_sunflowers (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
/en/vase_with_fifteen_sunflowers (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
      </p>
      <p>
        /en/two_cut_sunflowers (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
/en/four_cut_sunflowers (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
...
      </p>
      <p>Cl4
inCloud for Vincent Van Gogh
(/en/vincent_van_gogh)</p>
      <p>ESSENTIAL
THEMATIC CLUSTER
PROXIMITY LINK
Topic Artwork</p>
      <p>
        Garden
/en/the_starry_night (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
/en/farmhouse_in_provence (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
/en/the_painter_of_sunflowers (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
/en/the_olive_trees (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
/en/spring_in_arles (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
/en/irises (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
...
      </p>
      <p>Cl5
The query result is the graph Gs = (Ns; Es) where a node
n 2 Ns, called linked data entity, can be an URI, a literal,
or a type value that satisfy the query selection, and an edge
e (ni; nj ) 2 Es, called property link, represents a property
relationship of R between the nodes ni; nj 2 Ns.</p>
      <p>Based on the RDF graph Gs, linked data aggregation is
articulated in two main steps, namely similarity evaluation
and thematic clustering.</p>
    </sec>
    <sec id="sec-4">
      <title>3.1 Similarity evaluation</title>
      <p>This step has the goal to analyze the graph Gs and to
generate an augmented linked data graph Gs+ where a
similarity link is added between each pair of matching linked
data entities in Ns. To this end, the level of a nity
between the entities of Ns is evaluated as follows. Given two
linked data entities ni, nj 2 Ns, the linked data a nity
(ni; nj ) 2 [0; 1] denotes the level of similarity of ni and
nj based on the commonalities of their terminological
equipments. Each linked data entity n 2 Ns is associated with
a terminological equipment Termn = fterm1; : : : ; termmg
where termj , with 1 j m, is a term appearing in the
label of a node adjacent to n in Gs, or a term appearing
in the label of n itself. Before inclusion in a terminological
equipment, each term is submitted to a normalization
procedure for word-lemma extraction and for compound-term
tokenization [4, 15].</p>
      <p>The a nity of two linked data entities ni, nj 2 Ns is
calculated as the Dice coe cient over their terminological
equipments as follows:
(ni; nj ) =
2 j termx termy j
j Termni j + j Termnj j
where termx termy denotes that termx 2 Termni and
termy 2 Termnj are matching terms according to a string
matching metric that considers the structure of the terms
termx and termy. For calculation, we employ our
matching system HMatch 2.0, where state-of-the-art metrics for
string matching (e.g., I-Sub, Q-Gram, Edit-Distance, and
JaroWinkler) are implemented [2]. A similarity link e (ni; nj ) is
established between the linked data entities ni and nj i
(ni; nj ) th where th 2 (0; 1] is a matching threshold
denoting the minimum level of similarity required to consider
two linked data entities as matching entities.</p>
    </sec>
    <sec id="sec-5">
      <title>3.2 Thematic aggregation</title>
      <p>This step has the goal to analyze the graph Gs+ obtained
through similarity evaluation and to identify/mine a set
CL of thematic clusters. Given a graph Gs+, a cluster Cl
= f(n1; f1) ; : : : ; (nh; fh)g is a set of linked data entities
n1; : : : ; nh 2 Ns that are more similar to each other than
to the other entities of Ns. Each entity nj belonging to Cl
is associated with a corresponding frequency fj which
denotes the number of occurrences of nj in Cl.</p>
      <p>Clusters are determined by exploiting the graph Gs+ and
by detecting those node regions that are highly
interconnected through property/similarity links. The problem of
thematic aggregation is analogous to the problem of cluster
calculation, also known as module, community, or cohesive
group, in graph theory. For this reason, for thematic
aggregation, we rely on a clique percolation method (CPM) [13].
The CPM is based on the notion of k-clique which
corresponds to a complete (fully-connected) sub-graph of k nodes
within the graph Gs+. Two k-cliques are de ned as adjacent
k-cliques if they share k 1 nodes. The CPM determines
clusters from k-cliques. In particular, a cluster, or more
precisely, a k-clique-cluster, is de ned as the union of all
kcliques that can be reached from each other through a series
of adjacent k-cliques. As a consequence, a typical
k-cliquecluster is composed of several cliques (with size k) that
tend to share many of their nodes. Since the cliques of a
graph can share one or more nodes, we observe that a node
can belong to several clusters, and thus clusters can
overlap. In our approach, we employ the CPM implemented
in the CFinder tool2. Although the determination of the full
set of cliques of a graph is widely believed to be a
nonpolynomial problem, CFinder proves to be e cient when
applied to graphs like those considered in our approach. Such
an algorithm is based on rst locating all complete
subgraphs of Gs+ that are not part of larger complete subgraphs,
and then on identifying existing k-clique-clusters by
carrying out a standard component analysis of the clique-clique
overlap matrix [6]. As a result, CFinder produces the full set
CL of k-clique-clusters existing in the graph Gs+ for all the
possible values of k. A linked data entity ni belonging to
a cluster Cl 2 CL is represented as a pair (ni; fi) where
the frequency value fi denotes the number of cliques of Cl
which the entity nj belongs to (see Example of Figure 2).
The entities of a cluster are represented with di erent sizes,
proportional to the corresponding frequency values
according to a visualization manner \a la tag-cloud"3.</p>
    </sec>
    <sec id="sec-6">
      <title>LINKED DATA ABSTRACTION</title>
      <p>The goal of linked data abstraction techniques is to build
an inCloud, namely a high-level view on top of linked data
clusters by synthesizing them through essentials. inCloud
clusters are also featured by a level of prominence and by
proximity relations that denote the level of overlapping of
the di erent clusters.
4.1</p>
    </sec>
    <sec id="sec-7">
      <title>Essential abstraction</title>
      <p>An essential Essi is a concise and convenient summary of
a thematic cluster Cli and it is de ned as a pair of the form
Essi = (Ci; Di) where Ci is the category associated with
Cli and Di is a descriptor associated with Cli. A category
Ci is a set composed by the labels of the most frequent
types of the linked data entities in Cli, while a descriptor
Di is a set composed by the most frequent terms in the
terminological equipments of the entities in Cli. If more
than one most equally-frequent type and/or term exist, they
are all inserted in Ci and Di, respectively. In the example
of Figure 2, the cluster Cl4 corresponds to a very focused
theme expressed by the essential category Topic Artwork (the
most frequent type of the entities in the cluster) and by
the essential descriptor Sun ower (the most frequent term in
the terminological equipments of the entities in Cl4). In
cases where many entities are equally frequent in a cluster,
the abstracted essential is less focused and contains more
terms. This is the case for example of the cluster Cl3 of
Figure 2, representing persons and visual artists in uenced
by Van Gogh. In this case, the most frequent terms used
as descriptors are the names of the people involved in the
cluster, which are all equally frequent in the cluster.
4.2</p>
    </sec>
    <sec id="sec-8">
      <title>Prominence evaluation</title>
      <p>Clusters (and related essentials) in an inCloud are
differently relevant with respect to the original search target.</p>
      <sec id="sec-8-1">
        <title>2Available at http://www.cfinder.org/.</title>
        <p>3For a more readable visualization of highly-populated
clusters, the representation of less-frequent linked data entities
can be omitted.</p>
        <p>In order to represent this fact, we introduce the notion of
prominence of a cluster, namely a value Pi 2 [0; 1]. The
higher Pi is, the higher is also the prominence of Cli in the
inCloud. In our approach, the level of prominence of a
cluster is higher when the cluster is very focused on its theme
and its contents are homogeneous. In particular, we
formalize two cluster properties that are variability and density.</p>
        <p>Variability vi is the degree of overlap among the cliques
of the cluster Cli. For a linked data entity nj 2 Ns+, we
call fj the frequency of nj, that is the number of cliques of
Cli that contain nj. Variability vi is measured by a coe
cient of variation, which is the ratio between the standard
deviation of the linked data entity frequencies in Cli and
the arithmetic mean of those frequencies, as follows (with f
denoting the arithmetic mean value of frequencies):
vi =</p>
        <p>v
1 uu
f t Ni
1
1</p>
        <p>Ni
X(fi
i=1
f )2</p>
        <p>According to this de nition, high values of vi denote a
low degree of overlap in the cliques of the cluster Cli, while
low values of vi denote a high degree of overlap in the Cli
cliques.</p>
        <p>Density di of a cluster Cli is the degree of interconnection
among the linked data entities of Cli. The density coe
cient di = 2 Ri=Ni(Ni 1) is the ratio between the number
Ri of links in the cluster Cli and the maximum number of
possible links. According to this de nition, high values of di
denote a high degree of interconnection among the cluster
Cli entities, while low values of di denote a low degree of
interconnection. The prominence Pi of a cluster Cli is
calculated on the basis of its variability and density as follows:
Pi =
2 (1
(1</p>
        <p>vi) di
vi) + di</p>
        <p>According to this approach, most prominent clusters are
those which are more focused and homogeneous with respect
to their theme. We graphically represent cluster prominence
by drawing circles proportional to the prominence values of
the corresponding clusters. In our example of Figure 2,
clusters Cl1 and Cl4 are more prominent (larger circles) because
they are more focused and homogeneous. On the opposite,
clusters like Cl3, which collect several entities of di erent
types are considered less prominent (smaller circle).
However, other options are possible for the evaluation of
prominence in case of speci c application needs. A rst option
is to consider a cluster to be more prominent as it is more
close to the seed s of interest. In this case, the prominence
Pi of a cluster Cli is evaluated by taking into account the
average value of similarity between the linked data entities
in the cluster Cli and s, weighted by the frequency of each
entity ni in Cli, as follows:</p>
        <p>Ni</p>
        <p>P
Pi = p=1
(ni; s) fi
Ni
P fi
p=1
where fi denotes the frequency of the linked data entity ni
in the cluster Cli. Another option is to consider the
prominence Pi of a cluster Cli as proportional to the dimension
Ni of Cli and to the size ki of the smaller clique in Cli, as
follows: Pi = 2 Ni ki=Ni + ki.
tension to the multi-repository exploration and to the
multiseed extraction can be performed.
4.3</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Proximity relations</title>
      <p>In an inCloud, clusters (and consequently their associated
essentials) are connected by reciprocal proximity relations,
which represent the degree of overlapping between them.
In particular, given two clusters Cli and Clj, the degree of
proximity Xij =j Cli \ Clj j = j Cli j between Cli and Clj is
proportional to the number of linked data entities common
to Cli and Clj over the number of linked data entities in Cli.
The greater the level of overlapping between Cli and Clj,
the higher the degree of their proximity relation. Proximity
relations are graphically represented by arrows with
thickness proportional to the proximity degree. In Figure 2, we
can see how proximity relations connect those clusters that
are more semantically related to each other, such as Cl2,
Cl4, and Cl5 which all represent di erent types of artworks
by Vincent van Gogh.</p>
    </sec>
    <sec id="sec-10">
      <title>USING INCLOUDS FOR THEMATIC EX</title>
    </sec>
    <sec id="sec-11">
      <title>PLORATION</title>
      <p>In this section, we discuss how inClouds can be exploited
for thematic exploration of linked data and we provide some
considerations about the applicability of the inCloud
approach in the large-scale scenario.
5.1</p>
    </sec>
    <sec id="sec-12">
      <title>Thematic exploration through inClouds</title>
      <p>An inCloud enables di erent exploration modalities that
can be switched on according to the speci c user preferences.
In particular, the following modalities are de ned.</p>
      <p>Exploration-by-essential. This is the most intuitive
exploration modality and it is based on cluster essentials.
A user can consider each essential as a sort of
instantaneous picture of the associated cluster and linked data
therein contained, thus allowing the user to rapidly
choose the most preferred one for starting the
exploration.</p>
      <p>Exploration-by-prominence. This modality allows the
user to organize the exploration according to the
prominence values associated with the clusters. The idea is
to support the user in moving throughout the clusters
according to their relevance with respect to the set of
considered linked data. As discussed in Section 4,
different criteria can be used to calculate the prominence
value. The capability to switch from one criterion to
another allows the user to dynamically re-organize the
inCloud in light of a di erent notion of cluster
prominence.</p>
      <p>Exploration-by-proximity. This modality enables the
user to choose a cluster and to browse its
constellation, by exploiting the proximity relations. When a
user is exploring a certain cluster, the proximity
relations provide indication of its fully/partially
overlapping neighbors, thus suggesting the possible
exploration of clusters that are somehow related in content.
5.2</p>
    </sec>
    <sec id="sec-13">
      <title>Linked data exploration in-the-large</title>
      <p>The presented inCloud approach can be also exploited for
applicability in the large scale scenario. In particular,
exExtension to multi-repository exploration. For a more
complete visualization of the available linked data about
a certain search target, multiple RDF repositories can be
queried to originate a unique, comprehensive inCloud. In
the Linked Data Cloud, the property owl:sameAs is used to
denote when a linked data entity ni belonging to a certain
RDF repository R and another entity nj belonging to a
different repository R0 refer to the same real-world object. In
a multi-repository scenario, the construction of the graph
Gs can take into account the owl:sameAs relations as a sort
of \natural join" operation. The idea is to start the
construction of Gs by querying an initial repository R and to
exploit the owl:sameAs relations to extend the linked data
extraction to other RDF repositories. In particular, the URIs
connected by a owl:sameAs relation are collapsed in a unique
linked data entity of Gs and the extraction/ ltering
operations described in Section 3 are applied to the whole set of
linked data extracted by the considered RDF repositories.
Extension to multi-seed extraction. In some cases, the
user can be interested in exploring the available linked data
about more than one seed of interest. In this framework,
the inCloud mechanism can be used to build a
comprehensive thematic picture that takes into account all the seeds
of interest. In a multi-seed scenario, the starting point is a
set of seeds S = fs1; : : : ; skg. The graph Gs is built by
executing the extraction/ ltering operations of Section 3 for
each element si 2 S. Depending on the seeds of interest,
one or more portions of the graph Gs can be disjoint from
the rest of the graph. In particular, when the seeds in S
are about completely di erent arguments, a separate
independent cluster is generated through aggregation for each
si 2 S. In such a limit case, the usefulness of the inCloud
mechanism for exploration is in the capability of providing
an e ective synthetic essential for each seed si 2 S and in
calculating the relative prominence of each seed with respect
to the others.</p>
      <p>We stress that linked data exploration in-the-large can
require the execution of thematic aggregation techniques over
a starting RDF graph Gs containing a huge number of nodes
(e.g., thousands of linked data entities). The clique
percolation method we use for cluster calculation best performs
when a small-medium number of nodes in the graph Gs is
considered (e.g., hundreds of linked data entities). For
example, in our tests, the CPM over a graph Gs containing
200 nodes takes an execution time of 200ms (considering a
matching threshold th=0.9). For linked data exploration
inthe-large, when 1.000 (or more) nodes are considered, more
e cient clustering algorithms, like hierarchical clustering,
can be exploited (see [3] for further details).
6.</p>
    </sec>
    <sec id="sec-14">
      <title>RELATED WORK</title>
      <p>Problems and solutions more strictly related to our work
are focused either on improving search and retrieval of
information in the Linked Data cloud [14] or on browsing
and presentation of linked data contents [5]. Search and
retrieval is moving from traditional information lookup to
exploratory search, de ned as the activity of nding and
understanding knowledge about a topic of interest by
exploiting aggregation and learning of information in a
social context [11]. In this respect, for example, Sig.ma
(Semantic Information MAshup) [16] retrieves and integrates
linked data, starting from a single URI, by querying the
Web of Data and applying machine learning to the data
found. In a similar direction, structured and
collaborative search engines are being emerging as a promising
solution for presenting the query results in a sort of
structured form and focusing on the understanding of the user
information need. Examples in this eld are Wolfram
Alpha (http://www.wolframalpha.com), Google Wonder Wheel
(http://www.googlewonderwheel.com), and YAGO2 (http:
//www.mpi-inf.mpg.de/yago-naga/yago). Another
category of related work includes approaches aiming at
presenting linked data in a more intuitive way. Examples of
solutions in this respect are [8, 12] and Freebase Parallax (http://
www.freebase.com/labs/parallax/), where tools that help
users in exploring DBpedia and Freebase are presented, not
only via directed links in the RDF dataset, but also via
newly discovered knowledge associations and visual
navigation. These tools exploit aggregation techniques in order
to combine related topics in uni ed nodes, providing also a
textual description of each node. In other approaches, like
Marbles (http://www5.wiwiss.fu-berlin.de/marbles) and
LESS (http://less.aksw.org), information about resources
of interest is presented exploiting HTML and RSS and by
using di erent colors to distinguish sources.</p>
      <p>With respect to the related work, our contribution regards
the use of data similarity, proximity, and prominence
techniques for inCloud construction, to move from a basic, at
organization of linked data to a high-level, thematic view of
them. Moreover, the proposed techniques allow the di
erent themes/topics to directly emerge from the original linked
data and their mutual links, by suggesting also an intuitive
visualization of data contents in terms of essentials, which
synthesize the contents of thematic clusters.</p>
    </sec>
    <sec id="sec-15">
      <title>CONCLUDING REMARKS</title>
      <p>In this paper, we presented inClouds, high-level views of
linked data enabling their thematic exploration. Ongoing
work is focused on nalizing the development of a web
application fully covering the steps of linked data aggregation
and abstraction required for inCloud construction. By
exploiting an initial prototype implementation, we run some
experiments concerning user evaluation of inClouds based
on standard user-oriented evaluation methods for
interactive web search interfaces and systems [10]. Initial results
are promising and inClouds are seen by real users as a valid
support to the satisfaction of users information needs [3].
Moreover, ongoing research activity regards the extension
of the inCloud approach to consider additional kinds of
web data contents, like microdata, microblogging posts, and
news. The idea is to propose inClouds as a comprehensive
exploration tool considering also actual, up-to-date social
web information about the search target for possible fruition
in the framework of event-promoting applications.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <article-title>Linked Data - The Story So Far</article-title>
          .
          <source>Int. Journal on Semantic Web and Information Systems</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ):1{
          <fpage>22</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Castano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrara</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Montanelli</surname>
          </string-name>
          .
          <article-title>Matching Ontologies in Open Networked Systems: Techniques and Applications</article-title>
          .
          <source>Journal on Data Semantics</source>
          , V:
          <volume>25</volume>
          {
          <fpage>63</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Castano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrara</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Montanelli</surname>
          </string-name>
          .
          <article-title>Structured Data Clouding across Multiple Webs</article-title>
          .
          <source>Technical report, Universita degli Studi di Milano</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Castano</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Varese</surname>
          </string-name>
          .
          <article-title>Next Generation Data Technologies for Collective Computational Intelligence, chapter Building Collective Intelligence through Folksonomy Coordination</article-title>
          , pages
          <volume>87</volume>
          {
          <fpage>112</fpage>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Davies</surname>
          </string-name>
          , J. Hat eld, C. Donaher, and
          <string-name>
            <surname>J. Zeitz.</surname>
          </string-name>
          <article-title>User Interface Design Considerations for Linked Data Authoring Environments</article-title>
          .
          <source>In Proc. of the WWW Int. Workshop on Linked Data on the Web (LDOW</source>
          <year>2010</year>
          ), Raleigh,
          <string-name>
            <surname>NC</surname>
          </string-name>
          , USA,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. Everitt. Cluster</given-names>
            <surname>Analysis</surname>
          </string-name>
          .
          <source>Edward Arnold, London, UK, 3rd edition</source>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>W.</given-names>
            <surname>Halb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Raimond</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Hausenblas</surname>
          </string-name>
          .
          <article-title>Building Linked Data for both Humans and Machines</article-title>
          .
          <source>In Proc. of the WWW Int. Workshop on Linked Data on the Web (LDOW</source>
          <year>2008</year>
          ), Beijing, China,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hirsch</surname>
          </string-name>
          et al.
          <article-title>Interactive Visualization Tools for Exploring the Semantic Graph of Large Knowledge Spaces</article-title>
          .
          <source>In Proc. of the IUI Int</source>
          .
          <article-title>Workshop on Visual Interfaces to the Social and the Semantic Web, Sanibel Island</article-title>
          , USA,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          .
          <article-title>Weaving the Pedantic Web</article-title>
          .
          <source>In Proc. of the WWW Int. Workshop on Linked Data on the Web (LDOW</source>
          <year>2010</year>
          ), Raleigh,
          <string-name>
            <surname>NC</surname>
          </string-name>
          , USA,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Leclercq</surname>
          </string-name>
          .
          <article-title>he perceptual evaluation of information systems using the construct of user satisfaction: case study of a large french group</article-title>
          .
          <source>ACM SIGMIS Database</source>
          ,
          <volume>38</volume>
          (
          <issue>2</issue>
          ):
          <volume>27</volume>
          {
          <fpage>60</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Marchionini. Exploratory</surname>
          </string-name>
          <article-title>Search: from Finding to Understanding</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>49</volume>
          (
          <issue>4</issue>
          ):
          <volume>41</volume>
          {
          <fpage>46</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mirizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ragone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          .
          <article-title>Semantic Wonder Cloud: Exploratory Search in DBpedia</article-title>
          .
          <source>In Proc. of the ICWE 2nd Int. Workshop on Semantic Web Information Management (SWIM</source>
          <year>2010</year>
          ), pages
          <fpage>138</fpage>
          {
          <fpage>149</fpage>
          , Vienna, Austria,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Palla</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Derenyi</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Farkas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Vicsek</surname>
          </string-name>
          .
          <article-title>Uncovering the Overlapping Community Structure of Complex Networks in Nature and Society</article-title>
          . Nature,
          <volume>435</volume>
          :
          <fpage>814</fpage>
          {
          <fpage>818</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Petrelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mazumdar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dadzie</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Ciravegna</surname>
          </string-name>
          <article-title>. Multi Visualization and Dynamic Query for E ective Exploration of Semantic Data</article-title>
          .
          <source>In Proc. of the 8th Int. Semantic Web Conference</source>
          , pages
          <volume>505</volume>
          {
          <fpage>520</fpage>
          ,
          <string-name>
            <surname>Chantilly</surname>
            ,
            <given-names>VA</given-names>
          </string-name>
          , USA,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sorrentino</surname>
          </string-name>
          et al.
          <article-title>Schema Normalization for Improving Schema Matching</article-title>
          .
          <source>In Proc. of the 28th Int. ER Conference</source>
          , pages
          <volume>280</volume>
          {
          <fpage>293</fpage>
          ,
          <string-name>
            <surname>Gramado</surname>
          </string-name>
          , Brazil,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>G.</given-names>
            <surname>Tummarello</surname>
          </string-name>
          et al.
          <source>Sig. ma: Live Views on the Web of Data. Web Semantics: Science, Services and Agents on the World Wide Web</source>
          ,
          <volume>8</volume>
          (
          <issue>4</issue>
          ):
          <volume>355</volume>
          {
          <fpage>364</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>