<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Semantic Clouding Approach for Cross-Webs Interoperability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Milano</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>castano</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ferrara</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>montanelli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>varese}@dico.unimi.it</string-name>
        </contrib>
      </contrib-group>
      <fpage>22</fpage>
      <lpage>26</lpage>
      <abstract>
        <p>The classical vision of the Web as a merely publishing environment for information-consuming users is being replaced by a plural vision where multiple webs, like Web 2.0, Social Web, and Semantic Web, co-exist and interoperate to make information sharing more effective and socially pervasive. In this paper, we propose a semantic clouding approach for the construction of cross-web, disciplined, and intuitive information organization structures called i-clouds. An overview of the proposed semantic clouding approach is presented in the paper, as well as an example of i-cloud over real web resources about a movie dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Over the recent years, the classical vision of the Web as a merely publishing
environment for information-consuming users is being replaced by a plural vision
where multiple webs, like Web 2.0, Social Web, and Semantic Web, co-exist and
interoperate to make information sharing more effective and socially pervasive. The
experience of active research projects like OKKAM and Linked Data places the
accent on the growing need to recognize identity and similarity relations between data
descriptions provided by different web sources in different domains. The variety of
webs data, spanning from textual tags to RDF(S) structural descriptions up to formal
OWL instances, makes the above mentioned need of identity/similarity recognition
even more crucial and challenging. In such a complex scenario, a new generation of
information search techniques is required to cope with the following needs: i) the
capability to span across multiple Webs, to properly consider the wide variety of
available web resources and pieces of knowledge by properly assessing their
information contribution nature; ii) the capability to anticipate the user needs by
providing a focused but comprehensive set of web resources prominent for his/her
target; iii) the capability to semantically organize all retrieved prominent resources
into an intuitive and coherent structure [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ].
      </p>
      <p>
        In this paper, we propose a semantic clouding approach for the construction of
cross-web, disciplined, and intuitive information organization structures called
iclouds. An i-cloud is built to organize all the web resources about a certain target
entity of interest into a graph on the basis of their level of prominence and reciprocal
closeness. An overview of the proposed semantic clouding approach is presented in
the paper, as well as an example of i-cloud over real web resources about movies. A
more technical discussion about construction and formal properties of i-clouds is
provided in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Semantic clouding of web resources</title>
      <p>An i-cloud is built around a certain target entity, which is a keyword-based
representation of a topic of interest, namely a real-world object/person, an event, a
situation, or any similar subject that can be of interest for the user. The notions of
closeness, and prominence are define for an i-cloud to capture how similar web
resources are each other and with respect to the target entity of the i-cloud, and the
relative importance of a resource within the i-cloud, respectively. The following
properties characterize i-clouds:
 Cross-webness. An i-cloud collects web resources coming from multiple webs to
provide a comprehensive picture of all the available information, both objective
and subjective, about the specified target entity for which the i-cloud is built.
 Discipliness. The web resources in an i-cloud are not only those directly related to
target entity (i.e., those trivially matching the target entity) but also those that are
in some way related to the target and are close to it.
 Intuitiveness. The i-cloud organization borrows the graphical representation
commonly used for folksonomies and tag-clouds. This supports the user in
browsing the i-cloud more effectively according to closeness and prominence of
the web resources therein contained.</p>
      <p>For i-cloud construction, we propose a semantic clouding approach articulated in
three main phases (see Fig. 1): acquisition of web resources, classification of web
resources, and clouding of web resources.</p>
      <p>Acquisition of web resources. For semantic clouding, all the different web
resources are acquired from their respective source webs according to a reference data
model called WDI model. The WDI model is capable of dealing with a variety of web
resources. In particular, a WDI representation is provided for tagged resources that
are resources from social annotation systems like Delicious, microdata resources that
are resources from microblogging systems like Twitter, and semantic web resources
that are resources from RDF/OWL knowledge repositories like Freebase. Each web
resource wr is stored in a support repository called WDI repository in the form of a
web data item wdi(wr) where terminological, structural and logical information about
wr are properly represented.</p>
      <p>
        Classification of web resources. The acquired web resources are grouped together
according to their level of closeness. To this end, tailored matching techniques have
been developed in the framework of the HMatch 2.0 system [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. First, term and
structural matching techniques are adopted to calculate the level of closeness CC(wdii,
wdij) between any pair of web data items wdii and wdij in the WDI repository. Then, a
hierarchical clustering procedure is adopted to determine a closeness tree where all
the wdis are properly grouped according to their closeness coefficients previously
calculated.
      </p>
      <p>Clouding of web resources. Given a target entity e specified by the user, an
icloud is built for e by firstly extracting from the WDI repository a ground set of web
data items that syntactically match e. Given a closeness threshold, the wdis in the
ground set are used to select a number of candidate clusters in the closeness tree,
namely all the clusters containing the wdis of the ground set and the wdis whose
closeness is greater than or equal to the threshold. The candidate clusters originate the
graph structure of the resulting i-cloud through a graph construction procedure. A
labelling function is finally applied to assign to nodes and edges of the i-cloud their
corresponding closeness and prominence values, respectively.</p>
    </sec>
    <sec id="sec-3">
      <title>An example of i-cloud for cross-web interoperability</title>
      <p>As an example, we consider the i-cloud of Fig. 2 where a number of web resources
related to the target entity “Star Wars” are shown. We can observe that resources in
the i-cloud are not only those directly related to this popular movie, such as the titles
of the six movies of the Star Wars saga, but also resources that are close to the movie
saga even if not directly matching the target, such as some of the most important
characters in the movies. The dimension of each node in the i-cloud is proportional to
the prominence of the corresponding web resource for “Star Wars” and the edges
connecting the nodes are labelled with their closeness degree. We observe that
different kinds of web resources populate the “Star Wars” i-cloud. In particular, this
icloud is built over resources acquired from the Delicious annotation systems (e.g.,
wdi(delicious1)), the Twitter microblogging system (e.g., wdi(twitter1)), and the Freebase
Linked Data repository (e.g., wdi(iimb1)).</p>
      <p>
        In this example of i-cloud, the prominence of the various web resources is
calculated through a popularity-based mechanism. This means that the prominence of
a resource wr depends on the “centrality” of wr with respect to the i-cloud and it
corresponds to the degree of connection of wr with the other nodes in the graph of the
i-cloud [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Other techniques can be used for prominence computation based on the
provenance of the web resources in the i-cloud (e.g., [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
A more detailed description of the approach and related support techniques can be
found in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. A prototype for i-cloud construction has been developed on top of the
HMatch 2.0 environment (http://islab.dico.unimi.it/hmatch) and it has been evaluated on the
OAEI-IIMB 2010 dataset (http://www.instancematching.org/oaei/imei2010.html). The
positive results we obtained during evaluation encourages to continue working on
icloud research issues. In particular, a focused search application is being developed in
the domain of tourism and entertainment related to the city of Milan.
5
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kuo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hentrich</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Good</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilkinson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Tag Clouds for Summarizing Web Search Results</article-title>
          .
          <source>In: Proc. of the 16th Int. Conference on World Wide Web (WWW</source>
          <year>2007</year>
          ). Banff, Alberta, Canada,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Koutrika</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zadeh</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Molina</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Data Clouds: Summarizing Keyword Search Results over Structured Data</article-title>
          .
          <source>In: Proc. of the 12th Int. Conference on Extending Database Technology (EDBT</source>
          <year>2009</year>
          ). Saint Petersburg, Russia,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Castano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montanelli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Semantic Data Clouding across Multiple Webs</article-title>
          . Submitted to Information Systems. Elsevier,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Castano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montanelli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Dealing with Matching Variability of Semantic Web Data Using Contexts</article-title>
          .
          <source>In: Proc. of the 22nd Int. Conference on Advanced Information Systems Engineering (CAiSE'10)</source>
          . Hammamet, Tunisia,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Easley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleinberg</surname>
          </string-name>
          , J.: Networks, Crowds, and
          <article-title>Markets: Reasoning About a Highly Connected World</article-title>
          . Cambridge University Press,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gil</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Artz</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Towards Content Trust of Web Resources</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>5</volume>
          (
          <issue>4</issue>
          ),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>