<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Distributional Semantic Web Stack</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andr´e Freitas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edward Curry</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Siegfried Handschuh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight Centre for Data Analytics, National University of Ireland</institution>
          ,
          <addr-line>Galway</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science and Mathematics, University of Passau</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The capacity of distributional semantic models (DSMs) to discover similarities over large scale heterogeneous and poorly structured data brings them as a promising universal and low-effort framework to support semantic approximation and knowledge discovery. This position paper explores the role of distributional semantics in the Semantic Web vision, based on state-of-the-art distributional-relational models, categorizing and generalizing existing approaches into a Distributional Semantic Web stack.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Distributional semantics is based on the idea that semantic information can be
extracted from lexical co-occurrence from large-scale data corpora. The
simplicity of its vector space representation, its ability to automatically derive meaning
from large-scale unstructured and heterogeneous data and its built-in
semantic approximation capabilities are bringing distributional semantic models as a
promising approach to bring additional flexibility into existing knowledge
representation frameworks.</p>
      <p>Distributional semantic approaches are being used to complement the
semantics of structured knowledge bases, generating hybrid distributional-relational
models. These hybrid models are built to support semantic approximation, and
can be applied to selective reasoning mechanisms, reasoning over incomplete
KBs, semantic search, schema-agnostic queries over structured knowledge bases
and knowledge discovery.
Distributional semantic models (DSMs) are semantic models which are based on
the statistical analysis of co-occurrences of words in large corpora. Distributional
semantics allows the construction of a quantitative model of meaning, where the
degree of the semantic association between different words can be quantified
in relation to a reference corpus. With the availability of large Web corpora,
comprehensive distributional models can effectively be built.</p>
      <p>DSMs are represented as a vector space model, where each dimension
represents a context C for the linguistic or data context in which the target term T
occurs. A context can be defined using documents, co-occurrence window sizes
(number of neighboring words or data elements) or syntactic features. The
distributional interpretation of a target term is defined by a weighted vector of the
contexts in which the term occurs, defining a geometric interpretation under a
distributional vector space. The weights associated with the vectors are defined
using an associated weighting scheme W, which can re-calibrates the relevance
of more generic or discriminative contexts. A semantic relatedness measure S
between two words in the dataset can be calculated by using different
similarity/distance measures such as the cosine similarity or Euclidean distance. As
the dimensionality of the distributional space can grow large, dimensionality
reduction approaches d can be applied.</p>
      <p>Different DSMs are built by varying the parameters of the tuple (T , C, W, d, S).
Examples of distributional models are Latent Semantic Analysis, Random
Indexing, Dependency Vectors, Explicit Semantic Analysis, among others.
Distributional semantic models can be specialized to different application areas using
different corpora.
3</p>
      <p>Distributional-Relational Models (DRMs)
Distributional-Relational Models (DRMs) are models in which the semantics of
a structured knowledge base (KB) is complemented by a distributional semantic
model.</p>
      <p>A Distributional-Relational Model (DRM) is a tuple (DSM, KB, RC, F , H, OP),
where: DSM is the associated distributional semantic model ; KB is the
structured dataset, with elements E and tuples Ω; RC is the reference corpora which
can be unstructured, structured or both. The reference corpora can be internal
(based on the co-occurrence of elements within the KB) or external (a separate
reference corpora); F is a map which translates the elements ei ∈ E into vectors
−→
ei in the the distributional vector space V SDSM using the natural language
label and the entity type of ei; H is a set of threshold values for S above which
−→
two terms are considered to be equivalent; OP is a set of operations over ei in
V SDSM and over E and Ω in the KB. The set of operations may include search,
query and graph navigation operations using the distance measure S.</p>
      <p>
        The DRM supports a double perspective of semantics, keeping the
finegrained precise semantics of the structured KB but also complementing it with
the distributional model. Two main categories of DRMs and associated
applications can be distinguished:
Semantic Matching &amp; Commonsense Reasoning: In this category the RC
is unstructured and it is distinct from the KB. The large-scale unstructured RC
is used as a commonsense knowledge base. Freitas &amp; Curry [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] define a DRM
(τ − Space) for supporting schema-agnostic queries over the structured KB:
terms used in the query are projected into the distributional vector space and
are semantically matched with terms in the KB via distributional semantics
using commonsense information embedded on large scale unstructured corpora
RC. In a different application scenario, Freitas et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] uses the τ − Space to
support selective reasoning over commonsense KBs. Distributional semantics is
used to select the facts which are semantically relevant under a specific reasoning
context, allowing the scoping of the reasoning context and also coping with
incomplete knowledge of commonsense KBs. Pereira da Silva &amp; Freitas [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] used
the τ − Space to support approximate reasoning on logic programs.
Knowledge Discovery: In this category, the structured KB is used as a
distributional reference corpora (where RC = KB). Implicit and explicit semantic
associations are used to derive new meaning and discover new knowledge. The
use of structured data as a distributional corpus is a pattern used for knowledge
discovery applications, where knowledge emerging from similarity patterns in the
data can be used to retrieve similar entities and expose implicit associations. In
this context, the ability to represent the KB entities’ attributes in a vector space
and the use of vector similarity measures as way to retrieve and compare similar
entities can define universal mechanisms for knowledge discovery and semantic
approximation. Novacek et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] describe an approach for using web data as a
bottom-up phenomena, capturing meaning that is not associated with explicit
semantic descriptions, applying it to entity consolidation in the life sciences
domain. Speer et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] proposed AnalogySpace, a DRM over a commonsense KB
using Latent Semantic Indexing targeting the creation of the analogical closure
of a semantic network using dimensional reduction. AnalogySpace was used to
reduce the sparseness of the KB, generalizing its knowledge, allowing users to
explore implicit associations. Cohen et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] introduced PSI, a
predicationbased semantic indexing for biomedical data. PSI was used for similarity-based
retrieval and detection of implicit associations.
4
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Distributional Semantic Web Stack</title>
      <p>DRMs provide universal mechanisms which have fundamental features for
semantic systems: (i) built-in semantic approximation for terminological and
instance data; (ii) ability to use large-scale unstructured data as commonsense
knowledge, (iii) ability to detect emerging implicit associations in the KB, (iv)
simplicity of use supported by the vector space model abstraction, (v)
robustness with regard to poorly structured, heterogeneous and incomplete data. These
features provide a framework for a robust and easy-to-deploy semantic
approximation component grounded on large-scale data. Considering the relevance of
these features in the deployment of semantic systems in general, this paper
synthesizes its vision by proposing a Distributional Semantic Web stack abstraction
(Figure 1), complementing the Semantic Web stack. At the bottom of the stack,
unstructured and structured data can be used as reference corpora together
with the target KB (RDF(S)). Different elements of the distributional model
are included as optional and composable elements of the architecture. The
approximate search and query operations layer access the DSM layer, supporting
users with semantically flexible search and query operations. A graph navigation
layer defines graph navigation algorithms (e.g. such as spreading activation,
bi-directional search) using the semantic approximation and the distributional
information from the layers below.</p>
      <p>Application
Graph Navigation Model
Approximate Search / Query</p>
      <p>c0
cn context space semantic relatedness =</p>
      <p>cos(θ) = daughter . child = 0.234
Distributional Semantic Model
Reference Data
same tCeormrpus context c termcontext c
Reality
..</p>
      <p>distribution of
term associations
... a child (plural: children) is a human between ...</p>
      <p>...is the third child and only son of Prince ...
... was the first son and last child of King ...</p>
      <p>...</p>
      <p>Cosine similarity
Euclidean distance</p>
      <p>Jaccard distance</p>
      <p>Chebyshev
Correlation</p>
      <p>...</p>
      <p>TF/IDF</p>
      <p>Okapi BM25
Mutual Information</p>
      <p>T-Test
χ2
...</p>
      <p>Acknowledgment: This publication was supported in part by Science Foundation
Ireland (SFI) (Grant Number SFI/12/RC/2289) and by the Irish Research Council.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Curry</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <article-title>Natural Language Queries over Heterogeneous Linked Data Graphs: A Distributional-Compositional Semantics Approach</article-title>
          .
          <source>In Proc. of the 19th Intl. Conf. on Intelligent User Interfaces (IUI)</source>
          .
          <article-title>(</article-title>
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Pereira da Silva,
          <string-name>
            <given-names>J.C.</given-names>
            ,
            <surname>Freitas</surname>
          </string-name>
          <string-name>
            <surname>A.</surname>
          </string-name>
          ,
          <article-title>Towards An Approximative Ontology-Agnostic Approach for Logic Programs</article-title>
          ,
          <source>In Proc. of the 8th Intl. Symposium on Foundations of Information and Knowledge Systems</source>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Pereira Da Silva,
          <string-name>
            <given-names>J.C.</given-names>
            ,
            <surname>Curry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Buitelaar</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          ,
          <article-title>A Distributional Semantics Approach for Selective Reasoning on Commonsense Graph Knowledge Bases</article-title>
          .
          <source>In Proc. of the 19th Int .Conf. on Applications of Natural Language to Information Systems (NLDB)</source>
          .
          <article-title>(</article-title>
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Speer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Havasi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lieberman</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <article-title>AnalogySpace: Reducing the Dimensionality of Common Sense Knowledge</article-title>
          .
          <source>In Proc. of the 23rd Intl. Conf. on Artificial Intelligence</source>
          ,
          <fpage>548</fpage>
          -
          <lpage>553</lpage>
          . (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Novacek</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Handschuh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S..</given-names>
          </string-name>
          <article-title>Getting the Meaning Right: A Complementary Distributional Layer for the Web Semantics</article-title>
          .
          <source>In Proc. of the Intl. Semantic Web Conference</source>
          ,
          <volume>504</volume>
          -
          <fpage>519</fpage>
          . (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schvaneveldt</surname>
            ,
            <given-names>R.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rindflesch</surname>
          </string-name>
          , T.C..
          <article-title>Predication-based Semantic Indexing: Permutations as a Means to Encode Predications in Semantic Space</article-title>
          .
          <source>T. AMIA Annu Symp Proc.</source>
          ,
          <volume>114</volume>
          -
          <fpage>118</fpage>
          . (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Turney</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pantel</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <article-title>From frequency to meaning: vector space models of semantics</article-title>
          .
          <source>J. Artif. Int. Res.</source>
          ,
          <volume>37</volume>
          (
          <issue>1</issue>
          ),
          <fpage>141</fpage>
          -
          <lpage>188</lpage>
          . (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Speer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Havasi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lieberman</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <article-title>AnalogySpace: Reducing the Dimensionality of Common Sense Knowledge</article-title>
          .
          <source>In Proc. of the 23rd Intl. Conf. on Artificial Intelligence</source>
          ,
          <fpage>548</fpage>
          -
          <lpage>553</lpage>
          . (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>