<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Algebraic compositional models for semantic similarity in ranking and clustering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Paolo Annesi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio Storch</string-name>
          <email>storch@uniroma2.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danilo Croce</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Basili</string-name>
          <email>basilig@info.uniroma2.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Computer Science, University of Roma Tor Vergata</institution>
          ,
          <addr-line>Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Although distributional models of word meaning have been widely used in Information Retrieval achieving an e ective representation and generalization schema of words in isolation, the composition of words in phrases or sentences is still a challenging task. Di erent methods have been proposed to account on syntactic structures to combine words in term of algebraic operators (e.g. tensor product) among vectors that represent lexical constituents. In this paper, a novel approach for semantic composition based on space projection techniques over the basic geometric lexical representations is proposed. In the geometric perspective here pursued, syntactic bi-grams are projected in the so called Support Subspace, aimed at emphasizing the semantic features shared by the compound words and better capturing phrase-speci c aspects of the involved lexical meanings. State-of-the-art results are achieved in a well known benchmark for phrase similarity task and the generalization capability of the proposed operators is investigated in a cross-linguistic scenario, i.e. in the English and Italian Language.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        With the rapid development of the World Wide Web and the spread of
humangenerated contents, Information Retrieval (IR) has many challenges in
discovering and exploiting those rich and huge information resources. Semantic search [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
improves search precision and recall by understanding user's intent and the
contextual meaning of concepts in documents and queries. Semantic search extends
the scope of traditional information retrieval paradigms from mere document
retrieval to entity and knowledge retrieval, improving the conventional IR
methods by looking at a di erent perspective, i.e. the meaning of words. However,
the language richness and its intrinsic relation to the world and human activities
make semantic search a very complex task. In a IR system, a user can express
its speci c user need with a natural language query like "... buy a car ...". This
request can be satis ed by documents expressing the abstract concept of buying
something and in particular the focus of the action is a car. This information
can be expressed inside a document collection in many di erent forms, e.g. the
quasi-synonymic expression "... purchase an automobile ...". Accounting on
lexical overlap with respect to the original query, a Bag-of-word based system would
instead retrieve di erent documents, containing expressions such as "... buy a
bag ..." or "... drive a car ...". A proper semantic generalization is thus needed,
in order to derive the correct composition of the target words, i.e. an action like
buy and an object like car.
      </p>
      <p>
        While compositional approaches to language understanding have been largely
adopted, semantic tasks are still challenging for research in Natural Language
Processing. Traditional logic-based approaches (as the Montague's approach in
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) rely on Frege's principle for which the meaning of a sentence is a
function of the meanings of its parts [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The resulting theory allows an algebra
on the discrete propositional symbols to represent the meaning of arbitrarily
complex expressions. Despite the fact that they are formally well de ned,
logicbased approaches have limitations in the treatment of ambiguity, vagueness and
cognitive aspects intrinsically connected to natural language.
      </p>
      <p>
        On the other hand, distributional models early introduced by Schutze [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
rely on the Word Space model. Here semantic uncertainty is managed through
the statistical analysis of large scale corpora. Linguistic phenomena are then
modeled according to a geometrical perspective, i.e. points in a high-dimensional
space representing semantic concepts, such as words, and can be learned from
corpora, in such a way that similar, or related, concepts are near each another
in the space. Methods for constructing representations for phrases or sentences
through vector composition has recently received a wide attention in literature
(e.g. [
        <xref ref-type="bibr" rid="ref15 ref23">15, 23</xref>
        ]). However, vector-based models typically represent isolated words
and ignore grammatical structure [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Such models have thus a limited
capability to model compositional operations over phrases and sentences.
      </p>
      <p>
        In order to overcome these limitations a so-called compositional
distributional semantics (DCS) model is needed and its development is still object of
on-going and controversial research (e.g. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]). A compositional model based
on distributional analysis should provide semantic information consistent with
the meaning assignment that is typical of human subjects. For example, it should
support synonymy and similarity judgments on phrases, rather than only on
single words. The objective should be a measure of similarity between
quasisynonymic complex expressions, such as "... buy a car ..." vs. "... purchase an
automobile ...". Another typical bene t should be a computational model for
entailment, so that the representation for " ... buying something ..." should be
implied by the expression "... buying a car ..." but not by "... buying time ...".
Distributional compositional semantics (DCS) need thus a method to de ne: (1)
a way to represent lexical vectors u and v, for words u; v dependent on the phrase
(r; u; v) (where r is a syntactic relation, such as verb-object), and (2) a metric
for comparing di erent phrases according to the selected representations u, v.
Existing models are still controversial and provide general algebraic operators
(such as tensor products) over lexical vectors.
      </p>
      <p>In this paper, we focus on the geometry of latent semantic spaces by
proposing a novel distributional model for semantic composition. The aim is to model
semantics of syntactic bigrams as projections in lexically-driven subspaces.
Distances in such subspaces (called Support Spaces) emphasize the role of common
features that constraint in "parallel" the interpretation of the involved lexical
meanings and better capture phrase-speci c aspects. In the following evaluations,
operators will be employed to compose word pairs involved in speci c syntactic
structures. This resulting compositions will be evaluated according two di erent
perspectives. First, similarity among compositions will be evaluated with respect
to human annotators' judgments. Then, the operators generalization capability
will be measured in order to prove their applicability in semantic search complex
systems. Moreover the robustness of this Support Spaces based will be con rmed
in a cross-linguistic scenario, i.e. in the English and Italian Language.</p>
      <p>While Section 2 discusses existing methods of compositional distributional
semantics, Section 3 presents our model based on support spaces. Experiments
in Section 4 are used to show the bene cial impact of the proposed model and
the contribution to semantic search systems. Finally, Section 5 derives the
conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        While compositional semantics allows to govern the recursive interpretation of
sentences or phrases, traditional vector space models (as in IR [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]) and, mostly,
semantic space models, such as LSA ([
        <xref ref-type="bibr" rid="ref13 ref7">7, 13</xref>
        ]), represent lexical information in
metric spaces where individual words are represented according to the
distributional analysis of their co-occurrences over a large corpus. Such models are
based on the distributional hypotesis which assumes that words occurring within
similar contexts are semantically similar (Harris in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
      </p>
      <p>
        Semantic spaces have been widely used for representing the meaning of words
or other lexical entities (e.g. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]), with successful applications in lexical
disambiguation ([
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]) or harvesting thesauri (as in Lin [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]). In this work we will refer
to the so-called word-based spaces, in which words are represented by
probabilistic information of their co-occurences calculated in a xed range window over
all sentences. In such models, vector components correspond to the entries f of
the vocabulary V (i.e. to features that are individual words). Weigths are
associated with each component, using di erent estimators of their correlation. In some
works (e.g. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]) pure co-occurrence counts are adopted as weighting functions
fi, where i = 1; :::; N and N = jV j; in other works (e.g. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]), statistical
functions like the pointwise mutual information between the target word w and the
captured co-occurences in the window are used, i.e. pmi(w; i) = log2 p(pw(w);pf(if)i) .
      </p>
      <p>
        A vector w = (pmi1; :::; pmiN ) models a word w and it is thus built over all
the words fi belonging to the dictionary. When w and f never co-occur in any
window their pmi is by default set to 0. Weights of vector components depend
on the size of the co-occurrence window and express the global statistics in the
entire corpus. Larger values of the adopted window size aim to capture topical
similarity (as in the document based models of IR), while smaller sizes
(usually between the 1-3 surrounding words) lead to representation better suited
for paradigmatic similarities between word vectors w. Cosine similarity between
vectors w1 and w2 is modeled as the normalized scalar product, i.e. hw1;w2i
kw1kkw2k
that expresses topical or paradigmatic similarity according to the di erent
representations (e.g. window sizes). Notice that dimensionality reduction methods,
such as LSA [
        <xref ref-type="bibr" rid="ref13 ref7">7, 13</xref>
        ] are also applied in some studies, to capture second order
dependencies between features f , i.e. applying semantic smoothing to possibly
sparse input data. Applications of an LSA-based representation to Frame
Induction or Semantic Role Labeling are presented in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], respectively.
      </p>
      <p>
        The main limitation of distributional models of lexical semantic is their
noncompositional nature: they are based on statistics related to the occurences of
the individual words in the corpus. In such models, the semantic of topological
similarity functions is thus de ned only for the comparison between
individual words. That is the reason why distributional methods can not compute the
meanings of phrases (and sentences) as e ectively as they do indeed over
individual words. Distributional methods have been recently extended to better
account compositionality, in the so called distributional compositional semantics
(DCS) approaches. Mitchell and Lapata in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] follow Foltz [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and assume that
the contribution of the syntactic structure can be ignored, while the meaning
of a phrase is simply the commutative sum of the meanings of its constituent
words. More formally, [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] de nes the composition p = u v of vectors u and
v through an additive class of composition functions expressed by:
p+ = u + v
(1)
This perspective clearly leads to a variety of e cient yet shallow models of
compositional semantics compared in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. For example pointwise multiplication
is de ned by the multiplicative function:
p = u
v
(2)
where the symbol represents multiplication of the corresponding components,
i.e. pi = ui vi. Point-wise multiplication seems to best correspond with the
intended e ects of syntactic interaction, as experiments in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] demonstrate. In
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], the concept of a structured vector space is introduced, where each word is
associated with a set of vectors corresponding to di erent syntactic dependencies.
Every word is thus expressed by a tensor, and tensor operations are imposed.
      </p>
      <p>
        The main di erences among these studies lies in (1) the lexical vector
representation selected (e.g. some authors do not even commit to any representation,
but generically refer to any lexical vector, as in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) as well as in (2) the adopted
compositional algebra, i.e. the system of operators de ned over such vectors.
Generally, proposed operators do not depend on the involved lexical items, but
a general purpose algebra is adopted. Since compositional structures are highly
lexicalized, and the same syntactic relation triggers to very di erent semantic
relations with respect to the di erent involved words, a proposal that makes the
compositionality operators dependent on individual lexical vectors is hereafter
discussed.
      </p>
    </sec>
    <sec id="sec-3">
      <title>A quantitative model for compositionality</title>
      <p>In order to determine the semantic analogies and di erences between two phrases,
such as "... buy a car ..." and "... buy time ...", a distributional compositional
model is employed as follows. The involved lexicals are buy, car and time, while
their corresponding vector representation will be denoted by wbuy wcar and
wtime. The major result of most studies on DCS is the de nition of the function
that associates with wbuy and wcar a new vector wbuy car = wbuy wcar.</p>
      <p>We consider this approach misleading since vector components in the word
space are tied to the syntactic nature of the composed words and the new vector
wbuy car should not have the same type of the original vectors. Notice also that
the components of wbuy and wcar express all their contexts, i.e. interpretations,
and thus senses, of buy and car in the corpus. Algebric operations are thus open
to misleading contributions, brought by not-null feature scores of buyi vs. carj
(i 6= j) that may correspond to senses of buy and car that are not related to the
speci c phrase "buy a car ". On the contrary, in a composition, such as the
verbobject pair (buy; car), the word car in uences the interpretation of the verb buy
and viceversa. The model here proposed is based on the assumption that this
in uence can be expressed via the operation of projection into a subspace, i.e.
a subset of original features fi. A projection is a mapping (a selection function)
over the set of all features. A subspace generated by a projection function
local to the (buy; car) phrase can be found such that only the features speci c
to the phrase meaning are selected and the irrelevant ones are neglected. The
resulting subspace has to preserve the compositional semantics of the phrase and
it is called support subspace of the underlying word pair.</p>
      <p>Consider the bigram composed of the words Buy-Car Buy-Time
buy and car and their vectorial representation cheap::Adj consume::V
in a co-occurrence N dimensional Word Space. insurance::N enough::Adj
Table 1 reports the k = 10 features with the rent::V waste::V
highest contributions of the point wise product lease::V save::In
of the pairs (buy,car) and (buy,time). The sup- dealer::N permit::N
port space thus selects the most important fea- motorcycle::N stressful::Adj
tures for both words, e.g. buy:V and car:N. No- hire::V spare::Adj
tice that this captures the conjunctive nature of auto::N save::V
the scalar product to which contributions come california::Adj warner::N
from feature with non zero scores in both vec- tesco::N expensive::Adj
tors. It is clear that the two pairs give rise to dif- Table 1. Features
correspondferent support subspaces: the main components ing to dimensions in the k=10
related with buy car refer mostly to the automo- dimensional support space of
bile commerce area unlike the ones related with bigrams buy car and buy time
buy time mostly referring to the time wasting or
saving. Similarity judgments about a pair can be thus better computed within
its support subspace.</p>
      <p>More formally k dimensional support subspace for a word pair (u; v) (with
k N ) is the subspace spanned by the subset of n k indexes Ik(u; v) =
fi1; :::; ing for which Ptn=1 uit vit is maximal. Given two pairs the similarity
between syntactic equivalent words (e.g. nouns with nouns, verbs with verbs)
is measured in the support subspace derived by applying a speci c projection
function. Compositional similarity between buy car and the latter pairs (e.g.
buy time) is thus estimated by (1) immersing wbuy and wtime in the selected
". . . buy car . . . " support subspace and (2) estimating similarity between
corresponding arguments of the pairs locally in that subspace. Therefore the similarity
between syntactic equivalent words (e.g. car with time) within these new
subspace is measured.</p>
      <p>Therefore given a pair (u; v), a unique matrix Mkuv = (mkuv)ij is de ned for a
given projection k(u; v) into the k-dimensional support space of any pair (u; v)
according to the following de nition:
(mkuv)ij =
(1 i i = j 2 Ik(u; v)</p>
      <p>0 otherwise.</p>
      <p>The vector u~ projected in the support subspace can be thus estimated through
the following matrix operation:
u~ =
k(u; v)
u~ = Mkuvu</p>
      <p>A special case of the projection matrix is given when no k limitation is
imposed to the dimension and all the positive addends in the scalar product are
taken. Notice also that two pairs p1 = (u; v) and p2 = (u0; v0) give rise to two
di erent projections denoted by M1k and M2k and de ned as:
(Left projection)
1k =
k(u; v)</p>
      <p>(Right projection)
It is also possible to de ne a unique symmetric projection
the combined matrix M1k2 as follows:
2k =</p>
      <p>k(u0 ; v0 ) (5)
1k2 corresponding to</p>
      <p>M1k2 = (M1k + M2k) (M1kM2k)
where the mutual components that satisfy Eq. 3 are employed as M1k2.
As 1 is the projection in the support subspace for the pair p1, it is possible to
immerse the latter pair p2 by applying Eq. 4. This results in the two
vectors M1ku0 and the M1kv0 . It follows that a compositional similarity judgment
between two phrase over the rst pair support subspace can be expressed as:
(3)
(4)
(6)
(7)
(p1)(p1; p2) =
(1 )(p1; p2) =
hM1ku; M1ku0 i
M1ku</p>
      <p>M1ku0
hM1kv; M1kv0 i
M1kv</p>
      <p>
        M1kv0
bination of (1 )(p1; p2) and (2 )(p1; p2) as:
where rst cosine similarity between syntactically correlated vectors in the
selected support subspaces are computed and then a composition function , such
as the sum or the product, is applied. Compositional function over the
latter support subspace evoked by the pair p2 can be correspondingly denoted by
(2 )(p1; p2). A symmetric composition function can thus be obtained as a
com(12)(p1; p2) =
(1 )(p1; p2)
(2 )(p1; p2)
(8)
where the composition function (again the sum or the product) between the
similarities over the left and right support subspaces is applied. Notice how the
left and right composition operators ( ) may di er from the overall composition
operator . More details are discussed in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Evaluation</title>
      <p>
        This experimental evaluation aims to estimate the e ectiveness of the proposed
class of projection based methods in capturing similarity judgments over phrases
and syntactic structures. In particular, a rst evaluation is carried out to measure
the correlation of the operator outcomes with judgments provided by human
annotators. The generalization capability of the operators is measured in the
second evaluation in order to prove their applicability in semantic search complex
systems. Moreover the latter experiments are carried out in a cross-language
setting, i.e. for english and italian datasets.
The rst evaluation is carried out over the dataset proposed by [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], which is part
of the GEMS 2011 Shared Evaluation. It consists of a list of 5,833 adjective-noun
(AdjN), verb-object (VO) or noun-noun (NN) pairs, rated with scores ranging from
1 The corpus is developed by the WaCky community and it is available in the Wacky
project web page at http://medialab.di.unipi.it/Project/QA/wikiCoNLL.bz2
1 to 7. In Table 2, examples of pairs and scores are shown. The correlation of
the similarity judgements outputed by a DCS model against the human
judgements is computed using Spearman's , a non-parametric measure of statistical
dependence between two variables proposed by [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
In this second evaluation, the generalization capability of the employed operators
will be investigated. A verb (e.g. perform) can be more or less semantically
close to another verb (e.g. other verbs like solve, or produce) depending on the
context in which it appears. The verb-object (VO) composition speci es the verb's
meaning by expressing one of its selectional preferences, i.e. its object. In this
scenario, we expect that a pair such as perform task will be more similar to
solve issue, as they both re ect an abstract cognitive action, with respect to
a pair like produce car, i.e. a concrete production. This kind of generalization
capability is crucial to e ectively use this class of operators in a QA scenario by
enabling to rank results according to the complex representations of the question.
Moreover, both English and Italian languages can be considered to demonstrate
the impact in a cross language setting. Figure 4 shows a manually developed
dataset. It consists of 24 VO word pairs in English and Italian, divided into 3
di erent semantic classes: Cognitive, Ingest Liquid and Fabricate.
      </p>
      <p>English Italian
perform task svolgere compito
solve issue risolvere questione
handle problem gestire problema
use method applicare metodo
suggest idea suggerire idea
determine solution trovare soluzione
spread knowledge divulgare conoscenza
start argument iniziare ragionamento
drink water bere acqua
ingest syrup ingerire sciroppo</p>
      <p>pour beer versare birra
swallow saliva inghiottire saliva
assume alcohol assumere alcool
taste wine assaggiare vino
sip liquor assaporare liquore
take co ee prendere ca
produce car produrre auto
complete construction completare costruzione
fabricate toy fabbricare giocattolo
build tower edi care torre
assemble device assemblare dispositivo
construct building costruire edi cio
manufacture product realizzare prodotto</p>
      <p>create artwork creare opera</p>
      <p>Table 4. Cross-linguistic dataset
Ingest Liquid
Fabricate
Semantic Class</p>
      <p>Cognitive</p>
      <p>This evaluation aims to measure how the proposed compositional operators
group together semantically related word pairs, i.e. those belonging to the same
class, and separate the unrelated pairs. Figure 1 shows the application of two
models, the Additive (eq. 1) and Support Subspace (Eq. 8) ones that achieve
the best results in the previous experiment. The two languages are reported in
di erent rows. Similarity distribution between the geometric representation of
verb pair, with no composition, has been investigated as a baseline. For each
language, the similarity distribution among the possible 552 verb pairs is estimated
and two distributions of the infra and intra-class pairs are independently
plotted. In order to summarize them, a Normal Distribution N ( ; 2) of mean
and variance 2 are employed. Each point represents the percentage p(x) of
pairs in a group that have a given similarity value equal to x. In a given class,
the VO-VO pairs of a DCS operator are expected to increase this probability with
respect to the baseline pairs V-V of the same set. Viceversa, for pairs belonging
to di erent classes, i.e. intra-class pairs. The distributions for the baseline
control set (i.e. Verbs Only, V-V) are always depicted by dotted lines, while
DCS operators are expressed in continuous line.</p>
      <p>Notice that the overlap between the curves of the infra and intra-class
pairs corresponds to the amount of ambiguity in deciding if a pair is in the
same class. It is the error probability, i.e. the percentage of cases of one group
that by chance appears to have more probability in the other group. Although
the actions described by di erent classes are very di erent, e.g. Ingest Liquid
vs. Fabricate, most verbs are ambiguous: contextual information is expected
to enable the correct decision. For example, although the class Ingest Liquid
is clearly separated with respect to the others, a verb like assume could well be
classi ed in the Cognitive class, as in assume a position.</p>
      <p>(a)  English  AddiIve  </p>
      <p>(b)  English  Support  Subspace  
Verbs  Only  Rel  
Verbs  Only  Unrel  
AddiIve  Rel  </p>
      <p>AddiIve  Unrel  
Verbs  Only  Rel  
Verbs  Only  Unrel  
AddiIve  Rel  
AddiIve  Unrel  
Verbs  Only  Rel  
Verbs  Only  Unrel  
Sub-­‐space  Rel  
Sub-­‐space  Unrel  
Verbs  Only  Rel  
Verbs  Only  Unrel  
Sub-­‐space  Rel  
Sub-­‐space  Unrel  </p>
      <p>The outcome of the experiment is that DCS operators are always able to
increase the gap in the average similarity of the infra vs. intra-class pairs. It
seems that the geometrical representation of the verb is consistently changed as
most similarity distributions suggest. The compositional operators seem able to
decrease the overlap between di erent distributions, i.e. reduce the ambiguity.</p>
      <p>Figure 1 (a) and (c) report the distribution of the ML additive operator,
that achieves an impressive ambiguity reduction, i.e. the overlap between curves
is drastically reduced. This phenomenon is further increased when the Support
Subspace operator is employed as shown in Figure 1 (b) and (d): notice how
the mean value of the distribution of semantically related word is signi cantly
increased for both languages.</p>
      <p>The probability of error reduction can be computed against the control
groups. It is the decrease of the error probability of a DCS relative to the same
estimate for the control (i.e. V-V) group. It is a natural estimator of the
generalization capability of the involved operators. In Table 5 the intersection area for
all the models and the decrement of the relative probability of error are shown.
For English, the ambiguity reduction of the Support Subspace operator is of 91%
with respect to the control set. This is comparable with the additive operator
results, i.e. 92:3%. It con rms the ndings of the previous experiment where the
di erence between these operators is negligible. For Italian, the generalization
capability of support subspace operator is more stable, as its error reduction is
of 62:9% with respect to the additive model, i.e. 54:2%.</p>
      <sec id="sec-4-1">
        <title>Model</title>
        <p>VerbOnly</p>
      </sec>
      <sec id="sec-4-2">
        <title>English Italian</title>
        <p>Probability Ambiguity Probability Ambiguity
of Error Decrease of Error Decrease</p>
        <p>
          .401 - .222
In this paper, a distributional compositional semantic model based on space projection
guided by syntagmatically related lexical pairs is de ned. Syntactic bi-grams are here
projected in the so called Support Subspace and compositional similarity scores are
correspondingly derived. This represents a novel perspective on compositional models
over vector representations with respect to shallow vector operators (e.g. additive or
multiplicative operations) as proposed in literature, e.g. in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The presented approach
focuses on selecting the most important components for a speci c word pair involved
in a syntactic relation in order to have a more accurate estimator of their similarity.
        </p>
        <p>
          The proposed method have been evaluated over the well known dataset in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]
achieving results close to the average human interannotator agreement scores. A rst
applicability study of such compositional models in typical IR systems was carried
out. The operators' generalization capability was measured proving that compositional
operators can e ectively separate phrase structure in di erent semantic clusters. The
robustness of such operators has been also con rmed in a cross-linguistic scenario, i.e.
in the English and Italian Language. Future work on other compositional prediction
tasks (e.g. selectional preference modeling) and over di erent datasets will be carried
out to better assess and generalize the presented results.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Annesi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Storch</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
          </string-name>
          , R.:
          <article-title>Space projections as distributional models for semantic composition (2012), submitted for publication</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>B.</given-names>
            <surname>Coecke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Mathematical foundations for a compositional distributed model of meaning</article-title>
          .
          <source>Lambek Festschirft, Linguistic Analysis</source>
          , vol.
          <volume>36</volume>
          36 (
          <year>2010</year>
          ), http://arxiv.org/submit/10256/preview
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciaramita</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mika</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaragoza</surname>
          </string-name>
          , H.:
          <article-title>Towards semantic search</article-title>
          .
          <source>Natural Language and Information</source>
          Systems pp.
          <volume>4</volume>
          {
          <issue>11</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferraresi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanchetta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The wacky wide web: a collection of very large linguistically processed web-crawled corpora</article-title>
          .
          <source>Language Resources And Evaluation</source>
          <volume>43</volume>
          (
          <issue>3</issue>
          ),
          <volume>209</volume>
          {
          <fpage>226</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zamparelli</surname>
          </string-name>
          , R.:
          <article-title>Nouns are vectors, adjectives are matrices: representing adjective-noun constructions in semantic space</article-title>
          .
          <source>In: Proceedings of EMNLP 2010</source>
          . pp.
          <volume>1183</volume>
          {
          <fpage>1193</fpage>
          . EMNLP '
          <volume>10</volume>
          ,
          <string-name>
            <surname>Stroudsburg</surname>
          </string-name>
          , PA, USA (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giannone</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Annesi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
          </string-name>
          , R.:
          <article-title>Towards open-domain semantic role labeling</article-title>
          .
          <source>In: ACL</source>
          . pp.
          <volume>237</volume>
          {
          <issue>246</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Deerwester</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furnas</surname>
            ,
            <given-names>G.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harshman</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          :
          <article-title>Indexing by latent semantic analysis</article-title>
          .
          <source>JASIS</source>
          <volume>41</volume>
          (
          <issue>6</issue>
          ),
          <volume>391</volume>
          {
          <fpage>407</fpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Erk</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pad</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A structured vector space model for word meaning in context (</article-title>
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Foltz</surname>
            ,
            <given-names>P.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kintsch</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          , L, T.K.:
          <article-title>The measurement of textual coherence with latent semantic analysis (</article-title>
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Frege</surname>
          </string-name>
          , G.:
          <article-title>Uber sinn und bedeutung</article-title>
          .
          <source>Zeitschrift fur Philosophie und philosophische Kritik</source>
          <volume>100</volume>
          ,
          <issue>25</issue>
          {50, translated, as `On Sense and Reference', by Max Black
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Grefenstette</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sadrzadeh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Experimental support for a categorical compositional distributional model of meaning</article-title>
          .
          <source>CoRR abs/1106</source>
          .4058 (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>Z.S.</given-names>
          </string-name>
          : Mathematical Structures of Language. Wiley, New York, NY, USA (
          <year>1968</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dutnais</surname>
          </string-name>
          , S.T.:
          <article-title>A solution to platos problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge</article-title>
          .
          <source>Psychological</source>
          review pp.
          <volume>211</volume>
          {
          <issue>240</issue>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Automatic retrieval and clustering of similar word</article-title>
          .
          <source>In: Proceedings of COLING-ACL</source>
          . Montreal, Canada (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Mitchell, J.,
          <string-name>
            <surname>Lapata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Vector-based models of semantic composition</article-title>
          .
          <source>In: In Proceedings of ACL-08: HLT</source>
          . pp.
          <volume>236</volume>
          {
          <issue>244</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. Mitchell, J.,
          <string-name>
            <surname>Lapata</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Composition in distributional models of semantics</article-title>
          .
          <source>Cognitive Science</source>
          <volume>34</volume>
          (
          <issue>8</issue>
          ),
          <volume>1388</volume>
          {
          <fpage>1429</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Montague</surname>
          </string-name>
          , R.: Formal Philosophy: Selected Papers of Richard Montague. Yale University Press (
          <year>1974</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Pantel</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Document clustering with committees</article-title>
          .
          <source>In: SIGIR-02</source>
          . pp.
          <volume>199</volume>
          {
          <issue>206</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Pennacchiotti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>D.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Automatic induction of framenet lexical units</article-title>
          .
          <source>In: EMNLP</source>
          . pp.
          <volume>457</volume>
          {
          <issue>465</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A vector space model for automatic indexing</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>18</volume>
          ,
          <issue>613</issue>
          {
          <fpage>620</fpage>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21. Schutze, H.:
          <article-title>Word space</article-title>
          . In: Hanson,
          <string-name>
            <given-names>S.J.</given-names>
            ,
            <surname>Cowan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.D.</given-names>
            ,
            <surname>Giles</surname>
          </string-name>
          , C.L. (eds.) NIPS 5, pp.
          <volume>895</volume>
          {
          <fpage>902</fpage>
          . Morgan Kaufmann Publishers, San Mateo CA (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. Schutze, H.:
          <article-title>Automatic Word Sense Discrimination</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>24</volume>
          ,
          <issue>97</issue>
          {
          <fpage>124</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Turney</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pantel</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>From frequency to meaning: Vector space models of semantics</article-title>
          .
          <source>Journal of arti cial intelligence research 37</source>
          ,
          <volume>141</volume>
          (
          <year>2010</year>
          ), doi:10.1613/jair.2934
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>