<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Encoding syntactic dependencies using Random Indexing and Wikipedia as a corpus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pierpaolo Basile</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Annalina Caputo</string-name>
          <email>acaputog@di.uniba.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Computer Science, University of Bari "Aldo Moro" Via Orabona</institution>
          ,
          <addr-line>4, I-70125, Bari</addr-line>
          ,
          <country country="IT">ITALY</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Distributional approaches are based on a simple hypothesis: the meaning of a word can be inferred from its usage. The application of that idea to the vector space model makes possible the construction of a WordSpace in which words are represented by mathematical points in a geometric space. Similar words are represented close in this space and the de nition of \word usage" depends on the de nition of the context used to build the space, which can be the whole document, the sentence in which the word occurs, a xed window of words, or a speci c syntactic context. However, in its original formulation WordSpace can take into account only one de nition of context at a time. We propose an approach based on vector permutation and Random Indexing to encode several syntactic contexts in a single WordSpace. We adopt WaCkypedia EN corpus to build our WordSpace that is a 2009 dump of the English Wikipedia (about 800 million tokens) annotated with syntactic information provided by a full dependency parser. The e ectiveness of our approach is evaluated using the GEometrical Models of natural language Semantics (GEMS) 2011 Shared Evaluation data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background and motivation</title>
      <p>
        Distributional approaches usually rely on the WordSpace model [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. An overview
can be found in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. This model is based on a vector space in which points are
used to represent semantic concepts, such as words.
      </p>
      <p>The core idea behind WordSpace is that words and concepts are represented
by points in a mathematical space, and this representation is learned from text
in such a way that concepts with similar or related meanings are near to one
another in that space (geometric metaphor of meaning). The semantic similarity
between concepts can be represented as proximity in an n-dimensional space.
Therefore, the main feature of the geometric metaphor of meaning is not that
meanings can be represented as locations in a semantic space, but rather that
similarity between word meanings can be expressed in spatial terms, as proximity
in a high-dimensional space.</p>
      <p>
        One of the great virtues of WordSpaces is that they make very few
languagespeci c assumptions, since just tokenized text is needed to build semantic spaces.
Even more important is their independence from the quality (and the quantity)
of available training material, since they can be built by exploiting an entirely
unsupervised distributional analysis of free text. Indeed, the basis of the WordSpace
model is the distributional hypothesis [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], according to which the meaning of
a word is determined by the set of textual contexts in which it appears. As a
consequence, in distributional models words can be represented as vectors built
over the observable contexts. This means that words are semantically related as
much as they are represented by similar vectors. For example, if \basketball"
and \tennis" occur frequently in the same context, say after \play", they are
semantically related or similar according to the distributional hypothesis.
      </p>
      <p>Since co-occurrence is de ned with respect to a context, co-occurring words
can be stored into matrices whose rows represent the terms and columns
represent contexts. More speci cally, each row corresponds to a vector representation
of a word. The strength of the semantic association between words can be
computed by using cosine similarity.</p>
      <p>
        A weak point of distributional approaches is that they are able to encode
only one de nition of context at a time. The type of semantics represented in a
WordSpace depends on the context. If we choose documents as context we obtain
a semantics di erent from the one we would obtain by selecting sentences as
context. Several approaches have investigated the aforementioned problem: [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] use
a representation based on third-order tensors and provide a general framework
for distributional semantics in which it is possible to represent several aspects
of meaning using a single data structure. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] adopt vector permutations as a
means to encode order in WordSpace, as described in Section 2. BEAGLE [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
is a very well-known method to encode word order and context information in
WordSpace. The drawback of BEAGLE is that it relies on a complex model to
build vectors which is computational expensive. This problem is solved by [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] in
which the authors propose an approach similar to BEAGLE, but using a method
based on Circular Holographic Reduced Representations to compute vectors.
      </p>
      <p>
        All these methods tackle the problem of representing word order in WordSpace,
but they do not take into account syntactic context. A valuable attempt in this
direction is described in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In this work, the authors propose a method to
build WordSpace using information about syntactic dependencies. In particular,
they consider syntactic dependencies as context and assign di erent weights to
each kind of dependency. Moreover, they take into account the distance between
two words into the graph of dependencies. The results obtained by the authors
support our hypothesis that syntactic information can be useful to produce
effective WordSpace. Nonetheless, their methods are not able to directly encode
syntactic dependencies into the space.
      </p>
      <p>This work aims to provide a simple approach to encode syntactic relations
dependencies directly into the WordSpace, dealing with both the scalability
problem and the possibility to encode several context information. To achieve that
goal, we developed a strategy based on Random Indexing and vector
permutations. Moreover, this strategy opens new possibilities in the area of semantic
composition as a result of the inherent capability of encoding relations between
words.</p>
      <p>The paper is structured as follows. Section 2 describes Random Indexing,
the strategy for building our WordSpace, while details about the method used
to encode syntactic dependencies are reported in Section 3. Section 4 describes
a rst attempt to de ne a model for semantic composition which relies on our
WordSpace. Finally, the results of the evaluation performed using the GEMS
2011 Shared Evaluation data1 is presented in Section 5, while conclusions are
reported in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Random Indexing</title>
      <p>
        We exploit Random Indexing (RI), introduced by Kanerva [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], for creating a
WordSpace. This technique allows us to build a WordSpace with no need for
(either term-document or term-term) matrix factorization, because vectors are
inferred by using an incremental strategy. Moreover, it allows to solve e ciently
the problem of reducing dimensions, which is one of the key features used to
uncover the \latent semantic dimensions" of a word distribution.
      </p>
      <p>RI is based on the concept of Random Projection according to which high
dimensional vectors chosen randomly are \nearly orthogonal".</p>
      <p>Formally, given an n m matrix A and an m k matrix R made up of k
m-dimensional random vectors, we de ne a new n k matrix B as follows:
Bn;k = An;m Rm;k
k &lt;&lt; m
(1)</p>
      <p>
        The new matrix B has the property to preserve the distance between points.
This property is known as Johnson-Lindenstrauss lemma: if the distance between
two any points of A is d, then the distance dr between the corresponding points
in B will satisfy the property that dr = c d. A proof of that property is reported
in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>Speci cally, RI creates a WordSpace in two steps (in this case we consider
the document as context):
1. a context vector is assigned to each document. This vector is sparse,
highdimensional and ternary, which means that its elements can take values in
f-1, 0, 1g. A context vector contains a small number of randomly distributed
non-zero elements, and the structure of this vector follows the hypothesis
behind the concept of Random Projection;
2. context vectors are accumulated by analyzing terms and documents in which
terms occur. In particular, the semantic vector for a term is computed as the
sum of the context vectors for the documents which contain that term.
Context vectors are multiplied by term occurrences or other weighting functions,
for example log-entropy.</p>
      <p>Formally, given a collection of documents D whose vocabulary of terms is V
(we denote with dim(D) and dim(V ) the dimension of D and V , respectively)
the above steps can be formalized as follows:</p>
      <sec id="sec-2-1">
        <title>1 Available on line:</title>
        <p>http://sites.google.com/site/geometricalmodels/shared-evaluation
1. 8di 2 D, i = 0; ::; dim(D) we built the correspondent randomly generated
context vector as:
where n dim(D), ri 2 f 1; 0; 1g and !rj contains only a small number of
elements di erent from zero;
2. the WordSpace is made up of all term vectors !tj where:
(2)
(3)
(4)
and wj is the weight assigned to tj in di.</p>
        <p>By considering a xed window W of terms as context, the WordSpace is built
as follows:
1. a context vector is assigned to each term;
2. context vectors are accumulated by analyzing terms which co-occur in a
window W . In particular, the semantic vector for each term is computed as
the sum of the context vectors for terms which co-occur in W .</p>
        <p>It is important to point out that the classical RI approach can handle only
one context at a time, such as the whole document or the window W .</p>
        <p>
          A method to add information about context (word order) in RI is proposed in
[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. The authors describe a strategy to encode word order in RI by permutation
of coordinates in random vector. When the coordinates are shu ed using a
random permutation, the resulting vector is nearly orthogonal to the original one.
That operation corresponds to the generation of a new random vector. Moreover,
by applying a predetermined mechanism to obtain random permutations, such as
elements rotation, it is always possible to reconstruct the original vector using
the reverse permutations. By exploiting this strategy it is possible to obtain
di erent random vectors for each context2 in which the term occurs. Let us
consider the following example \The cat eats the mouse". To encode the word
order for the word \cat" using a context window W = 3, we obtain:
&lt; cat &gt;= (
1the) + (
        </p>
        <p>+1eat)+
+(
+2the) + (</p>
        <p>+3mouse)
!rj = (ri1; :::; rin)
!tj = wj X !ri
tdji22dDi
where nx indicates a rotation by n places of the elements in the vector x.
Indeed, the rotation is performed by n right-shifting steps.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Encoding syntactic dependencies</title>
      <p>Our idea is to encode syntactic dependencies, instead of words order, in the
WordSpace using vector permutations.
2 In the case in point the context corresponds to the word order
A syntactic dependency between two words is de ned as:
dep(head; dependent)
(5)
(6)
(7)
(8)
(9)
(10)
where dep is the syntactic link which connects the dependent word to the head
word. Generally speaking, dependent is the modi er, object or complement,
while head plays a key role in determining the behavior of the link. For example,
subj(eat; cat) means that \cat" is the subject of \eat". In that case the head
word is \eat", which plays the role of verb.</p>
      <p>The key idea is to assign a permutation function to each kind of syntactic
dependencies. Formally, let D be the set of all dependencies that we take into
account. The function f : D ! returns a schema of vector permutation for
each dep 2 D. Then, the method adopted to construct a semantic space that
takes into account both syntactic dependencies and Random Indexing can be
de ned as follows:
1. a random context vector is assigned to each term, as described in Section 2
(Random Indexing);
2. random context vectors are accumulated by analyzing terms which are linked
by a dependency. In particular the semantic vector for each term ti is
computed as the sum of the permuted context vectors for the terms tj which
are dependents of ti and the inverse-permuted vectors for the terms tj which
are heads of ti. The permutation is computed according to f . If f (d) = n
the inverse-permutation is de ned as f 1(d) = n: the elements rotation
is performed by n left-shifting steps.</p>
      <p>
        Adding permuted vectors to the head word and inverse-permuted vectors to the
corresponding dependent words allows to encode the information about both
heads and dependents into the space. This approach is similar to the one
investigated by [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for encoding relations between medical terms.
      </p>
      <p>To clarify, we provide an example. Given the following de nition of f :
and the sentence \The cat eats the mouse", we obtain the following dependencies:
The semantic vector for each word is computed as:
f (subj) =
+3
f (obj) =</p>
      <p>+7
det(the; cat)</p>
      <p>subj(eat; cat)
obj(eat; mouse)</p>
      <p>det(the; mouse)
&lt; eat &gt;= ( +3cat) + ( +7mouse)
&lt; cat &gt;= (</p>
      <p>3eat)
&lt; mouse &gt;= (
7eat)
{ eat:
{ cat:
{ mouse:</p>
      <p>In the above examples, the function f does not consider the dependency det.</p>
    </sec>
    <sec id="sec-4">
      <title>Compositional semantics</title>
      <p>In this section we provide some initial ideas about semantic composition relying
on our WordSpace. Distributional approaches represent words in isolation and
they are typically used to compute similarities between words. They are not able
to represent complex structures such as phrases or sentences. In some
applications, such as Question Answering and Text Entailment, representing text by
single words is not enough. These applications would bene t from the
composition of words in more complex structures. The strength of our approach lies on
the capability of codify syntactic relations between words overcoming the \word
isolation" issue.</p>
      <p>
        Recent work in compositional semantics argue that tensor product ( ) could
be useful to combine word vectors. In [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] some preliminary investigations about
product and tensor product are provided, while an interesting work by Clark
and Pulman [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] proposes an approach to combine symbolic and distributional
models. The main idea is to use tensor product to combine these two aspects,
but the authors do not describe a method to represent symbolic features, such
as syntactic dependencies. Conversely, our approach deals with symbolic
features by encoding syntactic information directly into the distributional model.
The authors in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] propose a strategy to represent a sentence like \man reads
magazine" by tensor product:
man
subj
read
obj
magazine
(11)
      </p>
      <p>They also propose a solid model for compositionality, but they do not provide
a strategy to represent symbolic relations, such as subj and obj. Indeed, they
state: \How to obtain vectors for the dependency relations - subj, obj, etc.
is an open question". We believe that our approach can tackle this problem by
encoding the dependency directly in the space, because each semantic vector in
our space contains information about syntactic roles.</p>
      <p>The representation based on tensor product is useful to compute sentence
similarity. For example, given the previous sentence and the following one:
\woman browses newspaper", we want to compute the similarity between those two
sentences. The sentence \woman browses newspaper", using the compositional
model, is represented by:
woman
subj
browse
obj
newspaper
(12)</p>
      <p>Finally, we can compute the similarity between the two sentences by inner
product, as follows:
(man subj read obj magazine) (woman subj browse obj newspaper)
(13)</p>
      <p>Computing the similarity requires to calculate the tensor product between
each sentence element and then compute the inner product. This task is complex,
but exploiting the following property of the tensor product:
(w1
the similarity between two sentences can be computed by taking into account
the pairs in each dependency and multiplying the inner products as follows:
man woman</p>
      <p>read browse
magazine newspaper
(15)</p>
      <p>
        According to the property above mentioned, we can compute the
similarity between sentences without using the tensor product. However, some open
questions arise. This simple compositional strategy allows to compare sentences
which have similar dependency trees. For example, the sentence \the dog bit
the man" cannot can be compared to \the man was bitten by the dog". This
problem can be easily solved by identifying active and passive forms of a verb.
When two sentences have di erent trees, Clark and Pulman [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] propose to adopt
the convolution kernel [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This strategy identi es all the possible ways of
decomposing the two trees, and sums up the similarities between all the pairwise
decompositions. It is important to point out that, in a more recent work, Clark
et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] propose a model based on [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] combined with a compositional theory
for grammatical types, known as Lambek's pregroup semantics, which is able
to take into account grammar structures. However, this strategy does not allow
to encode grammatical roles into the WordSpace. This peculiarity makes our
approach di erent. A more recent approach to distributional semantics and tree
kernel can be found in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] where authors propose a tree kernel that exploits
distributional features to compute similarity between words.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>
        The goal of the evaluation is to prove the capability of our approach in
compositional semantics task exploiting the dataset proposed by Mitchell and Lapata
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which is part of the \GEMS 2011 Shared Evaluation". The dataset is a list of
two pairs of adjective-noun/verb-object combinations or compound nouns.
Humans rated pairs of combinations according to similarity. The dataset contains
5,833 rates which range from 1 to 7. Examples of pairs follow:
support offer help provide 7
old person right hand 1
where the similarity between o er-support and provide-help (verb-object) is
higher than the one between old-person and right-hand (adjective-noun). As
suggested by the authors, the goal of the evaluation is to compare the system
performance against humans scores by Spearman correlation.
5.1
      </p>
      <sec id="sec-5-1">
        <title>System setup</title>
        <p>
          The system is implemented in Java and relies on some portions of code publicly
available in the Semantic Vectors package [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. For the evaluation of the system,
we build our WordSpaces using the WaCkypedia EN corpus3.
3 Available on line: http://wacky.sslmit.unibo.it/doku.php?id=corpora
        </p>
        <p>
          WaCkypedia EN is based on a 2009 dump of the English Wikipedia (about
800 million tokens) and includes information about: PoS, lemma and a full
dependency parse performed by MaltParser [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>Our approach involves some parameters. We set the random vector
dimension to 4,000 and the number of non-zero elements in the random vector equal
to 10. We restrict the WordSpace to the 500,000 most frequent words. Another
parameter is the set of dependencies that we take into account. In this
preliminary investigation we consider the four dependencies described in Table 1 which
reports also the kind of permutation4 applied to each dependency.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Results</title>
        <p>
          In this section, we provide the results of semantic composition. Table 2 reports
the Spearman correlation between the output of our system and the scores given
by the humans. Table 2 shows results for each type of combination: verb-object,
adjective-noun and compound nouns. Moreover, Table 2 shows the results
obtained when two other corpora were used for building the WordSpace: ukWaC
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] and TASA.
        </p>
        <p>ukWaC contains 2 billion words and is constructed from the Web by limiting
the crawling to the .uk domain and using medium-frequency words from the
BNC corpus as seeds. We use only a portion of ukWaC corpus consisting of
7,025,587 sentences (about 220,000 documents).</p>
        <p>
          The TASA corpus contains a collection of English texts that is approximately
equivalent to what an average college-level student has read in his/her lifetime.
More details about results on ukWaC and TASA corpora are reported in an our
previous work [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          It is important to underline that syntactic dependencies in ukWaC and TASA
are extracted using MINIPAR5 [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] instead of the MaltParser adopted by
WaCkypedia EN.
        </p>
        <p>The results show that WaCkypedia EN provides a signi cant improvement
with respect to TASA and ukWaC. This result is mainly due to two factors: (1)
the WordSpace built using WaCkypedia EN contains more words and
dependencies; (2) MaltParser produces more accurate dependencies than MINIPAR.
However, considering adjective-noun relation, TASA corpus obtains the best
re</p>
        <sec id="sec-5-2-1">
          <title>4 The number of rotations is randomly chosen. 5 MINIPAR is available at http://webdocs.cs.ualberta.ca/ lindek/minipar.htm</title>
          <p>Corpus Combination</p>
          <p>verb-object 0.257
WaCkypedia EN caodmjepctoiuven-dnonuonuns 00..235446
overall 0.299
verb-object 0.160
TASA caodmjepctoiuven-dnonuonuns 00..423453
overall 0.186
verb-object 0.190
ukWaC caodmjepctoiuven-dnonuonuns 00..310539</p>
          <p>overall 0.179
sult and generally all corpora obtain their best performance in this relation.
Probably, it is easier to discriminate this kind of relation than others.</p>
          <p>Another important point, is that TASA corpus provides better results than
ukWaC in spite of the huger number of relations encoded in ukWaC. We believe
that texts in ukWaC contain more noise because they are extracted from the
Web.</p>
          <p>
            As future research, we plan to conduct an experiment similar to the one
proposed in [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ], which is based on the same dataset used in our evaluation. The idea
is to use the composition functions proposed by the authors in our WordSpace,
and compare them with our compositional model. In order to perform a fair
evaluation, our WordSpace should be built from the BNC corpus. Nevertheless, the
obtained results seem to be encouraging and the strength of our approach relies
on the capability of capturing syntactic relations in a semantic space. We
believe that the real advantage of our approach, that is the possibility to represent
several syntactic relations, leaves some room for exploration.
6
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this work, we propose an approach to encode syntactic dependencies in
WordSpace using vector permutations and Random Indexing. WordSpace is
built relying on WaCkypedia EN corpus extracted from English Wikipedia pages
which contains information about syntactic dependencies. Moreover, we propose
an early attempt to use that space for semantic composition of short phrases.</p>
      <p>The evaluation using the GEMS 2011 shared dataset provides encouraging
results, but we believe that there are open points which deserve more
investigation. In future work, we have planned a deeper evaluation of our WordSpace
and a more formal study about semantic composition.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferraresi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanchetta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The WaCky Wide Web: A collection of very large linguistically processed Web-crawled corpora</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>43</volume>
          (
          <issue>3</issue>
          ),
          <volume>209</volume>
          {
          <fpage>226</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lenci</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Distributional memory: A general framework for corpusbased semantics</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>36</volume>
          (
          <issue>4</issue>
          ),
          <volume>673</volume>
          {
          <fpage>721</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semeraro</surname>
          </string-name>
          , G.:
          <article-title>Encoding syntactic dependencies by vector permutation</article-title>
          .
          <source>In: Proceedings of the GEMS 2011 Workshop on GEometrical Models of Natural Language Semantics</source>
          . pp.
          <volume>43</volume>
          {
          <fpage>51</fpage>
          . Association for Computational Linguistics, Edinburgh,
          <string-name>
            <surname>UK</surname>
          </string-name>
          (
          <year>July 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coecke</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sadrzadeh</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A compositional distributional model of meaning</article-title>
          .
          <source>In: Proceedings of the Second Quantum Interaction Symposium (QI2008)</source>
          . pp.
          <volume>133</volume>
          {
          <issue>140</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pulman</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Combining symbolic and distributional models of meaning</article-title>
          .
          <source>In: Proceedings of the AAAI Spring Symposium on Quantum Interaction</source>
          . pp.
          <volume>52</volume>
          {
          <issue>55</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Widdows</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schvaneveldt</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Rind esch, T.:
          <article-title>Logical leaps and quantum connectives: Forging paths through predication space</article-title>
          .
          <source>In: AAAI-Fall 2010 Symposium on Quantum Informatics for Cognitive, Social, and Semantic Processes</source>
          . pp.
          <volume>11</volume>
          {
          <issue>13</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
          </string-name>
          , R.:
          <article-title>Structured lexical similarity via convolution kernels on dependency trees</article-title>
          .
          <source>In: Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>1034</volume>
          {
          <fpage>1046</fpage>
          . Association for Computational Linguistics, Edinburgh, Scotland, UK. (
          <year>July 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dasgupta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An elementary proof of the Johnson-Lindenstrauss lemma</article-title>
          .
          <source>Tech. rep.</source>
          ,
          <source>Technical Report TR-99-006</source>
          , International Computer Science Institute, Berkeley, California, USA (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>De Vine</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bruza</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Semantic Oscillations:
          <article-title>Encoding Context and Structure in Complex Valued Holographic Vectors. Quantum Informatics for Cognitive, Social, and Semantic Processes (QI</article-title>
          <year>2010</year>
          )
          <article-title>(</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <source>Mathematical Structures of Language</source>
          . New York: Interscience (
          <year>1968</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Haussler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Convolution kernels on discrete structures</article-title>
          .
          <source>Technical Report UCSCCRL-99-10</source>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mewhort</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Representing word meaning and order information in a composite holographic lexicon</article-title>
          .
          <source>Psychological review 114(1)</source>
          ,
          <volume>1</volume>
          {
          <fpage>37</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kanerva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Sparse Distributed Memory</article-title>
          . MIT Press (
          <year>1988</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Dependency-based evaluation of MINIPAR. Treebanks: building and using parsed corpora (</article-title>
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Mitchell, J.,
          <string-name>
            <surname>Lapata</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Composition in distributional models of semantics</article-title>
          .
          <source>Cognitive Science</source>
          <volume>34</volume>
          (
          <issue>8</issue>
          ),
          <volume>1388</volume>
          {
          <fpage>1429</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Nivre</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nilsson</surname>
          </string-name>
          , J.:
          <article-title>Maltparser: A data-driven parser-generator for dependency parsing</article-title>
          .
          <source>In: Proceedings of LREC</source>
          . vol.
          <volume>6</volume>
          , pp.
          <volume>2216</volume>
          {
          <issue>2219</issue>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Pado</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Dependency-based construction of semantic space models</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>33</volume>
          (
          <issue>2</issue>
          ),
          <volume>161</volume>
          {
          <fpage>199</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Sahlgren</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The Word-Space Model: Using distributional analysis to represent syntagmatic and paradigmatic relations between words in high-dimensional vector spaces</article-title>
          .
          <source>Ph.D. thesis</source>
          , Stockholm: Stockholm University, Faculty of Humanities, Department of Linguistics (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Sahlgren</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holst</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanerva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Permutations as a means to encode order in word space</article-title>
          .
          <source>In: Proceedings of the 30th Annual Meeting of the Cognitive Science Society (CogSci'08)</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. Schutze, H.:
          <article-title>Word space</article-title>
          . In: Hanson,
          <string-name>
            <given-names>S.J.</given-names>
            ,
            <surname>Cowan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.D.</given-names>
            ,
            <surname>Giles</surname>
          </string-name>
          , C.L. (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          . pp.
          <volume>895</volume>
          {
          <fpage>902</fpage>
          . Morgan Kaufmann Publishers (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Widdows</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Semantic vector products: Some initial investigations</article-title>
          .
          <source>In: The Second AAAI Symposium on Quantum Interaction</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Widdows</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferraro</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Semantic Vectors: A Scalable Open Source Package and Online Technology Management Application</article-title>
          .
          <source>In: Proceedings of the 6th International Conference on Language Resources and Evaluation (LREC</source>
          <year>2008</year>
          )
          <article-title>(</article-title>
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>