<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MEDEA: Merging Event knowledge and Distributional vEctor Addition</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ludovica Pannitto</string-name>
          <email>ellepannitto@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Lenci</string-name>
          <email>alessandro.lenci@unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CoLing Lab, University of Pisa</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. The great majority of compositional models in distributional semantics present methods to compose distributional vectors or tensors in a representation of the sentence. Here we propose to enrich the best performing method (vector addition, which we take as a baseline) with distributional knowledge about events, outperforming our baseline.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Compositional Distributional</title>
    </sec>
    <sec id="sec-2">
      <title>Semantics: Beyond vector addition</title>
      <p>
        Composing word representations into larger
phrases and sentences notoriously represents a
big challenge for distributional semantics
        <xref ref-type="bibr" rid="ref13">(Lenci,
2018)</xref>
        . Various approaches have been proposed
ranging from simple arithmetic operations on
word vectors
        <xref ref-type="bibr" rid="ref18 ref8">(Mitchell and Lapata, 2008)</xref>
        , to
algebraic compositional functions on higher-order
objects
        <xref ref-type="bibr" rid="ref2 ref6">(Baroni et al., 2014; Coecke et al., 2010)</xref>
        ,
as well as neural networks approaches
        <xref ref-type="bibr" rid="ref17 ref20">(Socher et
al., 2010; Mikolov et al., 2013)</xref>
        .
      </p>
      <p>
        Among all proposed compositional functions,
vector addition still shows the best performances
on various tasks
        <xref ref-type="bibr" rid="ref1 ref19 ref3">(Asher et al., 2016; Blacoe and
Lapata, 2012; Rimell et al., 2016)</xref>
        , beating more
complex methods, such as the Lexical Functional
Model
        <xref ref-type="bibr" rid="ref2">(Baroni et al., 2014)</xref>
        . However, the success
of vector addition is quite puzzling from the
linguistic and cognitive point of view: the meaning
of a complex expression is not simply the sum of
the meaning of its parts, and the contribution of
a lexical item might be different depending on its
syntactic as well as pragmatic context.
      </p>
      <p>The majority of available models in literature
assumes the meaning of complex expressions like
sentences to be a vector (i.e., an embedding)
projected from the vectors representing the content
of its lexical parts. However, as pointed out by
Erk and Pado´ (2008), while vectors serve well the
cause of capturing the semantic relatedness among
lexemes, this might not be the best choice for
more complex linguistic expressions, because of
the limited and fixed amount of information that
can be encoded. Moreover events and situations,
expressed through sentences, are by definition
inherently complex and structured semantic objects.
Actually, assuming the equation “meaning is
vector” is eventually too limited even at the lexical
level.</p>
      <p>
        Psycholinguistic evidence shows that lexical
items activate a great amount of generalized event
knowledge (GEK)
        <xref ref-type="bibr" rid="ref10 ref7 ref9">(Elman, 2011; Hagoort and
van Berkum, 2007; Hare et al., 2009)</xref>
        , and that this
knowledge is crucially exploited during online
language processing, constraining the speakers’
expectations about upcoming linguistic input
        <xref ref-type="bibr" rid="ref16">(McRae and Matsuki, 2009)</xref>
        . GEK is concerned
with the idea that the lexicon is not organized as
a dictionary, but rather as a network, where words
trigger expectations about the upcoming input,
influenced by pragmatic knowledge along with
lexical knowledge. Therefore sentence
comprehension can be phrased as the identification of the
event that best explains the linguistic cues used in
the input
        <xref ref-type="bibr" rid="ref12 ref19">(Kuperberg and Jaeger, 2016)</xref>
        .
      </p>
      <p>In this paper, we introduce MEDEA, a
compositional distributional model of sentence meaning
which integrates vector addition with GEK
activated by lexical items. MEDEA is directly
inspired by the model in Chersoni et al. (2017a) and
relies on two major assumptions:
lexical items are represented with
embeddings within a network of syntagmatic
relations encoding prototypical knowledge about
events;
the semantic representation of a sentence is
a structured object incrementally
integrating the semantic information cued by lexical
items.</p>
      <p>We test MEDEA on two datasets for
compositional distributional semantics in which addition
has proven to be very hard to beat. At least, before
meeting MEDEA.</p>
    </sec>
    <sec id="sec-3">
      <title>2 Introducing MEDEA</title>
      <p>MEDEA consists of two main components: i.) a</p>
      <sec id="sec-3-1">
        <title>Distributional Event Graph (DEG) that models a</title>
        <p>fragment of semantic memory activated by lexical
units (Section 2.1); ii.) a Meaning Composition
Function that dynamically integrates information
activated from DEG to build a sentence semantic
representation (Section 2.2).
2.1</p>
      </sec>
      <sec id="sec-3-2">
        <title>Distributional Event Graph</title>
        <p>We assume a broad notion of event, corresponding
to any configuration of entities, actions,
properties, and relationships. Accordingly, an event
can be a complex relationship between entities, as
the one expressed by the sentence The student read
a book, but also the association between an
individual and a property, as expressed by the noun
phrase heavy book.</p>
        <p>In order to represent the GEK cued by
lexical items during sentence comprehension, we
explored a graph based implementation of a
distributional model, for both theoretical and
methodological reasons: in graphs, structural-syntactic
information and lexical information can naturally
coexist and be related, moreover vectorial
distributional models often struggle with the
modeling of dynamic phenomena, as it is often difficult
to update the recorded information, while graphs
are more suitable for situations where relations
among items change overtime. The data structure
would ideally keep track of each event
automatically retrieved from corpora, thus indirectly
containing information about schematic or
underspecified events, by abstracting over one or more
participants from each recorded instance. Events are
cued by all the potential participants to the event.</p>
        <p>The nodes of DEG are lexical embeddings, and
edges link lexical items participating to the same
events (i.e., its syntagmatic neighbors). Edges are
weighted with respect to the statistical salience of
the event given the item. Weights, expressed in
terms of a statistical association measure such as
Local Mutual Information, determine the event
activation strength by linguistic cues.</p>
        <p>In order to build DEG, we automatically
harvested events from corpora, using syntactic
relations as an approximation of semantic roles of
event participants. From a dependency parsed
sentence we identified an event by selecting a
semantic head (verb or noun) and grouping all its
syntactic dependents together (Figure 1). Since we
expect each participant to be able to trigger the
event and consequently any of the other
participants, a relation can be created and added to the
graph from each subset of each group extracted
from sentence.
The resulting structure is therefore a weighted
hypergraph, as it contains relations holding among
groups of nodes, and a labeled multigraph, since
each edge or hyperedge is labeled in order to
represent the syntactic pattern holding in the group.</p>
        <p>As graph nodes are embeddings, given a lexical
cue w, DEG can be queried in two modes:
retrieving the most similar nodes to w (i.e.,
its paradigmatic neighbors), using a standard
vector similarity measure like the cosine
(Table 1, top row);
retrieving the closest associates of w (i.e., its
syntagmatic neighbors), using the weights on
the graph edges (Table 1, bottom row).</p>
        <p>essay/N, anthology/N, novel/N, author/N,
para. neighbors publish/N, biography/N, autobiography/N,
nonfiction/N, story/N, novella/N
publish/V, write/V, read/V,
synt. neighbors include/V, child/N, series/N,</p>
        <p>have/V, buy/V, author/N, contain/V
In MEDEA, we model sentence comprehension
as the creation of a semantic representation SR,
which includes two different yet interacting
information tiers that are equally relevant in the
overall representation of sentence meaning: i.)
the lexical meaning component (LM), which is a
context-independent tier of sentence meaning that
accumulates the lexical content of the sentence,
as traditional models do; ii.) an active context
(AC), which aims at representing the most
probable event, in terms of its participants, that can be
reconstructed from DEG portions cued by lexical
items. This latter component corresponds to the
GEK activated by the single lexemes (or by other
contextual elements) and integrated into a
semantically coherent structure representing the sentence
interpretation. It is incrementally updated during
processing, when a new input is integrated into
existing information.</p>
      </sec>
      <sec id="sec-3-3">
        <title>2.2.1 Active Context</title>
        <p>Each lexical item in the input activates a portion of
GEK that is integrated into the current AC through
a process of mutual re-weighting that aims at
maximizing the overall semantic coherence of the SR.</p>
        <p>At the outset, no information is contained in the
AC of the sentence. When new lexeme -
syntactic role pair hwi; rii (e.g., student - nsbj) are
encountered, expectations about the set of upcoming
roles in the sentences are generated from DEG
(figure 2). These include: i.) expectations about the
role filled by the lexeme itself, which consists of
its vector (and possibly its p-neighbours); ii.)
expectations about sentence structure and other
participants, which are collected in weighted list of
vectors of its s-neighbours.</p>
        <p>These expectations are then weighted with
respect to what is already in the AC, and the AC is
similarly adapted to the ewly retrieved
information: each weighted list is represented with the
weighted centroid of its top elements, and each
element of a weighted lists is re-ranked
according to its cosine similarity with the correspondent
centroid (e.g., the newly retrieved weighted list of
subjects is ranked according to the cosine
similarity of each item in the list with the weighted
centroid of subjects available in AC).</p>
        <p>
          The final semantic representation of a sentence
consists of two vectors, the lexical meaning
vector (LM!) and the event knowledge vector (A!C),
which is obtained by composing the weighted
centroids of each role in AC.
We wanted to evaluate the contribution of
activated event knowledge in a sentence
comprehension task. For this reason, among the many
existing datasets concerning entailment or
paraphrase detection, we chose RELPRON
          <xref ref-type="bibr" rid="ref19">(Rimell et
al., 2016)</xref>
          , a dataset of subject and object
relative clauses, and the transitive sentence
similarity dataset presented in Kartsaklis and Sadrzadeh
(2014). These two datasets show an intermediate
level of grammatical complexity, as they involve
complete sentences (while other datasets include
smaller phrases), but have fixed length structures
featuring similar syntactic constructions (i.e.,
transitive sentences). The two datasets differ with
respect to size and construction method.
        </p>
        <p>RELPRON consists of 1,087 pairs, split in
development and test set, made up by a target noun
labeled with a syntactic role (either subject
or direct object) and a property expressed as
[head noun] that [verb] [argument]. For
instance, here are some example properties for
the target noun treaty:
(1) a. OBJ treaty/N: document/N that
delegation/N negotiate/V
b. SBJ treaty/N: document/N that grant/V
in</p>
        <p>dependence/N</p>
      </sec>
      <sec id="sec-3-4">
        <title>Transitive sentence similarity dataset consists</title>
        <p>of 108 pairs of transitive sentences, each
annotated with human similarity judgments
collected through the Amazon Mechanical
Turk platform. Each transitive sentence in
composed by a triplet subject verb object.
Here are two pairs with high (2) and low (3)
similarity scores respectively:
(2) a. government use power</p>
        <p>b. authority exercise influence
(3) a. team win match</p>
        <p>b. design reduce amount
3.2</p>
      </sec>
      <sec id="sec-3-5">
        <title>Graph implementation</title>
        <p>
          We tailored the construction of the DEG to this
kind of simple syntactic structures, restricting it
to the case of relations among pairs of event
participants. Relations were automatically
extracted from a 2018 dump of Wikipedia, BNC,
and ukWaC corpora, parsed with the Stanford
CoreNLP Pipeline
          <xref ref-type="bibr" rid="ref15">(Manning et al., 2014)</xref>
          .
        </p>
        <p>Each h(word1; word2); (r1; r2)i pair was then
weighted with a smoothed version of Local
Mutual Information1:</p>
        <p>LMI (w1; w2; r1; r2) = f(w1; w2; r1; r2)log(P^(wP^1()wP^1;(ww22;)rP1^;(rr21);r2))
where:
f (x)</p>
        <p>P^ (x) = Px f (x)
Each lexical node in DEG was then represented
with its embedding. We used the same training
parameters as in Rimell et al. (2016),2, since we
wanted our model to be directly comparable with
their results on the dataset. While Rimell et al.
(2016) built the vectors from a 2015 download of
Wikpedia, we needed to cover all the lexemes
contained in the graph and therefore we used the same
corpora from which the DEG was extracted.</p>
        <p>
          We represented each property in RELPRON as
a triplet ((hn; r); (w1; r1); (w2; r2)) where hn is
the head noun, w1 and w2 are the lexemes that
1The smoothed version (with = 0:75) was chosen in
order to alleviate PMI’s bias towards rare words
          <xref ref-type="bibr" rid="ref14">(Levy et al.,
2015)</xref>
          , which arises especially when extending the graph to
more complex structures than pairs.
        </p>
        <p>
          2lemmatized 100-dim vectors with skip-gram with
negative sampling (SGNS
          <xref ref-type="bibr" rid="ref17">(Mikolov et al., 2013)</xref>
          ), setting
minimum item frequency at 100 and context window size at 10.
(1)
(2)
compose the proper relative clause, and each
element of the triplet is associated with its syntactic
role in the property sentence.3 Likewise, each
sentence of the transitive sentences dataset is a triplet
((w1; nsbj); (w2; root); (w3; dobj)).
3.3
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>Active Context implementation</title>
        <p>In MEDEA, the SR is composed of two vectors:
LM!, as the sum of the word embeddings (as
this was the best performing model in
literature, on the chosen datasets);
A!C, obtained by summing up all the
weighted centroids of triggered participants.
Each lexeme - syntactic role pair is used to
retrieve its 50 top s-neighbors from the graph.
The top 20 re-ranked elements were used to
build each weighted centroid. These
threshold were choosen empirically, after a few
trials with different (i.e., higher) thresholds (as
in Chersoni et al. (2017b)).</p>
        <p>We provide an example of the re-weighting
process with the property document that store
maintains, whose target is inventory: i.) at first the head
noun document is encountered: its vector is
activated as event knowledge for the object role of
the sentence and constitutes the contextual
information in AC against which GEK is re-weighted;
ii.) store as a subject triggers some direct object
participants, such as product, range, item,
technology, etc. If the centroid were built from the top of
this list, the cosine similarity with the target would
be around 0:62; iii.) s-neighbours of store are
reweighted according to the fact that AC contains
some information about the target already, (i.e.,
the fact that it is a document). The re-weighting
process has the effect of placing on top of the list
elements that are more similar to document. Thus,
now we find collection, copy, book, item, name,
trading, location, etc., improving the cosine
similarity with the target, that goes up to 0:68; iv.)
the same happens for maintain: its s-neighbors are
retrieved and weighted against the complete AC,
improving their cosine similarity with inventory,
from 0:55 to 0:61.
3.4</p>
      </sec>
      <sec id="sec-3-7">
        <title>Evaluation</title>
        <p>We evaluated our model on RELPRON
development set using Mean Average Precision (MAP), as
3The relation for the head noun is assumed to be the same
as the target relation (either subject of direct object of the
relative clause).
in Rimell et al. (2016). We produced the
compositional representation of each property in terms
of SR, and then ranked for each target all the 518
properties of the dataset portion, according to their
similarity to the target. Our main goal was to
evaluate the contribution of event knowledge,
therefore the similarity between the target vector and
the property SR was measured as the sum of the
cosine similarity of the target vector with the LM!
of the property, and the cosine similarity of the
target vector with the A!C cued by each property. As
shown in Table 2, the full MEDEA model (last
column) achieves top performance, above the simple
additive model LM.</p>
        <p>verb
arg
hn+verb
hn+arg
verb+arg
hn+verb+arg</p>
        <p>LM
0,18
0,34
0,27
0,47
0,42
0,51</p>
        <p>RELPRON</p>
        <p>AC LM+AC
0,18 0,20
0,34 0,36
0,28 0,29
0,45 0,49
0,28 0,39
0,47 0,55</p>
        <p>For the transitive sentences dataset, we
evaluated the correlation of our scores with human
ratings with Spearman’s . The similarity between
a pair of sentences s1; s2 is defined as the cosine
between their LM vectors plus the cosine between
their EK vectors. MEDEA is in the last column of
Table 3 and again outperforms simple addition.
sbj
root
obj
sbj+root
sbj+obj
root+obj
sbj+root+obj</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>We provided a basic implementation of a
meaning composition model, which aims at being
incremental and cognitively plausible. While still
relying on vector addition, our results suggest that
distributional vectors do not encode sufficient
information about event knowledge, and that, in line
with psycholinguistic results, activated GEK plays
an important role in building semantic
representations during online sentence processing.</p>
      <p>Our ongoing work focuses on refining the way
in which this event knowledge takes part in the
processing phase and testing its performance on
more complex datasets: while both RELPRON and
the transitive sentences dataset provided a straight
forward mapping between syntactic label and
semantic roles, more naturalistic datasets show a
much wider range of syntactic phenomena that
would allow us to test how expectations jointly
work on syntactic structure and semantic roles.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Nicholas</given-names>
            <surname>Asher</surname>
          </string-name>
          , Tim Van de Cruys, Antoine Bride, and Ma´rta Abrusa´n.
          <year>2016</year>
          .
          <article-title>Integrating Type Theory and Distributional Semantics: A Case Study on Adjective-Noun Compositions</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>42</volume>
          (
          <issue>4</issue>
          ):
          <fpage>703</fpage>
          -
          <lpage>725</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          , Raffaela Bernardi, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Zamparelli</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Frege in Space: A Program of Compositional Distributional Semantics</article-title>
          .
          <source>Linguistic Issues in Language Technology</source>
          ,
          <volume>9</volume>
          (
          <issue>6</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>William</given-names>
            <surname>Blacoe</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mirella</given-names>
            <surname>Lapata</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A comparison of vector-based representations for semantic composition</article-title>
          .
          <source>In Proceedings of the 2012 joint conference on empirical methods in natural language processing and computational natural language learning</source>
          , pages
          <fpage>546</fpage>
          -
          <lpage>556</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Emmanuele</given-names>
            <surname>Chersoni</surname>
          </string-name>
          , Alessandro Lenci, and
          <string-name>
            <given-names>Philippe</given-names>
            <surname>Blache</surname>
          </string-name>
          . 2017a.
          <article-title>Logical metonymy in a distributional model of sentence comprehension</article-title>
          .
          <source>In Sixth Joint Conference on Lexical and Computational Semantics (* SEM</source>
          <year>2017</year>
          ), pages
          <fpage>168</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Emmanuele</given-names>
            <surname>Chersoni</surname>
          </string-name>
          , Enrico Santus, Philippe Blache, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          . 2017b.
          <article-title>Is structure necessary for modeling argument expectations in distributional semantics?</article-title>
          <source>In 12th International Conference on Computational Semantics (IWCS</source>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Bob</given-names>
            <surname>Coecke</surname>
          </string-name>
          , Stephen Clark, and
          <string-name>
            <given-names>Mehrnoosh</given-names>
            <surname>Sadrzadeh</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Mathematical foundations for a compositional distributional model of meaning</article-title>
          .
          <source>Technical report.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey L Elman</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Lexical knowledge without a lexicon? The mental lexicon, 6(1</article-title>
          ):
          <fpage>1</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Katrin</given-names>
            <surname>Erk</surname>
          </string-name>
          and Sebastian Pado´.
          <year>2008</year>
          .
          <article-title>A structured vector space model for word meaning in context</article-title>
          .
          <source>In Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>897</fpage>
          -
          <lpage>906</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Peter</given-names>
            <surname>Hagoort</surname>
          </string-name>
          and Jos van Berkum.
          <year>2007</year>
          .
          <article-title>Beyond the sentence given</article-title>
          .
          <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source>
          ,
          <volume>362</volume>
          (
          <issue>1481</issue>
          ):
          <fpage>801</fpage>
          -
          <lpage>811</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Mary</given-names>
            <surname>Hare</surname>
          </string-name>
          , Michael Jones, Caroline Thomson, Sarah Kelly, and
          <string-name>
            <surname>Ken McRae</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Activating event knowledge</article-title>
          .
          <source>Cognition</source>
          ,
          <volume>111</volume>
          (
          <issue>2</issue>
          ):
          <fpage>151</fpage>
          -
          <lpage>167</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Dimitri</given-names>
            <surname>Kartsaklis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mehrnoosh</given-names>
            <surname>Sadrzadeh</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>A study of entanglement in a categorical framework of natural language</article-title>
          .
          <source>In Proceedings of the 11th Workshop on Quantum Physics and Logic (QPL)</source>
          . Kyoto, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Gina R Kuperberg</surname>
            and
            <given-names>T Florian</given-names>
          </string-name>
          <string-name>
            <surname>Jaeger</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>What do we mean by prediction in language comprehension? Language, cognition</article-title>
          and neuroscience,
          <volume>31</volume>
          (
          <issue>1</issue>
          ):
          <fpage>32</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Distributional Models of Word Meaning</article-title>
          .
          <source>Annual Review of Linguistics</source>
          ,
          <volume>4</volume>
          :
          <fpage>151</fpage>
          -
          <lpage>171</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Omer</surname>
            <given-names>Levy</given-names>
          </string-name>
          , Yoav Goldberg, and
          <string-name>
            <given-names>Ido</given-names>
            <surname>Dagan</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Improving distributional similarity with lessons learned from word embeddings</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>3</volume>
          :
          <fpage>211</fpage>
          -
          <lpage>225</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Christopher D. Manning</surname>
            , Mihai Surdeanu, John Bauer, Jenny Finkel,
            <given-names>Steven J.</given-names>
          </string-name>
          <string-name>
            <surname>Bethard</surname>
          </string-name>
          , and
          <string-name>
            <surname>David McClosky</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The Stanford CoreNLP natural language processing toolkit. In Association for Computational Linguistics (ACL) System Demonstrations</article-title>
          , pages
          <fpage>55</fpage>
          -
          <lpage>60</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Ken</given-names>
            <surname>McRae</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kazunaga</given-names>
            <surname>Matsuki</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>People use their knowledge of common events to understand language, and do so as quickly as possible</article-title>
          .
          <source>Language and linguistics compass</source>
          ,
          <volume>3</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1417</fpage>
          -
          <lpage>1429</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Mitchell</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mirella</given-names>
            <surname>Lapata</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Vector-based models of semantic composition</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Laura</given-names>
            <surname>Rimell</surname>
          </string-name>
          , Jean Maillard, Tamara Polajnar, and
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Clark</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Relpron: A relative clause evaluation data set for compositional distributional semantics</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>42</volume>
          (
          <issue>4</issue>
          ):
          <fpage>661</fpage>
          -
          <lpage>701</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Richard</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          , and Andrew Y Ng.
          <year>2010</year>
          .
          <article-title>Learning continuous phrase representations and syntactic parsing with recursive neural networks</article-title>
          .
          <source>In Proceedings of the NIPS-2010 Deep Learning and Unsupervised Feature Learning Workshop</source>
          , volume
          <volume>2010</volume>
          , pages
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>