<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Same same but different: Type and typicality in a distributional model of complement coercion</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandra Zarcone</string-name>
          <email>1zarcone@coli.uni-saarland.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Pad o´</string-name>
          <email>2pado@ims.uni-stuttgart.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Lenci</string-name>
          <email>3alessandro.lenci@unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universita ̈t des Saarlandes Universita ̈t Stuttgart Universita` degli Studi di Pisa Saarbr u ̈cken</institution>
          ,
          <addr-line>Germany Stuttgart, Germany Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>91</fpage>
      <lpage>94</lpage>
      <abstract>
        <p>We aim to model the results from a selfpaced reading experiment, which tested the effect of semantic type clash and typicality on the processing of German complement coercion. We present two distributional semantic models to test if they can model the effect of both type and typicality in the psycholinguistic study. We show that one of the models, without explicitly representing type information, can account both for the effect of type and typicality in complement coercion.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction: Complement Coercion</title>
      <p>
        Complement coercion (The author began the book
→ reading the book) has been shown to cause an
increase in processing cost
        <xref ref-type="bibr" rid="ref4 ref7">(Pylkka¨nen and
McElree, 2006; Katsika et al., 2012)</xref>
        , which has been
ascribed to a type clash between an event-selecting
verb (begin) and an entity-denoting object (book).
The increase in processing costs is found in
comparison with a baseline condition, where the same
verb is combined with an event-denoting object
(journey), which does not trigger a type clash.
      </p>
      <p>
        A second influence on processing cost is the
thematic fit or typicality of the fillers of the verb’s
argument slots
        <xref ref-type="bibr" rid="ref2 ref6">(Bicknell et al., 2010; Matsuki et
al., 2011)</xref>
        : high-typicality combinations are
processed more quickly than low-typicality ones (the
mechanic checked the brakes / the spelling).
      </p>
      <p>
        Distributional semantic models (DSMs) can
successfully model a range of psycholinguistic
phenomena, including the effect of typicality
on complement coercion
        <xref ref-type="bibr" rid="ref11">(Zarcone et al., 2012)</xref>
        .
However, they generally do not include a notion
of type. Can a DSM account for effects both of
type and typicality?
      </p>
      <p>In this paper, we consider experimental results
from a study on complement coercion in German
that manipulates both type and typicality. We
discuss the performance of existing DSMs and a
novel DSM combination. We also discuss how
type information can be emerge from
distributional information.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Manipulating Type and Typicality</title>
      <p>In a self-paced reading study on German
complement coercion (Zarcone et al., in preparation), we
have manipulated both type and typicality. The
dataset consists of 20 pairs of subjets (S) and
aspectual verbs (V). Each pair is combined with four
nominal objects (O) in SOV order:
[S Das Geburtstagskind] hat [O mit den Geschenken
[S The birthday boy] has [O with the presents
/ der Feier / der Suppe / der Schicht] [V angefangen].
/ party / soup / work shift] [V begun].</p>
      <p>The objects are: a high-typicality entity
(presents); a high-typicality event (party); a
lowtypicality entity (soup); and a low-typicality event
(work shift). The low-typicality objects are drawn
from the high-typicality objects of other S-V pairs.</p>
      <p>The self-paced reading study yielded the
following significant effects: (1) an effect of
typicality on reading times (t = 2.28, p = .02) at the
object region (indicating subject-object
integration), (2) an effect of object type on reading times
(t = −2.5, p = .01) at the verb region (the region
of the type clash), (3) an interaction of type and
thematic fit at the verb region (t = 2.04, p = .04).
Mean reading times per condition are reported in
Table 1. In sum, the study shows that
complement coercion involves both type and typicality.
Thus, computational models of complement
coercion need to account for both.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Modeling the Experimental Results</title>
      <p>Distributional semantic models (DSMs)
represent word meaning as high-dimensional vectors
recording co-occurrences with elements of their
high-fit entity
high-fit event
low-fit entity
low-fit event</p>
      <p>Object region
mit den Geschenken
with the presents
642
655
667
710
usage contexts. Semantic similarity is defined in
terms of a vector similarity metric such as cosine.</p>
      <p>
        Distributional Memory (DM, Baroni and Lenci
(2010)) is a DSM that includes syntactic
knowledge into the word representations. More
concretely, the TypeDM version of DM records
wordrelation-word tuples hw1 r w2i. The tuples are
weighted by Local Mutual Information
        <xref ref-type="bibr" rid="ref3">(Evert,
2005)</xref>
        , which can be employed to model
predicateargument typicality. For example, the weight of
hbook obj readi is higher than hlabel obj readi,
which in turn is higher than helephant obj readi.
TypeDM has been shown to be versatile and
effective in several semantic tasks, including predicting
verb-argument plausibility.
3.1
      </p>
      <p>Complement Coercion and DSMs.</p>
      <p>DM has been extended into the Expectation
Composition and Update model (ECU, Lenci (2011)),
a family of procedures that can be used to
predict the typicality of one sentence part given other
sentence parts. E.g., to model the typicality at
the verb region in a German sentence with SOV
word order (e.g. Das Geburtstagskind hat mit dem
Geschenk angefangen / The birthday boy has with
the present begun), ECU determines the thematic
fit for the verb given subject and object:
• compute an expectation for the verb given the
subject s, as the distribution over verbs v
defined by the weights of the tupleshs subj vi
• compute an expectation for the verb given the
object o, as the distribution over verbs v
defined by the weights of the tuplesho obj vi.
To combine the subject and object expectations,
we combine the two distributions component by
component, typically either by sum or products.
This distribution is then represented in a vector
space by computing the centroid or prototype of
the vectors of the 20 most expected verbs. Finally,
the thematic fit for a verbv given the subject s and
the object o is its cosine similarity to the centroid.
Subject</p>
      <p>Object</p>
      <p>
        Verb
JE
ECU. We call the models following the ECU
procedure SOV+ and SOV*, depending on their
combination operation (sum and product,
respectively). Simpler models only consider the
influence of subject or object on the verb (SV and OV
respectively), just by leaving out the combination
step. These models can successfully account for
reading time results on a dataset of complement
coercion in German that manipulates typicality but
not type
        <xref ref-type="bibr" rid="ref11">(Zarcone et al., 2012)</xref>
        .
      </p>
      <p>In order to test ECU on a dataset which
manipulates both type and typicality, we evaluate the
following ECU models on the complement
coercion data in (Zarcone et al., in preparation): SO
to model effects at the object given the subject;
SOV+, SOV* and OV to model effects at the verb.
We expect these models to account for the
typicality effect at the object (1), but not for the type
effects at the verb (2,3).</p>
      <p>The results are summarized in Table 2 (left and
middle). In accordance with our prediction, SO
correctly yields the typicality effect at the object
(F = 7.38, p &lt; 0.01). Neither SOV+, SOV*, nor
OV can model the type-typicality interaction at the
verb (3). Surprisingly, though, SOV* and OV yield
(2), an effect of type at the verb (F = 5.3228, p &lt;
0.05 and F = 20.388, p &lt; 0.001, respectively).
Joint Expectations. The reading time study
found that the subject-object typicality effects
linger at the verb, interacting with type. The main
shortcoming of ECU is its inability to model the
typicality effects at the verb. This is due to the
architecture of the SOV models (cf. Fig. 1, top): they
compute the expectations for the verb first from
the subject (SV) and update them with the object’s
expectations (OV). They ignore the interaction
between subject and object (SO) – the source of
typicality effects (1,3) – corresponding to the
assumption that this interaction should only matter at the
object. In order to account for this, we draw an
analogy to the concept of joint probability:
P (S, O, V )
(1) effect of typicality at the object region (SO interaction)
(2) effect of type at the verb region (type clash)
(3) type x thematic fit interaction at the verb region
non-compos.</p>
      <p>SO OV
X ×
× X
× ×
SOV+
×
×
×
which is equivalent (by the chain rule), to
which we can interpret distributionally as
motivation to reweight the typicality of the verb given
the object with the typicality of the object given
the subject, thus re-introducing the subject-object
interaction into the verb prediction (cf. Figure 1,
bottom).</p>
      <p>In the Joint Expectation (JE) model, the
thematic fit score assigned to the target verb is
influenced both by the verb’s thematic fit with the
object (the verb’s initial thematic fit score,
equivalent to the ECU weight for the hobject obj verbi
tuple) and by the object’s thematic fit with the
subject (equivalent to the ECU weight for the
hsubject verb objecti tuple), which in turn is
used to reweight the verb’s score.</p>
      <p>Similar to ECU, there is a choice of
combination operations in JE (sum or product). Since
JE can be formulated as a simple wrapper around
ECU, ECU can be used to compute the individual
components (e.g. SO, OV, or more complex ones)
and these then just need to be combined additively
(SO+OV) or multiplicatively (SO*OV).</p>
      <p>The right-hand side of Table 2 shows the results
for JE. SO+OV yields an effect of typicality (F =
6.777, p &lt; 0.05) but no effect of type (2) or
interaction (3). SO*OV yields two main effects of
(2) type (F = 7.2359, p &lt; 0.05) and typicality (F =
7.2359, p &lt; 0.01), although no interaction (3).</p>
      <p>Comparing the two models, we see that ECU
SO accounts for the results obtained at the object
(1), but the SOV models cannot explain the
interaction with typicality on the verb (2,3). JE (SO *
OV) models the effects of both type (2) and
typicality at the verb, but does not (yet) account for
their interaction (3).</p>
      <p>
        We found that the SO model successfully accounts
for the effect of typicality at the object. This is
not surprising: one of the most typical tasks
successfully performed by distributional models such
as ECU is predicting verb-argument plausibility,
and ECU had already been successful in modeling
effects of typicality on reading times in German
complement coercion
        <xref ref-type="bibr" rid="ref11">(Zarcone et al., 2012)</xref>
        .
      </p>
      <p>On the other hand, the ECU SOV models were
not able to account for the type–typicality
interaction at the verb. The JE model (SO * OV), which
we presented as an alternative to the ECU model
to better account for the typicality effects at the
verb, yielded effects of both type and typicality at
the verb, but did not account for their interaction.</p>
      <p>Our most surprising result is that the OV, SOV*,
and SO*OV models explain the effect of type. As
DSMs do not represent this concept explicitly, a
possible interpretation suggested by our results is
that type and typicality are not distinct categories,
but capture properties of predicate-argument
combinations at different granularity levels.</p>
      <p>
        Distributional models can account for types
because they emerge from the observed corpus
distributions. Specifically, for the aspectual verbs used
in the present data set, the distribution over their
objects – namely that event nouns occur much
more frequently that object nouns
        <xref ref-type="bibr" rid="ref12">(Zarcone et al.,
2013)</xref>
        – corresponds more naturally to an
interpretation in terms of types than of typicality. A
compositional distributional model where
semantic types emerge as patterns of behavior has the
advantage of relying on minimal assumptions
regarding the granularity of the type ontology, which
is intriguing, as pattern recognition is a key aspect
of human cognition
        <xref ref-type="bibr" rid="ref10 ref8 ref9">(Rumelhart and McClelland,
1987; Saffran et al., 1996; Tomasello, 2009)</xref>
        .
      </p>
      <p>In conclusion, the picture that emerges from
our experiments is one where (1) expectations for
predicate-argument combinations have a
hierarchical structure, with types as a high-level
distinction and typicality as a low-level distinction, (2)
both levels are different, but interact early during
processing, influencing reading times, and (3) both
type and typicality can emerge from the “same
same” distributional model.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This research was funded by the German Research
Foundation (DFG) as part of SFB 732
”Incremental Specification in Context” and SFB 1102
”Information Density and Linguistic Encoding”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Distributional Memory: a general framework for corpus-based semantics</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>36</volume>
          (
          <issue>4</issue>
          ):
          <fpage>673</fpage>
          -
          <lpage>721</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Klinton</given-names>
            <surname>Bicknell</surname>
          </string-name>
          , Jeffrey L Elman, Mary Hare,
          <string-name>
            <surname>Ken McRae</surname>
            ,
            <given-names>and Marta</given-names>
          </string-name>
          <string-name>
            <surname>Kutas</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Effects of event knowledge in processing verbal arguments</article-title>
          .
          <source>Journal of Memory and Language</source>
          ,
          <volume>63</volume>
          :
          <fpage>489</fpage>
          -
          <lpage>505</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Evert</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>The statistics of word cooccurrences</article-title>
          .
          <source>Ph.D. thesis</source>
          , Universita¨t Stuttgart.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Argyro</given-names>
            <surname>Katsika</surname>
          </string-name>
          , David Braze,
          <string-name>
            <given-names>Ashwini</given-names>
            <surname>Deo</surname>
          </string-name>
          , and Maria Mercedes Pin˜ango.
          <year>2012</year>
          .
          <article-title>Complement coercion: Distinguishing between type-shifting and pragmatic inferencing</article-title>
          .
          <source>The Mental Lexicon</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>58</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Composing and updating verb argument expectations: A distributional semantic model</article-title>
          .
          <source>In Proceedings of the 2nd Workshop on Cognitive Modeling and Computational Linguistics</source>
          , pages
          <fpage>58</fpage>
          -
          <lpage>66</lpage>
          , Portland, OR.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Kazunaga</given-names>
            <surname>Matsuki</surname>
          </string-name>
          , Tracy Chow, Mary Hare, Jeffrey L Elman,
          <string-name>
            <surname>Christoph Scheepers</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ken McRae</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Event-based plausibility immediately influences on-line language comprehension</article-title>
          .
          <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>
          ,
          <volume>37</volume>
          (
          <issue>4</issue>
          ):
          <fpage>913</fpage>
          -
          <lpage>934</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Liina</given-names>
            <surname>Pylkka</surname>
          </string-name>
          <article-title>¨nen and</article-title>
          <string-name>
            <surname>Brian McElree</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The syntax-semantics interface: On-line composition of sentence meaning</article-title>
          . In M. Traxler and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>A</article-title>
          . Gernsbacher, editors,
          <source>Handbook of Psycholinguistics</source>
          , pages
          <fpage>539</fpage>
          -
          <lpage>579</lpage>
          . Elsevier, Amsterdam, The Netherlands,
          <volume>2nd</volume>
          <fpage>edition</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>David E.</given-names>
            <surname>Rumelhart and James L. McClelland</surname>
          </string-name>
          .
          <year>1987</year>
          .
          <article-title>Learning the past tenses of English verbs. Implicit rules or parallel distributed processing</article-title>
          .
          <source>In Mechanisms of language acquisition</source>
          , pages
          <fpage>249</fpage>
          -
          <lpage>308</lpage>
          . Lawrence Erlbaum Associates, Hillsdale, NJ.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Jenny R Saffran</surname>
          </string-name>
          , Richard N Aslin, and Elissa L Newport.
          <year>1996</year>
          .
          <article-title>Statistical learning by 8-month-old infants</article-title>
          .
          <source>Science</source>
          ,
          <volume>274</volume>
          (
          <issue>5294</issue>
          ):
          <fpage>1926</fpage>
          -
          <lpage>1928</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Tomasello</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Constructing a language: A usage-based theory of language acquisition</article-title>
          . Harvard University Press, Cambridge, MA.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Zarcone</surname>
          </string-name>
          , Jason Utt, and Sebastian Pado´.
          <year>2012</year>
          .
          <article-title>Modeling covert event retrieval in logical metonymy: probabilistic and distributional accounts</article-title>
          .
          <source>In Proceedings of the 3rd Workshop on Cognitive Modeling and Computational Linguistics</source>
          , pages
          <fpage>70</fpage>
          -
          <lpage>79</lpage>
          , Montre´al, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Alessandra</given-names>
            <surname>Zarcone</surname>
          </string-name>
          , Alessandro Lenci, Sebastian Pado´, and
          <string-name>
            <given-names>Jason</given-names>
            <surname>Utt</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Fitting, not clashing! a distributional semantic model of logical metonymy</article-title>
          .
          <source>In Proceedings of the 10th International Conference on Computational Semantics</source>
          , Potsdam, Germany.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>