<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Distributional Semantics Approach to Implicit Language Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dimitrios Alikaniotis John N. Williams</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Theoretical and Applied Linguistics University of Cambridge</institution>
          <addr-line>9 West Road, Cambridge CB3 9DP</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2005</year>
      </pub-date>
      <fpage>81</fpage>
      <lpage>84</lpage>
      <abstract>
        <p>Copyright c by the paper's authors. Copying permitted for private and academic purposes. In Vito Pirrelli, Claudia Marzi, Marcello Ferro (eds.): Word Structure and Word Usage. Proceedings of the NetWordS Final Conference, Pisa, March 30-April 1, 2015, published at http://ceur-ws.org</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Vector-space models of semantics (VSMs) derive
word representations by keeping track of the
cooccurrence patterns of each word when found in
large linguistic corpora. By exploiting the fact that
similar words tend to appear in similar contexts
        <xref ref-type="bibr" rid="ref3">(Harris, 1954)</xref>
        , such models have been very
successful in tasks of semantic relatedness
        <xref ref-type="bibr" rid="ref11 ref6">(Landauer
and Dumais, 1997; Rohde et al., 2006)</xref>
        . A
common criticism addressed towards such models is
that those co-occurrence patterns do not explicitly
encode specific semantic features unlike more
traditional models of semantic memory
        <xref ref-type="bibr" rid="ref10 ref2">(Collins and
Quillian, 1969; Rogers and McClelland, 2004)</xref>
        .
Recently, however, corpus studies
        <xref ref-type="bibr" rid="ref1 ref4 ref5">(Bresnan and
Hay, 2008; Hill et al., 2013b)</xref>
        have shown that
some ‘core’ conceptual distinctions such as
animacy and concreteness are reflected in the
distributional patterns of words and can be captured by
such models
        <xref ref-type="bibr" rid="ref4 ref5">(Hill et al., 2013a)</xref>
        .
      </p>
      <p>
        In the present paper we argue that distributional
characteristics of words are particularly important
when considering concept availability under
implicit language learning conditions. Studies on
implicit learning of form-meaning connections have
highlighted that during the learning process a
restricted set of conceptual distinctions are available
such as those involving animacy and concreteness.
For example, in studies by
        <xref ref-type="bibr" rid="ref12">Williams (2005)</xref>
        (W)
and
        <xref ref-type="bibr" rid="ref7">Leung and Williams (2014)</xref>
        (L&amp;W) the
participants were introduced to four novel
determinerlike words: gi, ro, ul, and ne. They were
explicitly told that they functioned like the article ‘the’
but that gi and ro were used with near objects
and ro and ne with far objects. What they were
not told was that gi and ul were used with living
things and ro and ne with non-living things.
Participants were exposed to grammatical
determinernoun combinations in a training task and
afterwards given novel determiner-noun combinations
to test for generalisation of the hidden
regularity. W and L&amp;W report such a generalisation
effect even in participants who remained unaware
of the relevance of animacy to article usage –
semantic implicit learning.
        <xref ref-type="bibr" rid="ref9">Paciorek and Williams
(2015)</xref>
        (P&amp;W) report similar effects for a
system in which novel verbs (rather than determiners)
collocate with either abstract or concrete nouns.
However, certain semantic constraints on
semantic implicit learning have been obtained. In P&amp;W
generalisation was weaker when tested with items
that were of relatively low semantic similarity to
the exemplars received in training. In L&amp;W
Chinese participants showed implicit generalisation
of a system in which determiner usage was
governed by whether the noun referred to a long or
flat object (corresponding to the Chinese classifier
system) whereas there was no such implicit
generalisation in native English speakers. Based on
this evidence we argue that the implicit
learnability of semantic regularities depends on the degree
to which the relevant concept is reflected in
language use. By forming semantic representations
of words based on their distributional
characteristics we may be able to predict what would be
learnable under implicit learning conditions.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Simulation</title>
      <p>
        We obtained semantic representations using the
skip-gram architecture
        <xref ref-type="bibr" rid="ref8">(Mikolov et al., 2013)</xref>
        provided by the word2vec package,1 trained
with hierarchical softmax on the British National
Corpus or on a Chinese Wikipedia dump file of
comparable size. The parameters used were as
follows: window size: B5A5, vector dimensionality:
300, subsampling threshold: t = e−3 only for the
English corpus.
      </p>
      <p>The skip-gram model encapsulates the idea
of distributional semantics introduced above by
1https://code.google.com/p/word2vec/
0.8
n0.6
o
it
a
v
ti
c
A0.4
0.2
0.00
learning which contexts are more probable for a
given word. Concretely, it uses a neural network
architecture, where each word from a large
corpus is presented in the input layer and its context
(i.e. several words around it) in the output layer.
The goal of the network is to learn a configuration
of weights such that when a word is presented in
the input layer the nodes in the output that become
more activated correspond to those words in the
vocabulary, which had appeared more frequently
as its context.</p>
      <p>As argued above, the resulting representations
will carry, by means of their distributional
patterns, semantic information such as concreteness
or animacy. Consistent with the above
hypotheses, we predict that given a set of words in the
training phase, the degree to which one can
generalise to novel nouns will depend on how much
the relevant concepts are reeflcted in the former
words. If, for example, the words used during the
training session do not encode animacy based on
their co-occurrence statistics, albeit denoting
animate nouns, then generalising to other animate
nouns would be more difficult.</p>
      <p>In order to examine this prediction, we fed the
resulting semantic representations to a non-linear
0.150
1.00
0.75</p>
      <p>Abstract</p>
      <p>Concrete
classifier (a feedforward neural network) the task
of which was to learn to associate noun
representations to determiners or verbs, depending on the
study in question. During the training phase, the
neural network received as input the semantic
vectors of the nouns and the corresponding
determiners/verbs (coded as 1-in-N binary vectors, where
N is the number of novel non-words)2 in the
output vector. Using backpropagation with
stochastic gradient descent as the learning algorithm, the
goal of the network was to learn to discriminate
between grammatical and ungrammatical noun –
determiner/verb combinations. We hypothesise
that this could be possible if either specific
features of the input representation or a combination
of them contained the relevant concepts.
Considering the distributed nature of our semantic
representations, we explore the latter option by adding
a tanh hidden layer, the purpose of which was to
extract non-linear combinations of features of the
2All the studies reported use four novel non-words.
0.35
5
10
Epoch</p>
      <p>20
Grammatical
Ungrammatical
Chinese Grammatical
English Grammatical
250
Epoch</p>
      <p>500
Gr1a2m04matical
Ungrammatical
Abstract</p>
      <p>Concrete
English</p>
      <p>Chinese
input vector. We then recorded the generalisation
ability through time (epochs) of our classifier by
simply asking what would be the probability of
encountering a known determiner k with a novel
word w~ by taking the softmax function:
p(y = k|w~ ) =</p>
      <p>exp (netk)
Pk0∈K exp (netk0 )
.</p>
      <p>(1)
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion</title>
      <p>If the model has been successful in learning that
‘gi’ should be activated more given animate
concepts then the probability P (y = gi|w~ lion) would
be higher than P (y = ro|w~ lion). Fig. 1 shows the
performance of the classifier on the testing set of
W where, in the behavioural data, selection of the
grammatical item was significantly above chance
in a two alternative forced choice task for the
unaware group. The slopes of the gradients clearly
show that on such a task the model would favour
grammatical combinations as well.</p>
      <p>Figures 2-3 plot the results of two experiments
from P&amp;W which focused on the abstract/concrete
distinction. P&amp;W used a false memory task in the
generalisation phase, measuring learning by
comparing the endorsement rates between novel
grammatical and novel ungrammatical verb-noun pairs.</p>
      <p>It was reasoned that if the participants had some
knowledge of the system they would endorse more
novel grammatical sequences. Expt 1 (Fig. 2) used
generalisation items that were higher in
semantic similarity to trained items than was the case in
Expt 4 (Fig. 3). The behavioural results from the
unaware groups (bottom rows) show that this
manipulation resulted in larger grammaticality effects
on familiarity judgements in Expt 1 than Expt 4,
and also higher endorsements for concrete items
in general in Expt 1. Our simulation was able to
capture both of these effects.</p>
      <p>L&amp;W Expt 3 examined the learnability of a
system based on a long/flat distinction, which is
reflected in the distributional patterns of Chinese but
not of English. In Chinese, nouns denoting long
objects have to be preceded by a specific
classifier while flat object nouns by another. L&amp;W’s
training phase consisted of showing to participants
combinations of thin/flat objects with novel
determiners, asking them to judge whether the noun
was thin or flat. After a period of exposure,
participants were introduced to novel determiner – noun
combinations, which either followed the
grammatical system (control trials) or did not (violation
trials). Participants had significantly lower reaction
times (Fig. 4, bottom row) when presented with a
novel grammatical sequence than an
ungrammatical sequence, an effect not observed in the RTs
of the English participants. The corresponding
results of our simulations plotted in Fig. 4 show that
indeed the regularity was learnable when the
semantic model had only experienced a Chinese text,
but not when it experienced the English corpus.</p>
      <p>While more direct evidence is needed to support
our initial hypothesis, our results seem to point
to the direction that semantic information encoded
by the distributional characteristics of words when
found in large corpora can be important in
determining what could be implicitly learnable.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Bresnan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Hay</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Gradient grammar: An effect of animacy on the syntax of give in New Zealand and American English</article-title>
          . Lingua,
          <volume>118</volume>
          (
          <issue>2</issue>
          ):
          <fpage>245</fpage>
          -
          <lpage>259</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Collins</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Quillian</surname>
            ,
            <given-names>M. R.</given-names>
          </string-name>
          (
          <year>1969</year>
          ).
          <article-title>Retrieval time from semantic memory</article-title>
          .
          <source>Journal of Verbal Learning and Verbal Behavior</source>
          ,
          <volume>8</volume>
          (
          <issue>2</issue>
          ):
          <fpage>240</fpage>
          -
          <lpage>247</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>1954</year>
          ).
          <article-title>Distributional structure</article-title>
          .
          <source>Word</source>
          ,
          <volume>10</volume>
          (
          <issue>23</issue>
          ):
          <fpage>146</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiela</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Korhonen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2013a</year>
          ).
          <article-title>Concreteness and Corpora: A Theoretical and Practical Analysis</article-title>
          .
          <source>In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics</source>
          , pages
          <fpage>75</fpage>
          -
          <lpage>83</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korhonen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bentz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2013b</year>
          ).
          <article-title>A Quantitative Empirical Analysis of the Abstract/Concrete Distinction</article-title>
          .
          <source>Cognitive Science</source>
          ,
          <volume>38</volume>
          (
          <issue>1</issue>
          ):
          <fpage>162</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T. K.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S. T.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <article-title>A solution to Plato's problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge</article-title>
          .
          <source>Psychological Review</source>
          ,
          <volume>104</volume>
          (
          <issue>2</issue>
          ):
          <fpage>211</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Leung</surname>
            ,
            <given-names>J. H. C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J. N.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Crosslinguistic Differences in Implicit Language Learning</article-title>
          .
          <source>Studies in Second Language Acquisition</source>
          ,
          <volume>36</volume>
          (
          <issue>4</issue>
          ):
          <fpage>733</fpage>
          -
          <lpage>755</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G. S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Paciorek</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Semantic generalisation in implicit language learning</article-title>
          .
          <source>Journal of Experimental Psychology: Learning, Memory and Cognition.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Rogers</surname>
          </string-name>
          , T. T. and
          <string-name>
            <surname>McClelland</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Semantic Cognition: A Parallel Distributed Processing Approach</article-title>
          . MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Rohde</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonnerman</surname>
            ,
            <given-names>L. M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Plaut</surname>
            ,
            <given-names>D. C.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>An improved model of semantic similarity based on lexical co-occurrence</article-title>
          .
          <source>Communications of the ACM.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J. N.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Learning without awareness</article-title>
          .
          <source>Studies in Second Language Acquisition</source>
          ,
          <volume>27</volume>
          :
          <fpage>269</fpage>
          -
          <lpage>304</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>