<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modelling semantic transparency in English compound nouns</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Melanie J. Bell</string-name>
          <email>melanie.bell@anglia.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Schäfer</string-name>
          <email>post@martinschaefer.info</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Anglia Ruskin University</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">U.K.</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Friedrich Schiller University</institution>
          ,
          <addr-line>Jena</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>63</fpage>
      <lpage>65</lpage>
      <abstract>
        <p>Semantic transparency is known to play an important role in the storage and processing of complex words (e.g. Marslen-Wilson et al. 1994), and human raters of transparency achieve high levels of agreement (e.g. Frisson et al. 2008, Munro et al. 2010). In the case of noun-noun compounds, overall transparency is largely determined by the transparency of the individual constituents. For example, Reddy et al. (2011) showed that the perceived transparency of a compound is highly correlated with both the sum and the product of the perceived transparencies of its constituents. Furthermore, many psycholinguistic studies find significant effects for semantic transparency using a four-way distinction based on perceived constituent transparency: transparent-transparent (e.g. carwash), transparent-opaque (e.g. jailbird), opaque-transparent (e.g. strawberry) and opaque-opaque (e.g. hogwash) (Libben et al. 2003). Bell and Schäfer (2013) modelled the transparency of individual compound constituents and showed that shifted word senses reduce perceived transparency, while certain semantic relations between constituents increase it. However, this finding is problematic in at least two ways. Firstly, it is not clear whether there is a solid basis for establishing whether a specific word sense is shifted or not. For example, card in credit card is clearly shifted if viewed etymologically, but may not synchronically be perceived as shifted due to its frequent use. Secondly, work on conceptual combination by Gagné and collaborators has shown that relational information in compounds is accessed via the concepts associated with individual modifiers and heads, rather than independently of them (e.g. Spalding et al. 2010 for an overview). This leads to the hypothesis that it is not whether a specific word sense is etymologically shifted, nor whether a specific semantic relation is used per se, that makes a compound constituent more or less transparent; rather, it is Copyright © by the paper's authors. Copying permitted for private and academic purposes. In Vito Pirrelli, Claudia Marzi, Marcello Ferro (eds.): Word Structure and Word Usage. Proceedings of the NetWordS Final Conference, Pisa, March 30-April 1, 2015, published at http://ceur-ws.org</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>the degree of expectedness of a particular word
sense and a particular relation for a given
constituent. In this paper, we provide evidence in
support of this hypothesis: the more expected the
word sense and relation for a constituent, the
more transparent it is perceived to be.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Method</title>
      <p>
        We used the publicly available dataset described
in
        <xref ref-type="bibr" rid="ref7">Reddy et al. (2011)</xref>
        , which gives human
transparency ratings for a set of 90 compound types
and their constituents (N1 and N2), and
comprises a total of 7717 ratings. To model the
expectedness of word senses and semantic relations for
a given compound constituent, we used the
constituent families of the compounds, which we
extracted in a two step process. We took all
strings of exactly two nouns that follow an article
in the British National Corpus and which also
occur four times or more in the USENET corpus
        <xref ref-type="bibr" rid="ref8">(Shaoul and Westbury 2010)</xref>
        . From this set, we
extracted the positional constituent families for
all constituent nouns in the Reddy et al. dataset,
giving a total of 4553 compounds for the N1
families and 9226 for the N2 families. Each of
these compound types was coded for the
semantic relation between the constituents
        <xref ref-type="bibr" rid="ref3">(after Levi
1978)</xref>
        , and for the WordNet sense of the
constituent under consideration
        <xref ref-type="bibr" rid="ref6">(Princeton 2010)</xref>
        . We
then calculated the proportion of compound
types in each constituent family with each
semantic relation (relation proportion), and each
WordNet sense of the constituent in question
(synset proportion). We take these two measures
to reflect the expectedness of the respective
relations and WordNet senses of the constituents: if a
relation or sense occurs in a high proportion of
the constituent family, it is more expected. These
variables were used, along with other
quantitative measures, as predictors in ordinary least
squares regression models of perceived
constituent transparency. The final model for the
transparency of N1 is given in Table 1:
Intercept
relation proportion in N1family
log family size of N1
synset proportion in N1family
log synset count of N1
compound proportion in N1 family (token-based)
log frequency of N1
relation proportion * log family size
synset proportion * log synset count
compound proportion * log frequency N1
1 6
N
feo 5
z
i
s
ily 4
m
a
fg 3
o
l
All predictors in our model enter into significant
interactions, and these are shown graphically in
Figure 1, where the contour lines on the plots
represent perceived transparency of the first
constituent (N1). The first plot shows an interaction
between relation proportion and overall (log)
family size: for small families, relation
proportion plays little role, whereas for larger families,
in accordance with our hypothesis, the
transparency of N1 increases with the proportion of the
corresponding relation in the family. The second
plot shows the interaction between the synset
proportion and the total number of a
constituent’s senses (as listed in WordNet): only if there
is a sufficient number of different senses in the
family is their proportion a reliable predictor of
semantic transparency. There is also a small but
significant interaction between the log frequency
of a constituent and the proportion of the
constituent family (in terms of tokens) represented by
the compound in question: this shows that
transparency increases with frequency, but only in the
lower frequently ranges does the proportion in
the family play a role.
4
      </p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>Overall, the model provides clear evidence for
our hypothesis. N1 is rated as most transparent
when it is a frequent word, with a large family,
occurring with its preferred semantic relation and
most frequent sense, and with few other senses to
compete. We interpret the results as indicating
that compound constituents are perceived as
more transparent when they are more expected
(both generally and with a specific sense) and
when they occur in their most expected semantic
environments. In information theory, the less
expected an event, the greater its information
content: in so far as perceived transparency is a
reflection of expectedness, it can therefore also
be seen as the inverse of informativity.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>This work was made possible by three short visit
grants from the European Science Foundation
through NETWORDS - The European Network
on Word Structure (grants 4677, 6520 and 7027),
for which the authors are extremely grateful.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Bell</surname>
            ,
            <given-names>Melanie J.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Schäfer</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Semantic transparency: challenges for distributional semantics</article-title>
          . In Aurelie Herbelot, Roberto Zamparelli and Gemma Boleda eds.,
          <source>Proceedings of the IWCS 2013 workshop: Towards a formal distributional semantics</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . Potsdam:
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Frisson</surname>
            , Steven, Elizabeth Niswander-Klement and
            <given-names>Alexander</given-names>
          </string-name>
          <string-name>
            <surname>Pollatsek</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>The role of semantic transparency in the processing of English compound words</article-title>
          .
          <source>British Journal of Psychology</source>
          <volume>991</volume>
          ,
          <fpage>87</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Levi</surname>
            ,
            <given-names>Judith N.</given-names>
          </string-name>
          <year>1978</year>
          .
          <article-title>The syntax and semantics of complex nominals</article-title>
          . New York: Academic Press.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Marslen-Wilson</surname>
            , William, Lorraine K. Tyler, Rachelle Waksler and
            <given-names>Lianne</given-names>
          </string-name>
          <string-name>
            <surname>Older</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Morphology and meaning in the English mental lexicon</article-title>
          .
          <source>Psychological Review</source>
          <volume>101</volume>
          ,
          <issue>1</issue>
          :
          <fpage>3</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Munro</surname>
            , Robert, Steven Bethard, Victor Kuperman, Vicky Tzuyin Lai , Robin Melnick, Christopher Potts, Tyler Schnoebelen and
            <given-names>Harry</given-names>
          </string-name>
          <string-name>
            <surname>Tily</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Crowdsourcing and language studies: the new generation of linguistic data</article-title>
          .
          <source>In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech</source>
          and
          <article-title>Language Data with Amazon's Mechanical Turk</article-title>
          , pp.
          <fpage>122</fpage>
          -
          <lpage>130</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          Princeton University.
          <year>2010</year>
          . WordNet.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Reddy</surname>
            , Siva,
            <given-names>Diana McCarthy</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Suresh</given-names>
            <surname>Manandhar</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>An empirical study on compositionality in compound nouns</article-title>
          .
          <source>In Proceedings of The 5th International Joint Conference on Natural Language Processing 2011 IJCNLP</source>
          <year>2011</year>
          ,
          <string-name>
            <given-names>Chiang</given-names>
            <surname>Mai</surname>
          </string-name>
          , Thailand
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Shaoul</surname>
            , Cyrus and
            <given-names>Chris</given-names>
          </string-name>
          <string-name>
            <surname>Westbury</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>An anonymized multi-billion word USENET corpus</article-title>
          <year>2005</year>
          -2010 http://www.psych.ualberta.ca/˜westburylab/downl oads/usenet.download.html
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Spalding</surname>
          </string-name>
          , Thomas L.,
          <string-name>
            <surname>Christina L. Gagné</surname>
          </string-name>
          , Allison C.
          <article-title>Mullaly</article-title>
          and
          <string-name>
            <given-names>Hongbo</given-names>
            <surname>Ji</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Relation-based interpretation of noun-noun phrases: A new theoretical approach</article-title>
          .
          <source>Linguistische Berichte Sonderheft</source>
          <volume>17</volume>
          ,
          <fpage>283</fpage>
          -
          <lpage>315</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Wurm</surname>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            <given-names>H.</given-names>
          </string-name>
          <year>1997</year>
          .
          <article-title>Auditory processing of prefixed English words is both continuous and decompositional</article-title>
          .
          <source>Journal of Memory and Language</source>
          ,
          <volume>37</volume>
          ,
          <fpage>438</fpage>
          -
          <lpage>461</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>