<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Of crowds and corpora: A marriage of measures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emmanuel Keuleers</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paweł Mandera</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michaël Stevens</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marc Brysbaert</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Copyright © by the paper's authors. Copying permitted for private and academic purposes. In Vito Pirrelli, Claudia Marzi, Marcello Ferro (eds.): Word Structure and Word Usage. Proceedings of the NetWordS Final Conference</institution>
          ,
          <addr-line>Pisa</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Experimental Psychology Ghent University Henri Dunantlaan 2</institution>
          ,
          <addr-line>9000 Gent</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <fpage>10</fpage>
      <lpage>12</lpage>
      <abstract>
        <p>We discuss the relationship between a word's corpus frequency and its prevalence -the proportion of people who know the word- and show that they are complementary measures. We show that adding word prevalence as a predictor of lexical decision reaction time in the Dutch lexicon project increases explained variance by more than 10%. In addition, we show that, for the same dataset, word prevalence is the best independent predictor of word processing time.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Word frequency is one of the most important
measures in the cognitive study of word
processing, both theoretically and methodologically.
Its contribution in explaining behavioural
measures such as reaction time is so large that
researchers take great care in collecting large and
reliable corpora and in applying the best possible
word frequency estimates in their research.</p>
    </sec>
    <sec id="sec-2">
      <title>Where the corpus is weak the crowd is strong</title>
      <p>A drawback of frequency counts is that,
regardless of corpus size, lower counts are
unreliable. As an example, consider asking a
random sample of 100 people whether they
know each of the word types that occur just
once in a large corpus. Although frequency
for all these types is equal, the number of
judges knowing each word will vary from
zero to one hundred and, as the judges are
language users, words known to many of
them may be considered to occur more often
in language than words which are known by
fewer of them. Following this reasoning, the
estimate of the number of language users
who know a word, or word prevalence may
give a better indication of occurrence than
corpus frequency counts.
1.2</p>
    </sec>
    <sec id="sec-3">
      <title>Where the corpus is strong the crowd is weak</title>
      <p>On the other hand, consider presenting the
same random sample of people with words
from the language's core vocabulary. Since
these words will be known to all of the
judges, prevalence will be singularly high
and uninformative. In this case corpus counts
should be a much better estimate of
occurrence.
2</p>
    </sec>
    <sec id="sec-4">
      <title>Testing the prevalence measure</title>
      <p>To test the complementarity of prevalence
and frequency as measures of occurrence, we
used prevalence norms for Dutch collected
through a lexical decision task presented as
an online vocabulary test (Keuleers, Stevens,
Mandera, &amp; Brysbaert, in press). Each
participants saw 100 stimuli (about 70 words
and 30 nonwords) selected randomly from a
list of 54,319 words and 21,734 nonwords.
In the current analysis, we used the data of
190,771 participants who indicated that they
were living in Belgium, giving us about 250
observations per word. The score for a word
obtained by fitting a Rasch model –a
mathematical model simultaneously ranking
participants by ability and test-items by difficulty–
to the data was considered an
operationalization of its prevalence.</p>
      <p>In addition, we investigated the relationship
between prevalence and other typical
measures of word frequency. Table 1 gives an
overview of these correlations.</p>
      <p>Frequency Prevalence OLD 20 Length
icon ProjectTable 1 shows that the
correlation between prevalence and frequency was
relatively low (.34), giving further evidence
that prevalence is distinct from word
frequency and contextual diversity –a word's
document count– which correlates very
highly with word frequency.</p>
      <p>
        Finally, we used the data from the 7,885
items in the Dutch Lexicon Project
        <xref ref-type="bibr" rid="ref2 ref3">(Keuleers
et al., 2010)</xref>
        for which both frequency and
prevalence were available to examine the
contributions of Dutch corpus word
frequency
        <xref ref-type="bibr" rid="ref2 ref3">(SUBTLEX-NL, Keuleers et al.,
2010)</xref>
        and word prevalence on average
reaction times.
      </p>
      <p>In single variable analyses, log word
frequency explained about 36.13% of the
variance in reaction times and prevalence
explained about 33.03% of the variance in
reaction times.</p>
      <p>This was also made clear when both
measures were considered in the same analysis,
where both measures jointly explained 51.37
% of the variance in reaction times. The
unique contributions to explained variance
(eta-squared) were 27.39% for frequency and
23.87% for prevalence. In further analyses,
we found that including the quadratic trend
of word frequency and contextual diversity
did not substantially alter this pattern of
results.
3</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The results show that, next to word
frequency, prevalence is by far the most
important independent contributor to visual word
recognition times, suggesting that prevalence
should be included in any analysis where
word corpus frequency is considered to be
relevant. However, several questions remain
open. First, what is the influence of corpus
size on the relation between corpus word
frequency and prevalence and on the
contribution of prevalence to lexical processing?
Second, how well does prevalence perform on
others tasks and in other languages? Finally,
does the effect of prevalence on word
processing truly lie in a better measurement of
word occurrence or does it partly reflect an
independent property associated with the
learnability of a word?</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The text of this abstract is an early summary of find
ings from a larger study reported in the Quarterly
Journal of Experimental Psychology as Word
knowledge in the crowd: Measuring vocabulary size and
word prevalence in a massive online experiment.
(Keuleers, E., Stevens, M., Mandera, P., &amp; Brysbaert,
M., in press).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Balota</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yap</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hutchison</surname>
            ,
            <given-names>K. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cortese</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kessler</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loftis</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , … Treiman,
          <string-name>
            <surname>R.</surname>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>The English lexicon project</article-title>
          .
          <source>Behavior Research Methods</source>
          ,
          <volume>39</volume>
          (
          <issue>3</issue>
          ),
          <fpage>445</fpage>
          -
          <lpage>459</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Keuleers</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brysbaert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>New</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>SUBTLEX-NL: A new measure for Dutch word frequency based on film subtitles</article-title>
          .
          <source>Behavior Research M e t h o d s , 4</source>
          <volume>2</volume>
          (
          <issue>3</issue>
          ) ,
          <article-title>6 4 3 - 6 5 0</article-title>
          . doi:
          <volume>10</volume>
          .3758/BRM.42.3.
          <fpage>643</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Keuleers</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diependaele</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Brysbaert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Practice Effects in Large-Scale Visual Word Recognition Studies: A Lexical Decision Study on 14,000 Dutch Mono-</article-title>
          and
          <string-name>
            <surname>Disyllabic Words</surname>
          </string-name>
          and Nonwords. Frontiers in Psychology,
          <volume>1</volume>
          . doi:
          <volume>10</volume>
          .3389/fpsyg.
          <year>2010</year>
          .00174
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Keuleers</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lacey</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rastle</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Brysbaert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>The British Lexicon Project: Lexical decision data for 28,730 monosyllabic and disyllabic English words</article-title>
          .
          <source>Behavior Research Methods</source>
          ,
          <volume>44</volume>
          (
          <issue>1</issue>
          ),
          <fpage>287</fpage>
          -
          <lpage>304</lpage>
          . doi:
          <volume>10</volume>
          .3758/s13428-011-0118-4
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Keuleers</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevens</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandera</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Brysbaert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>(in press). Word knowledge in the crowd: Measuring vocabulary size and word prevalence in a massive online experiment</article-title>
          .
          <source>Quarterly Journal of Experimental Psychology. doi:10.1080/17470218</source>
          .
          <year>2015</year>
          .1022560
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>