<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Interpretation of Generalization in Masked Language Models: An Investigation Straddling Quantifiers and Generics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Collacciani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giulia Rambelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Bologna</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Generics are statements that express generalizations and are used to communicate generalizable knowledge. While generics convey general truths (e.g., Birds can fly ), they often allow for exceptions (e.g., penguins do not fly). Nonetheless, generics form the basis of how we communicate our commonsense about the world [1, 2]. We explored the interpretation of generics in Masked Language Models (MLMs), building on psycholinguistic experimental designs. As this interpretation requires a comparison with overtly quantified sentences, we investigated i) the probability of quantifiers, ii) the internal representation of nouns in generic vs. quantified sentences, and iii) whether the presence of a generic sentence as context inuflences quantifiers' probabilities. The outcomes confirm that MLMs are insensitive to quantification; nevertheless, they appear to encode a meaning associated with the generic form, which leads them to reshape the probability associated with various quantifiers when the generic sentence is provided as context.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Generics</kwd>
        <kwd>Quantifiers</kwd>
        <kwd>Masked Language Models</kwd>
        <kwd>Commonsense Knowledge</kwd>
        <kwd>Pragmatics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        there is any bias towards some of them (for ex- 2.2. Genericity in NLP
ample, a generic overgeneralization efect such
as that found in humans by Leslie et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
2. Are the hidden representations of generics similar
to those of quantified phrases? We extracted the
hidden representation of words in generic and
quantified sentences and compared their
representation pairwise to understand which
quantiifed nominal phrase approximates the meaning
of generics better.
3. Are LLMs showing the same prevalence efect as
humans? We reproduced the experimental design of
Cimpian et al. [7] Implied Prevalence task to test
whether the presence of the generic as a premise
impacts the probability of selected quantifiers.
      </p>
      <sec id="sec-1-1">
        <title>Most of the NLP literature on genericity has focused on</title>
        <p>the creation and annotation of resources for identifying
generic expressions as opposed to non-generic ones and,
based on these resources, on the development of
automatic annotation systems [10, 11, 12, among others]</p>
        <p>To the best of our knowledge, there are no studies
investigating the interpretation of generalizations by LLMs,
except for the recent work by Ralethe and Buys [13],
which addresses the generic overgeneralization efect
in BERT and RoBERTa. The authors argue that these
models sufer from overgeneralization by assessing how
many times one or more of all, every, most, some, few and
many are predicted in a masked sentence like [MASK]
lions have manes: the higher the rank of the quantifiers,
the stronger the LM exhibits the GOG efect. However,
the GOG efect refers to the acceptance of universally
quantified sentences, not just quantified ones; therefore,
we can speak of overgeneralization only when the
preferred quantifier is the universal one ( all or every). For
this reason, we first propose a similar task to evaluate
the probability distribution of various quantifiers,
distinguishing between them qualitatively.</p>
      </sec>
      <sec id="sec-1-2">
        <title>The data and code that we used for the experiments</title>
        <p>are publicly available1.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <sec id="sec-2-1">
        <title>2.1. Generics in Human Cognition</title>
        <p>
          Experimental evidence has revealed a generic bias for
which people tend to overgeneralize from the truth of a
generic to the truth of the corresponding universal state- 3. Materials and Methods
ment [
          <xref ref-type="bibr" rid="ref6">6, 8, 9</xref>
          ]. For example, people tend to accept the
statement All lions have manes as true, even though it is Data For this study, we selected the generic sentences
not, because they rely on the truth of the corresponding from the dataset of Allaway et al. [14]. The authors
exgeneric Lions have manes. This efect is known as the tracted 653 generics about objects, animals, and plants
generic overgeneralization (GOG) efect [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It is detected from Bhagavatula et al. [15] and annotated them into
only on certain categories of generics, namely those of three categories obtained unifying theories from
linguisminority characteristic and majority characteristic gener- tics and philosophy, by condensing the five types of
generics, i.e., generics that predicate properties that are true ics proposed by Leslie [16, 17] and Khemlani et al. [18]. In
for a minority or the majority of category members. quasi-definitional sentences, the property is essential
        </p>
        <p>Cimpian et al. [7] conducted a series of studies inves- to a concept, thus is considered a defining characteristic
tigating the relationship between genericity and preva- of the concept (e.g., triangles have three sides). In this
lence and found an inferential asymmetry in the meaning type of sentences, the generic is de facto equivalent to
of generics. People tend to judge a generic sentence about the corresponding universal quantified statement (e.g.,
a novel category as true even if they have been informed all triangles have three sides). In principled sentences,
that only a certain percentage of the kind (on average, up the property has a strong association with the concept.
to less than 70 percent) possess the property in question This category includes both properties that are viewed as
(Truth Condition task). However, when asked to estimate inherent, or connected in a principled way with a concept
how many members of the kind possess the property, (e.g., birds can fly. ), and properties that are uncommon
given the generic (Implied Prevalence task), they tend to and often dangerous (e.g., sharks attack swimmers); this
assign very high percentages (on average, very close to last case is the one that Leslie [16, 17] defines striking.
Fi100 percent). This study indicates that generic sentences nally, characterizing sentences express a non-accidental
require little evidence to be judged true but have substan- relationship between property and concept, based only
tial implications, since the properties they predicate tend on absolute or relative prevalence among category
memto be interpreted as applying to virtually all members of bers. These generics concern properties that are neither
the category. deeply connected to the concept nor striking, but occur in
the majority (majority characteristic generics for [16, 17],
e.g., Cars have radios) or in the minority of members of
1https://github.com/claudiacollacciani/ the category (minority characteristic generics for [16, 17],
Interpretation-of-Generalization-in-Masked-Language-Models e.g., Lions have manes).</p>
        <sec id="sec-2-1-1">
          <title>This analysis serves as a baseline to understand the</title>
          <p>
            probability distribution of quantifiers, if there is any bias
towards some of them (i.e., overgeneralization efect), and
possibly whether the belonging of sentences to diferent
categories impacts it (as observed for humans by Leslie
et al. [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]).
          </p>
        </sec>
        <sec id="sec-2-1-2">
          <title>From the original batch, we restricted our choice to</title>
          <p>207 generic sentences, picking only the ones in the bare
plural form (e.g., Tigers are striped), excluding
indefinite and definite singulars (e.g., A/The tiger is striped).
All these syntactic forms can express generic meanings,
but the bare plural is the only surface form in English
that gives rise to a generic interpretation unambiguously
[19]. For this and other reasons, this is considered as the
paradigmatic case, and it is the one that has been used in
the psycholinguistic experiments from which we draw
inspiration.</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>Results Figure 1 reports the quantifiers distributions</title>
          <p>for the base models, as the larger counterparts show a
similar trend (all boxplots are in Appendix B). Overall,
all models consider few the least likely. This outcome
reflects our expectations: as the selected sentences are
Models We experimented with BERT and RoBERTa, all generalizations, they are, in most cases, referable to
two bidirectional Masked Language Models (MLMs) a substantial number of category members, rarely to
based on the Transformer architecture. BERT [20] is ‘few’ members. Apart from that, BERT and RoBERTa
trained both on a masked language modeling task and show diferent probability distributions of quantifiers,
on a next sentence prediction task, as the model receives regarding some and all in particular. BERT models assign
sentence pairs in input and has to predict whether the a higher probability to the existential and proportional
second sentence is after the first one in the training data. quantifiers ( some, many, and most) than the universal
BERT has been trained on the BookCorpus and the En- quantifier all, and some is overall the most expected. The
glish Wikipedia for around 3300M tokens. We employed diferences among the quantifier scores are statistically
the bert-base-uncased and bert-large-uncased significant 3, with few exceptions (see Appendix B). It is
pre-trained versions, which difer in terms of parameters 3By relying on Wilcoxon Signed-Rank Test statistical test.
(110M and 340M parameters, respectively). On the other
hand, RoBERTa [21] has the same architecture as BERT;
however, it introduces several parameter optimization
choices, such as dynamic masking, a larger batch and
vocabulary size, and the removal of the next sentence
prediction objective. Another key diference is the larger
training corpus: RoBERTa was trained on 160GB of texts.</p>
          <p>We relied on the Huggingface’s Transformers2 Library
to load the models and carry on our experiments.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experiments</title>
      <sec id="sec-3-1">
        <title>4.1. Experiment 1: Probability distribution of Quantifiers in MLMs</title>
        <sec id="sec-3-1-1">
          <title>In the first place, we needed to assess what was the</title>
          <p>most expected quantifier for the sentences in our dataset.
Therefore, we modified the original generic sentences
by placing the special token [MASK] at the beginning of
each sentence, as in ‘[MASK] strawberries have a sweet
lfavor .’ Then, we computed the conditional log
probability of quantifiers few, some, many, most, and all in
the masked position, following previous works in
quantification [ 22, 13, 23]. The conditional log probability is
defined as
() =  (|)
(1)
where  are the words preceding and following the critical
word in the sentence.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>2https://huggingface.co/docs/transformers/index</title>
          <p>worth noticing that the reported distributions remain the
same even when we separate the analysis by sentence
categories. In other words, BERT models are not sensitive
to the pragmatic diferences of the selected sentences.</p>
          <p>
            Conversely, both all and some are the most expected
quantifiers for RoBERTa. Accordingly, the universal
quantifier all seems to be more expected by RoBERTa
than by BERT, but this distribution is not constant for
all three sentence conditions. In quasi-definitional and
characterizing sentences, some and all have the same
probability; alternatively, some is more expected in
principled sentences. We could draw that, for principled
sentences, the model prefers not to overgeneralize the
property to all members of the category. This behavior
seems to approximate Leslie et al. [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] results: people tend
to overgeneralize (i.e., to accept as true the universal
sentence corresponding to the generic) in the case of
characterizing sentences, while they do not
overgeneralize (correctly) in the case of striking sentences (included
in the principled category).
          </p>
          <p>Regardless, the observed trends could be determined
by the overall frequency of the quantifiers. As a
sanity check, we extracted their frequency from a large
corpus of English, enTenTen21 [24, 25]. We found
that  () &gt;  () &gt;  () &gt;
 ( ) &gt;  () (frequencies are reported
in Appendix A). This pattern confirms that few is not
the less probable because of a frequency efect but for
the properties of the sentence. Conversely, most is the
less frequent but has a probability score similar to the
more frequent many and all. Finally, some is overall the
most frequent quantifier. This observation could partially
reflect the probability outputs of BERT; however, it is not
the case for RoBERTa scores.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>4.2. Experiment 2: Representation of</title>
        <p>words in Generics and Quantified</p>
      </sec>
      <sec id="sec-3-3">
        <title>Sentences</title>
        <p>relation4.</p>
        <p>This study has a twofold aim: i) Identify the
quantiifer that shifts the noun representation closer to that of
the generic statement (if possible), and ii) Localize the
layers where quantification emerges. To the best of our
knowledge, no previous work has explored the internal
representations of quantified expressions in relation to
genericity.</p>
        <p>Results Figure 2 illustrates how the Spearman’s 
between the noun and the quantified version changes with
respect to each hidden layer in BERT and RoBERTa base
(see Appendix C for the plots of larger models and the
ones reporting correlations by sentence categories). The
ifrst layers do not show a diference among correlations,
meaning that representations characterized by the
diferent quantifiers are practically identical; this is expected,
as the context is limitedly attended by the model in these
layers. The following layers show a gradual change in
correlation values, but BERT and RoBERTa show
diferent patterns. For the first, we observe a slight decrease
in scores from layers 3 to 9 (but the correlation values
are still above 0.9). Conversely, the peak of the curve is
at layer 5 for RoBERTa ( 0.76), while the other internal</p>
        <sec id="sec-3-3-1">
          <title>4The authors reported that Spearman’s  is more robust to rogue</title>
          <p>dimensions contextual language models than cosine or Euclidean
similarity measures.</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>The architecture of MLMs allows us to follow the trans</title>
          <p>formations of each token throughout the neural network.</p>
          <p>Previous works in BERTology have reported that internal
representations, also known as contextualized
embeddings, encode syntactic and semantic properties in
diferent hidden layers [26]. However, it is complex to localize
semantic phenomena, as they spread across the entire
model [27]. For our purposes, we decided to compare the
contextualized embedding of a target token (strawberries)
in the generic sentence (strawberries have a sweet flavor )
with the embedding of the target token in each of the
corresponding quantified sentences (e.g., all strawberries
have..). Following Timkey and van Schijndel [28], for
each layer, we computed the similarity of the two
contextualized embeddings by relying on Spearman’s  cor- Figure 2: Spearman’s  against model layers.
layers have a constant value of around 0.73. In the last
intermediate layers (8-10), we observe some drifting away
between the values for the diferent quantifiers,
indicating that the contribution of the quantifiers difers from
one to another. For instance, in BERT-base and large the
noun quantified by some is the most similar to the
corresponding non-quantified. At the same time, RoBERTa
has generics similar to all, which could be interpreted as
an overgeneralization efect.</p>
          <p>Intriguingly, the most similar representations are the
most probable quantifiers of the previous experiment,
thus confirming that the choice of the quantifier is not a
frequency efect but is related to providing a
representation closer to that of the generic statement. However, the
values for the various quantifiers always remain close
to each other and follow the same trend, so it is hard
to disclose how the meaning of the quantifier afects
the noun representation. Finally, all models worsened
their performance at the last layer - the one producing
the most context-specific representations [ 29], indicating
that contextual information weakens the quantification
signal.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>4.3. Experiment 3: Implied Prevalence efects in MLMs</title>
        <p>As observed above, MLMs are not particularly sensitive
to quantifiers, and the probability choices are
independent of the sentence’s meaning. This outcome is mainly
due to the fact that these models are agnostic to world
knowledge. Therefore, we decided to test the relation
between quantification and generalization from a more
formal point of view: we examined how models interpret
generalizations aside from their content, that is, whether
they contain any linguistic information associated with
the form of generics.</p>
        <p>
          We reproduce the experimental design of Cimpian
et al. [7] Implied Prevalence task, in which people were
presented with a generic sentence about a novel animal
category and then asked to estimate how many members
of the category possess the characteristic predicated by
the generic (e.g., Information: Morseths have silver fur.
Question: What percentage of morseths do you think have
silver fur?). In this case, world knowledge is not called
into play, unlike in Leslie et al. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]’s experiments: the
categories employed are made up, and thus lack
associations to properties in the speakers’ mind. Since models
do not seem to encode the world knowledge necessary
to interpret generics on account of their content (with
the partial exception of RoBERTa), this experimental
design may be suitable for investigating instead the default
interpretation they associate with a generic form.
        </p>
        <p>We build the stimuli using the generic sentence as
the premise in the following way: Strawberries have a
sweet flavor means that [MASK] strawberries have a sweet
lfavor. As in Experiment 1, we compute the log
probability of quantifiers ( few, some, many, most, and all) in
the masked position. This last study should answer the
following question: Does the presence of a generic
sentence impact the quantifier preference, compared to the
a-contextualized version of the first analysis?</p>
        <sec id="sec-3-4-1">
          <title>Results Figure 3 reports the quantifiers distributions</title>
          <p>for the base models, as the larger counterparts show the
same trend (all boxplots are in Appendix D). Surprisingly,
the most expected quantifier is now all for all models5.
BERT shows an inversion in the ratio between the
probability associated with the universal quantifier all and that
associated with the other quantifiers (apart from few),
which in the first experiment were all more expected
than all, whereas now are all less expected. The pattern
exhibited by RoBERTa shows a less striking change than
BERT; however, the probability of all with respect to the
other quantifiers is still higher than in Experiment 1. As
an additional check, we performed the same test with a
very small dataset of 24 generic sentences about novel</p>
        </sec>
        <sec id="sec-3-4-2">
          <title>5The diferences among the quantifier scores are statistically signifi</title>
          <p>
            cant; a few exceptions are listed in Appendix D.
categories (non-words) from Cella et al. [
            <xref ref-type="bibr" rid="ref7">30</xref>
            ],6 to be sure
that the content of the sentences would not afect the
results. As we expected, we obtained the same pattern
as with our dataset.
          </p>
          <p>The results of this experiment, when compared with
the baseline for the choice of quantifiers in Experiment
1, suggest that the presence of a generic sentence as a
premise does indeed have an impact on the preference of
the quantifier by the models. When the generalization is
provided as context, the preferred quantifier becomes all.
This behavior mirrors that of people, observed in Cimpian
et al. [7] experiments and later replications: people tend
to estimate very high percentages (on average very close
to 100 percent) in the Implicit Prevalence task.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Discussion and Conclusions</title>
      <sec id="sec-4-1">
        <title>In this paper, we analyzed the interpretation of generics in MLMs through psycholinguistic experimental designs that exploited quantified expressions to investigate the understanding of generic ones.</title>
        <p>The first two experiments raise questions about
the codification of quantifiers, as it seems that the
models do not substantially exhibit a strong sensitivity
to quantifiers and do not encode a semantic diference
in the representation of quantification. Altogether,
our results suggest that the models do not appear
to contain the commonsense knowledge required
to interpret generics that difer in content through
quantifiers. However, they seem to have encoded a
meaning associated with the generic form, which leads
them to reshape the probability associated with various
quantifiers when the generic sentence is provided as
context. In the last experiment, we observed that the
models prefer the universal quantifier unanimously
if preceded by a generic utterance. People behave
similarly when tested on novel categories, that is,
non-word categories for which subjects have no prior
understanding. However, people can modulate their
interpretations of generalizations in a real language
setting through their world knowledge of real categories.
Regardless, MLMs tend to treat real and invented
categories equally, being agnostic to world knowledge.
For this reason, this could be a potentially harmful bias.</p>
      </sec>
      <sec id="sec-4-2">
        <title>The presented analysis has theoretical and methodologi</title>
        <p>cal implications. First, we observed that the language of
generalization is a complex phenomenon that is hard
to investigate in human processing and even more in
LLMs, mostly because the investigation of generics’
interpretation makes use of quantifiers, and language</p>
      </sec>
      <sec id="sec-4-3">
        <title>6The authors reproduced the experiment of Cimpian et al. [7], ob</title>
        <p>taining the same results.
models often fail in tasks related to quantification.
Another problem lies in the fact that it is dificult to
test autoregressive models (e.g., GPT family) on tasks
such as the one used in Experiment 1 because, as they
do not have access to the right context, they do not
have suficient information to modulate the probabilities
associated with the various quantifiers accordingly.
Finding ways to test autoregressive models in addition
to MLMs would be desirable.</p>
        <p>In this paper, we have not directly investigated
the aspect of the attention in MLMs. Future research
could address this aspect. Furthermore, future work
could involve the definition of alternative tasks for
investigating generalizations to make comparing models
and human interpretations easier. Psycholinguistic tests
on this phenomenon often rely on truth judgments,
and we should be cautious about comparing human
truth judgments with model outputs since they lack
commonsense knowledge comparable to that of humans.
Overall, further investigations are needed to clarify the
interpretation of generics in language models.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <sec id="sec-5-1">
        <title>This research is funded by the European Union (GRANT</title>
        <p>AGREEMENT: ERC-2021-STG-101039777). Views and
opinions expressed are however those of the author(s)
only and do not necessarily reflect those of the
European Union or the European Research Council Executive
Agency. Neither the European Union nor the granting
authority can be held responsible for them. We thank the
anonymous reviewers for their useful suggestions.
fect, Journal of Memory and Language 65 (2011) ence society. Austin, TX: Cognitive Science Society,
15–31. 2009.
[7] A. Cimpian, A. C. Brandone, S. A. Gelman, Generic [19] G. N. Carlson, Reference to kinds in English.,
Unistatements require little evidence for acceptance versity of Massachusetts Amherst, 1977.
but have powerful implications, Cognitive science [20] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT:
34 (2010) 1452–1482. Pre-training of Deep Bidirectional Transformers
[8] S. Khemlani, S.-J. Leslie, S. Glucksberg, Inferences for Language Understanding, in: Proceedings of
about members of kinds: The generics hypothesis, NAACL, 2019.</p>
        <p>Language and Cognitive Processes 27 (2012) 887– [21] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen,
900. O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov,
[9] S.-J. Leslie, S. A. Gelman, Quantified statements Roberta: A robustly optimized bert pretraining
apare recalled as generics: Evidence from preschool proach, arXiv preprint arXiv:1907.11692 (2019).
children and adults, Cognitive psychology 64 (2012) [22] M. Apidianaki, A. Garí Soler, ALL dolphins
186–214. are intelligent and SOME are friendly:
Prob[10] N. Reiter, A. Frank, Identifying generic noun ing BERT for nouns’ semantic properties and
phrases, in: Proceedings of the 48th annual meet- their prototypicality, in: Proceedings of the
ing of the association for computational linguistics, Fourth BlackboxNLP Workshop on Analyzing
2010, pp. 40–49. and Interpreting Neural Networks for NLP,
As[11] A. Friedrich, A. Palmer, M. P. Sørensen, M. Pinkal, sociation for Computational Linguistics, Punta
Annotating genericity: a survey, a scheme, and a Cana, Dominican Republic, 2021, pp. 79–94.
corpus, in: Proceedings of the 9th Linguistic Anno- URL: https://aclanthology.org/2021.blackboxnlp-1.
tation Workshop, 2015, pp. 21–30. 7. doi:10.18653/v1/2021.blackboxnlp-1.7.
[12] V. Govindarajan, B. V. Durme, A. S. White, Decom- [23] A. Gupta, Probing quantifier comprehension
posing generalization: Models of generic, habitual, in large language models, arXiv preprint
and episodic statements, Transactions of the As- arXiv:2306.07384 (2023).
sociation for Computational Linguistics 7 (2019) [24] A. Kilgarrif, Kovář; v.; rychlý, p.; suchomel, v. the
501–517. tenten corpus family, in: 7th International Corpus
[13] S. Ralethe, J. Buys, Generic overgeneralization in Linguistics Conference CL, 2013.
pre-trained language models, in: Proceedings of the [25] V. Suchomel, Better web corpora for corpus
lin29th International Conference on Computational guistics and nlp, Doctoral Theses. Brno: Masaryk
Linguistics, International Committee on Compu- University, Faculty of Informatics (2020).
tational Linguistics, Gyeongju, Republic of Korea, [26] A. Rogers, O. Kovaleva, A. Rumshisky, A
2022, pp. 3187–3196. URL: https://aclanthology.org/ primer in BERTology: What we know about
2022.coling-1.282. how BERT works, Transactions of the
Associa[14] E. Allaway, J. D. Hwang, C. Bhagavatula, K. McK- tion for Computational Linguistics 8 (2020) 842–
eown, D. Downey, Y. Choi, Penguins don’t fly: 866. URL: https://aclanthology.org/2020.tacl-1.54.
Reasoning about generics through instantiations doi:10.1162/tacl_a_00349.
and exceptions, arXiv preprint arXiv:2205.11658 [27] I. Tenney, D. Das, E. Pavlick, BERT rediscovers
(2022). the classical NLP pipeline, in: Proceedings of
[15] C. Bhagavatula, J. D. Hwang, D. Downey, R. Le Bras, the 57th Annual Meeting of the Association for
X. Lu, L. Qin, K. Sakaguchi, S. Swayamdipta, P. West, Computational Linguistics, Association for
ComY. Choi, I2D2: Inductive knowledge distillation putational Linguistics, Florence, Italy, 2019, pp.
with NeuroLogic and self-imitation, in: Proceed- 4593–4601. URL: https://aclanthology.org/P19-1452.
ings of the 61st Annual Meeting of the Association doi:10.18653/v1/P19-1452.
for Computational Linguistics (Volume 1: Long [28] W. Timkey, M. van Schijndel, All bark and no bite:
Papers), Association for Computational Linguis- Rogue dimensions in transformer language models
tics, Toronto, Canada, 2023, pp. 9614–9630. URL: obscure representational quality, in: Proceedings
https://aclanthology.org/2023.acl-long.535. of the 2021 Conference on Empirical Methods in
[16] S.-J. Leslie, Generics and the structure of the mind, Natural Language Processing, Association for
Com</p>
        <p>Philosophical perspectives 21 (2007) 375–403. putational Linguistics, Online and Punta Cana,
Do[17] S.-J. Leslie, Generics: Cognition and acquisition, minican Republic, 2021, pp. 4527–4546. URL: https:</p>
        <p>Philosophical Review 117 (2008) 1–47. //aclanthology.org/2021.emnlp-main.372. doi:10.
[18] S. Khemlani, S.-J. Leslie, S. Glucksberg, Generics, 18653/v1/2021.emnlp-main.372.
prevalence, and default inferences, in: Proceedings [29] K. Ethayarajh, How contextual are contextualized
of the 31st annual conference of the cognitive sci- word representations? Comparing the geometry</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>A. Quantifiers Frequencies in enTenTen21</title>
      <p>We extracted the frequencies from enTenTen21. The
corpus is made up of texts collected from the Internet
consisting of more than 60 billion tokens. The texts were
downloaded in October–December 2021 and January 2022. We
relied on the concordance tool provided by SketchEngine
to extract the frequencies in the form ‘[quantifier][noun]’.
some
all
many
few
most</p>
      <p>N. hits</p>
    </sec>
    <sec id="sec-7">
      <title>B. Experiment 1: Boxplots and</title>
    </sec>
    <sec id="sec-8">
      <title>Wilcoxon statistical analysis</title>
      <sec id="sec-8-1">
        <title>We report the boxplots for the base and large versions of BERT and RoBERTa for Experiment 1.</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>C. Experiment 2: A layer-wise analysis of MLMs representations</title>
      <sec id="sec-9-1">
        <title>We report the plots for the base and large variants of BERT and RoBERTa, with respect to each hidden layer.</title>
      </sec>
      <sec id="sec-9-2">
        <title>We also report the same plots with the correlation</title>
        <p>scores for each sentence category. While the trends are
the same for the three conditions, the values have a slight
diference in the means, with quasi-definitional sentences
having a higher correlation than the other two types.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>D. Experiment 3: Boxplots and</title>
    </sec>
    <sec id="sec-11">
      <title>Wilcoxon statistical analysis</title>
      <sec id="sec-11-1">
        <title>We report the boxplots for the base and large versions of BERT and RoBERTa for experiment 3.</title>
        <p>Model
Group1</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Hampton</surname>
          </string-name>
          ,
          <article-title>Generics as reflecting conceptual knowledge</article-title>
          , Recherches linguistiques de Vincennes (
          <year>2012</year>
          )
          <fpage>9</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.-J.</given-names>
            <surname>Leslie</surname>
          </string-name>
          ,
          <article-title>Carving up the social world with generics, Oxford studies in experimental philosophy 1 (</article-title>
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Chatzigoga</surname>
          </string-name>
          , Genericity,
          <source>in: The Oxford Handbook of Experimental Semantics and Pragmatics</source>
          , Oxford University Press,
          <year>2019</year>
          , pp.
          <fpage>156</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Krifka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Pelletier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Carlson</surname>
          </string-name>
          , A. ter Meulen, G. Chierchia, G. Link,
          <string-name>
            <surname>Genericity:</surname>
          </string-name>
          <article-title>An introduction</article-title>
          , in: G. N.
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>F. J.</given-names>
          </string-name>
          <string-name>
            <surname>Pelletier</surname>
          </string-name>
          (Eds.),
          <source>The Generic Book</source>
          , University of Chicago Press,
          <year>1995</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Tessler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. D.</given-names>
            <surname>Goodman</surname>
          </string-name>
          ,
          <article-title>The language of generalization</article-title>
          .,
          <source>Psychological review 126</source>
          (
          <year>2019</year>
          )
          <fpage>395</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.-J.</given-names>
            <surname>Leslie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khemlani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Glucksberg</surname>
          </string-name>
          ,
          <article-title>Do all ducks lay eggs? the generic overgeneralization efof BERT, ELMo, and GPT-2 embeddings</article-title>
          , in
          <source>: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>55</fpage>
          -
          <lpage>65</lpage>
          . URL: https://aclanthology.org/D19-1006. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1006.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Marchak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bianchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Gelman</surname>
          </string-name>
          ,
          <article-title>Generic language for social and animal kinds: An examination of the asymmetry between acceptance and inferences</article-title>
          ,
          <source>Cognitive Science 46</source>
          (
          <year>2022</year>
          )
          <article-title>e13209</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>