<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Assessing Language and Vision-Language Models on Event Plausibility</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maria Cassese</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Bondielli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Lenci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CoLing Lab, Department of Philology</institution>
          ,
          <addr-line>Literature, and Linguistics</addr-line>
          ,
          <institution>University of Pisa</institution>
          ,
          <addr-line>36 S. Maria St, Pisa, I-56126</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, University of Pisa</institution>
          ,
          <addr-line>3 Largo Bruno Pontecorvo, Pisa, 56127</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Transformer-based Language Models (LMs) excel in many tasks, but they appear to lack robustness in capturing crucial aspects of event knowledge due to their reliance on surface-level linguistic features and the mismatch between language descriptions and real-world occurrences. In this paper, we investigate the potential of Transformer-based Vision-Language Models (VLMs) in comprehending Generalized Event Knowledge (GEK), aiming to determine whether the inclusion of a visual component afects the mastery of GEK. To do so, we compare multimodal Transformer models with unimodal ones on a task evaluating the plausibility of curated minimal sentence pairs. We show that current VLMs generally perform worse than their unimodal counterparts, suggesting that VL pre-training strategies are not yet as efective to model semantic understanding and resulting models are more akin to bag-of-words in this context.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;multimodal semantics</kwd>
        <kwd>vision language models</kwd>
        <kwd>language models</kwd>
        <kwd>generalized event knowledge</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>CLiC-it 2023: 9th Italian Conference on Computational Linguistics,
Nov 30 — Dec 02, 2023, Venice, Italy
$ m.cassese4@studenti.unipi.it (M. Cassese);
alessandro.bondielli@unipi.it (A. Bondielli);
alessandro.lenci@unipi.it (A. Lenci)</p>
      <p>0009-0007-4765-1221 (M. Cassese); 0000-0003-3426-6643
(A. Bondielli); 0000-0001-5790-4308 (A. Lenci)</p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org)</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>The introduction of multimodal models in NLP stems
from the intrinsic limitations of computational models
that are trained exclusively on distributional statistics
extracted from textual data [9, 3, 10]. In fact, they lack
referential competence [11], which prevents them from
grounding linguistic structures onto real-world
experiences [12, 13].</p>
      <p>The earliest multimodal distributional models already
showed the ability to improve the semantic
representation of concrete concepts and properties [14], as well as
abstract verbs that lack direct perceptual information,
but benefit from integrating linguistic inputs and
perceptual information [15]. However, they proved to be less
efective in representing verbs, adjectives, and abstract
concepts [16].</p>
      <p>The introduction of Transformer-based VLMs such
as Visual-BERT [17] and FLAVA [18], and efective
techniques for Vision-Language Pre-training [19] paved the
way for new research in a multimodal setting. While
numerous studies has shown VLMs success on diferent
multimodal tasks, less efort has been put in analyzing
their diferences with unimodal counterparts on natural
language understanding (NLU). [20] show that both
dualstream and single-stream VLMs are equally capable of
preserving NLU capabilities. The analysis conducted by
[6] shows that multimodal models do not significantly
outperform the text-only variants in a language-only
setting. This was attributed to the use of narrow domain
data and direct extensions of NLP architectures. Our
work support these findings by focusing on
understanding the plausibility of events.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments</title>
      <p>Our goal is to evaluate the ability of LMs and VLMs to
predict the semantic plausibility of sentences with respect
to human judgements. In the following, we describe the
data used in the experiments (Sec. 3.1), the models we
considered in the evaluation (Sec. 3.2), and detail the
evaluation procedure itself (Sec. 3.3).
3.1. Data
We sourced our data from a number of existing datasets
containing pairs of sentences describing transitive event
distinguished by patient plausibility within the context
of the sentence. Plausibility is rated by humans and
expressed on a 1-7 Likert scale. Formally, each data point
consists of a plausible sentence  and its corresponding
implausible one  obtained through a modification of
.</p>
      <sec id="sec-3-1">
        <title>We considered the following datasets:</title>
      </sec>
      <sec id="sec-3-2">
        <title>DTFit [21]. It includes past tense sentences distin</title>
        <p>guished by patient prototipicality. For each
plausible sentence, the implausible (i.e., atypical) one
is obtained by replacing the patient with an
atypical filler for that role (e.g., The actor won the award
vs The actor won the battle).</p>
      </sec>
      <sec id="sec-3-3">
        <title>EventsAdapt [22]. It includes pairs of plausibleimplausible sentences where the implausible one is obtained by reversing the noun phrases (e.g.,</title>
        <p>The cop arrested the criminal vs. The criminal
arrested the cop). The dataset is divided into two
sub-datasets. In the former, henceforth referred
to as −  , both the agent and
the patient are animate. In the latter, denoted as
−  , the agent of the original
sentence is animate, while the patient is not. Thus
the implausible sentence is also semantically
impossible.</p>
      </sec>
      <sec id="sec-3-4">
        <title>EventsRev [23]. It includes concrete sentences describ</title>
        <p>ing events in the present progressive tense. Like
EventsAdapt, implausible sentences are obtained
by reversing the noun phrases, which in this case
always depict animate entities (e.g., The cat is
chasing the mouse vs. The mouse is chasing the
cat). Each sentence, both plausible and
implausible, is accompanied by an image depicting the
interaction between the two animated participants
described in the sentence. The images are simple
black and white drawings.</p>
      </sec>
      <sec id="sec-3-5">
        <title>As we are interested in considering also the efect of</title>
        <p>concreteness in VLMs’ ability to recognize plausibility,
we further grouped sentences of EventsAdapt (and its
subgroups) and DTFit into concrete and abstract ones.
We categorized the sentences based on the level of
concreteness of the verb, subject, and object in each sentence.
We chose to consider sentences that refer to abstract
concepts with high imageability as concrete (e.g., The priest
celebrated the marriage).</p>
        <p>To the best of our knowledge, none of the data used
in this study was included in the training set of the
evaluated models.</p>
        <sec id="sec-3-5-1">
          <title>3.2. Models</title>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>We test various popular multimodal VLMs and compare them with baseline unimodal LMs: BERT [24] and RoBERTa [25]. As for the VLMs, our analysis includes:</title>
      </sec>
      <sec id="sec-3-7">
        <title>VisualBERT [17]. A single-stream early fusion en</title>
        <p>
          coder model initialized from pre-trained
BERTbase weights and further trained on multimodal
datasets. Visual features are extracted from a
pretrained Faster R-CNN network [
          <xref ref-type="bibr" rid="ref4">26</xref>
          ] and fed into
the transformer model alongside the text.
        </p>
      </sec>
      <sec id="sec-3-8">
        <title>ViLT [28]. A single-stream model employing a BERT</title>
        <p>model for textual feature extraction and a ViT
model for visual feature extraction, respectively.</p>
        <p>Resulting representations are then concatenated
and fed into the final model.</p>
        <p>
          LXMERT [
          <xref ref-type="bibr" rid="ref5">27</xref>
          ]. A dual-stream early fusion encoder for plausible and implausible sentences, indicating that
model including some modality-specific lay- the model is less able to distinguish between them. Thus,
ers and allowing cross-attention in specific co- negative correlation values indicate good performances.
attention layers. Visual features are extracted We also analyzed the density of the distributions for PLLs.
with a Faster R-CNN network. This is essential to comprehend how humans and models
diferentiate between plausible and implausible classes,
aiding in evaluating sentence complexity and comparing
model behavior to humans’.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Discussion</title>
      <p>FLAVA [18]. A foundation VLM including an image We first verify the performances of VLMs in plausibility
encoder, a textual encoder, and a multimodal en- recognition via accuracy. Results of all models on the
coder. It is jointly pre-trained on both unimodal datasets are reported in Table 1.
and multimodal data, thus learning high-quality Both LMs and VLMs show significantly higher
pervisual and textual representations. It is capable formances on −  , where
implauof achieving both crossmodal alignment and mul- sible sentences describe impossible events, than on
timodal fusion objectives. −  where implausible sentences
deTo adapt multimodal models to the text-only task, we pict unlikely but not impossible events. On AN-AN
sensimply modified the inputs, e.g. by feeding them empty tences, BERT, RoBERTa, VisualBERT, and FLAVA
perimage tensors. FLAVA does not require to be adapted to formed above chance levels, while ViLT and LXMERT
text-only inputs, as it can directly be evaluated by using performed at chance. This indicates that extracting
inonly the textual encoder. formation about AN-AN sentence plausibility is
gener</p>
      <p>
        All models and their pre-trained weights are available ally challenging, and more so for VLMs. Among VLMs,
on Huggingface Transformers [
        <xref ref-type="bibr" rid="ref7">29</xref>
        ].1 FLAVA performs best, with results generally close to
RoBERTa.
      </p>
      <p>Going further, we consider EventsAdapt and we
pro3.3. Evaluation procedure vide the density plot of PLLs divided by plausibility for
To evaluate the ability of a LM (or VLM) to distinguish each model and human raters in Figure 2, and plot the
between plausible and implausible sentences, we first correlation between PLLs of plausible and implausible
have to compute a plausibility score for each sentence. sentences in Figure 1.</p>
      <p>
        Since we are dealing with bi-directional masked language Both LMs and VLMs do not clearly distinguish between
models, we can approximate this plausibility score via the two classes and exhibit very similar distributions for
pseudo-log-likelihood (PLL), defined as the sum of loga- plausible and implausible sentences. The complexity of
rithmic probabilities of each token based on the remain- the task afect the results as well: for tasks where
huing tokens in the sentence [
        <xref ref-type="bibr" rid="ref8">30</xref>
        ]. To avoid bias favoring mans have no dificulty in distinguishing between the
multi-token words, we apply an additional mask that cov- two classes, as the implausible sentence violates the verb
ers tokens to the right of the target, as proposed in [4]. To selection preferences (AN-IN), the models can better
idencompare PLL scores with human judgements expressed tify patterns that diferentiate the two sentences (Fig 1a);
on a Likert scale, we normalized both using a min-max for tasks where even humans are more uncertain (AN-AN),
scaler function. the models tend to assign very similar scores to the two
      </p>
      <p>First, we evaluated the models using an accuracy met- sentences (Fig. 1b). This is also clearly shown by the
denric. Specifically, considering all (, ) sentence pairs sity distribution plot in Figure 2. for the AN-IN case, the
for a dataset, we computed accuracy as the percentage density distribution for humans show a clear separation,
of cases for which  () &gt;  (). while models show more modest but still evident signs of</p>
      <p>To provide a more detailed analysis of the perfor- separation. The density distribution for AN-AN sentences
mances, we further evaluate the models via distribution shows a less separated distribution for human scores and
analyses. We used the Pearson correlation coeficient almost entirely overlapped distributions for models. One
between each model’s score for the plausible and implau- possible reason for this is that the grammaticality of a
sible sentences. More in detail, for each pair of (, ), sentence depends on syntactic rules that can be more
we plot the correlation between normalized  () easily detected through statistical inference. In contrast,
and  (). High correlation implies similar scores linguistic acceptability may depend on extralinguistic
information requiring multiple inference levels.</p>
      <p>
        The role of concreteness. The experiments show that is higher for VLMs than for LMs (0.06 for LMs, 0.09
VLMs do not show improved abilities to deal with GEK for VLMs); when considering the simpler sentences of
and event plausibility with respect to textual LMs. How- −  , the diferences are less marked.
ever, we could expect that this might also depend on On the other hand, multimodal models demonstrate
excelthe event concreteness, as concrete concepts are more lent recognition of abstract events in the DTFit dataset.
directly grounded on visual information than abstract Note however that abstract sentences are an order of
ones. The concreteness of an event depends on the the magnitude less than concrete ones in the dataset.
predicate itself, as well on its arguments. For instance, We also show a comparison of Pearson correlation
the verb to fight has a concrete use in the sentence The scores of results between − and
wrestler fought the opponent and an abstract use in The   , shown respectively in Figures 3a and 3b.
patient fought the cancer. Cognitive research has shown While VLMs exhibit high correlation values, i.e. less
that abstract concepts require more linguistic experiences prowess on the task, values for DTFit are generally
to be understood [
        <xref ref-type="bibr" rid="ref9">31</xref>
        ]. Thus they are generally more dif- lower, suggesting a better ability to assess
plausibilifcult to acquire and process. This is influenced by two ity. VLMs’ performance diference in the two datasets
main factors, namely imageability and familiarity [
        <xref ref-type="bibr" rid="ref10">32</xref>
        ]. may be due to how implausible sentences are generated.
For instance, the abstract verb to celebrate becomes more EventsAdapt uses noun phrase order reversal, while
DTconcrete in the context of The priest celebrated the wed- Fit only replaces the typical patient with an incompatible
ding because it is easier to form mental images of the one. If VLMs behave more like bag-of-words models,
event and it is very frequent in language use. they may struggle to recognize semantic diferences
be
      </p>
      <p>To evaluate how concreteness afects the models’ abil- tween sentences with the same words but diferent
orities, we first compute the accuracy on concrete and ab- der. This would explain their worse performances on
stract subsets of the dataset. Results are reported in Ta- −  .
ble 2. Multimodal models seem to perform worse on
abstract sentences with a higher degree of complexity: The impact of images Finally, we analyze whether
on the −  dataset, the average per- including images of the (im)plausible test events in the
formance gap between abstract and concrete sentences input is beneficial for VLMs. We provide accuracy scores</p>
      <sec id="sec-4-1">
        <title>Several interesting findings have emerged from our anal</title>
        <p>ysis. First, VLMs do not achieve significantly higher
accuracy values than unimodal ones in a semantic plausibility
recognition task. Second, we saw that performances of
VLMs is worse when dealing with more challenging
senfor VLMs on the EventsRev dataset in Table 3. Including tences represented by −  , exhibiting
event images does not lead to any improvement: perfor- lower accuracy and a high correlation between plausible
mances either remain the same or slightly degrade. and implausible sentences. Third, we saw that including
images of events in the input does not lead to improved</p>
        <sec id="sec-4-1-1">
          <title>Dataset VisualBERT LXMERT ViLT FLAVmAodel performances.</title>
          <p>0.76 0.66 0.76 0.79 We discuss a possible interpretation of these findings
+ 0.61 0.66 0.71 0.79in the following. First, the generally high correlation
between PLL scores for pairs of (, ) for VLMs suggest
Table 3 that these models struggle to recognize semantic
diferAccuracy of VLMs on EventsRev with ( + ) and without () ences, especially between sentences with diferent word
images in the input. orders (e.g., with subject-patient inversion), and
relationships between sentence components, like semantic roles.</p>
          <p>
            This may be further indication that VLMs model
lan4.1. Discussion guage in a bag-of-words fashion [7, 8]. The pre-training
method used in masked language modelling for VLMs,
adding visual features to language models already
specialized on linguistic tasks, may also compromise learning
as suggested by [
            <xref ref-type="bibr" rid="ref11">33</xref>
            ]. The high-dimensional space learnt
by these models could make it dificult to identify
se(a) 
−  .
          </p>
          <p>
            (b)   .
mantic errors. Moreover, models using pre-trained LM still behave similarly to Bag-of-Words models, regardless
weights for text processing may face limitations in the of the degree of concreteness of the events.
type of visual information they can capture during train- In the future, we plan to focus on the analysis of models
ing. Some models rely on object categories trained on with visual grounding as their training objective, such
bounding boxes. This is computationally expensive, and as PaLM-E [
            <xref ref-type="bibr" rid="ref12">34</xref>
            ], a large embodied multimodal language
the learned representations may not adequately capture model that directly incorporates real-world continuous
shapes and relationships. Other models, such as ViLT, sensor modalities into language processing. This may
that leverage ViT representations and use a linear func- shed more light into the abilities of large multimodal
tion to extract embeddings for image patches, are less models to achieve more human-level grounded language
costly but may result in lower-quality representations. understanding.
          </p>
          <p>
            These results are in line with [
            <xref ref-type="bibr" rid="ref11">33</xref>
            ].
          </p>
          <p>A possible explanation of why VLMs do not benefit
from including test images is that in this specific case Acknowledgments
(minimal sentence pairs with subject-object inversion)
the images for both sentences are very similar, and difer
only for the relationship between the entities. The visual
encoders of the models might be too weak to diferentiate
substantially similar images, leading the models to rely
on their LM priors and make random choices. Finally, we
saw that even the foundation Large VLM we considered –
FLAVA – does not show significantly improved accuracy
compared to other VLMs.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Research partially supported by the Italian Ministry of</title>
        <p>University and Research (MUR) in the framework of the
PON 2014-2021 “Research and Innovation" resources –
Innovation Action - DM MUR 1062/2021 - Title of the
Research: “Modelli semantici multimodali per l’industria
4.0 e le digital humanities.”, and by PNRR - M4C2 -
Investimento 1.3, Partenariato Esteso PE00000013 - “FAIR
Future Artificial Intelligence Research" - Spoke 1
“Humancentered AI", funded by the European Commission under
the NextGeneration EU programme.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Future Works</title>
      <sec id="sec-5-1">
        <title>In this paper, we presented a set of experiments aimed at</title>
        <p>evaluating the ability of VLMs to model event plausibility
in both language-only and vision-language tasks against
LMs. We find that VL pre-training does not lead to a
significant improvement compared to unimodal LMs in
this task aiming at testing their GEK. Specifically, we
observed that VLMs tend to perform worse when the
implausible sentence has a higher semantic complexity,
because it contains two animate nouns. Our analysis also
brings further support to argument that VLMs models
[3] P. Pedinotti, G. Rambelli, E. Chersoni, E. Santus, beddings from multi-modal data: Since you
probaA. Lenci, P. Blache, Did the cat drink the cofee? bly can’t see what I mean, in: Proceedings of the
challenging transformers with generalized event 2014 Conference on Empirical Methods in
Natuknowledge, arXiv preprint arXiv:2107.10922 (2021). ral Language Processing (EMNLP), Association for
[4] C. Kauf, A. A. Ivanova, G. Rambelli, E. Chersoni, Computational Linguistics, Doha, Qatar, 2014, pp.</p>
        <p>J. S. She, Z. Chowdhury, E. Fedorenko, A. Lenci, 255–265. doi:10.3115/v1/D14-1032.</p>
        <p>Event knowledge in large language models: the [16] R. Shekhar, S. Pezzelle, Y. Klimovich, A. Herbelot,
gap between the impossible and the unlikely, 2022. M. Nabi, E. Sangineto, R. Bernardi, FOIL it! find
doi:10.48550/ARXIV.2212.01488. one mismatch between image and language caption,
[5] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, in: Proceedings of the 55th Annual Meeting of the
L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, At- Association for Computational Linguistics (Volume
tention is all you need, in: Proceedings of the 31st 1: Long Papers), Association for Computational
International Conference on Neural Information Linguistics, Vancouver, Canada, 2017, pp. 255–265.
Processing Systems, NIPS’17, Curran Associates doi:10.18653/v1/P17-1024.</p>
        <p>Inc., Red Hook, NY, USA, 2017, p. 6000–6010. [17] L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, K.-W. Chang,
[6] T. Yun, C. Sun, E. Pavlick, Does vision-and- Visualbert: A simple and performant baseline for
vilanguage pretraining improve lexical grounding?, sion and language, arXiv preprint arXiv:1908.03557
in: Findings of the Association for Computational (2019).</p>
        <p>Linguistics: EMNLP 2021, Association for Computa- [18] A. Singh, R. Hu, V. Goswami, G. Couairon,
tional Linguistics, Punta Cana, Dominican Republic, W. Galuba, M. Rohrbach, D. Kiela, Flava: A
founda2021, pp. 4357–4366. doi:10.18653/v1/2021.f tional language and vision alignment model, 2022
indings-emnlp.370. IEEE/CVF Conference on Computer Vision and
Pat[7] S. Castro, O. Ignat, R. Mihalcea, Scalable perfor- tern Recognition (CVPR) (2021) 15617–15629.
mance analysis for vision-language models, in: [19] Z. Gan, L. Li, C. Li, L. Wang, Z. Liu, J. Gao, et al.,
Proceedings of the The 12th Joint Conference on Vision-language pre-training: Basics, recent
adLexical and Computational Semantics (*SEM 2023), vances, and future trends, Foundations and Trends®
Association for Computational Linguistics, Toronto, in Computer Graphics and Vision 14 (2022) 163–
Canada, 2023, pp. 284–294. 352.
[8] T. Thrush, R. Jiang, M. Bartolo, A. Singh, [20] T. Iki, A. Aizawa, Efect of visual extensions on
natuA. Williams, D. Kiela, C. Ross, Winoground: Prob- ral language understanding in vision-and-language
ing vision and language models for visio-linguistic models, in: Proceedings of the 2021 Conference
compositionality, 2022. arXiv:2204.03162. on Empirical Methods in Natural Language
Pro[9] J. Gordon, B. Van Durme, Reporting bias and cessing, Association for Computational Linguistics,
knowledge acquisition, in: Proceedings of the 2013 Online and Punta Cana, Dominican Republic, 2021,
Workshop on Automated Knowledge Base Con- pp. 2189–2196. doi:10.18653/v1/2021.emnlp
struction, AKBC ’13, Association for Computing -main.167.</p>
        <p>Machinery, New York, NY, USA, 2013, p. 25–30. [21] P. Vassallo, E. Chersoni, E. Santus, A. Lenci,
doi:10.1145/2509558.2509563. P. Blache, Event Knowledge in Sentence
Process[10] A. Lenci, M. Sahlgren, Distributional Semantics, ing: A New Dataset for the Evaluation of Argument</p>
        <p>Cambridge University Press, Cambridge, 2023. Typicality, in: LREC 2018 Workshop on Linguistic
[11] D. Marconi, Lexical Competence, The MIT Press, and Neurocognitive Resources (LiNCR), Miyazaki,</p>
        <p>Cambridge, MA, 1997. Japan, 2018.
[12] S. Harnad, The symbol grounding problem, Physica [22] E. Fedorenko, I. A. Blank, M. Siegelman, Z. Minerof,
D: Nonlinear Phenomena 42 (1990) 335–346. doi:ht Lack of selectivity for syntax relative to word
meantps://doi.org/10.1016/0167-2789(90)9 ings throughout the language network, Cognition
0087-6. 203 (2020) 104348. doi:https://doi.org/10.1
[13] E. M. Bender, A. Koller, Climbing towards nlu: On 016/j.cognition.2020.104348.</p>
        <p>meaning, form, and understanding in the age of [23] A. A. Ivanova, Z. Minerof, V. Zimmerer, N.
Kandata, in: Proc. ACL, Seattle, WA, 2020, pp. 5185– wisher, R. Varley, E. Fedorenko, The Language
5198. Network Is Recruited but Not Required for
Nonver[14] E. Bruni, G. Boleda, M. Baroni, N.-K. Tran, Dis- bal Event Semantics, Neurobiology of Language
tributional semantics in technicolor, Proceedings 2 (2021) 176–201. doi:10.1162/nol_a_00030.
of the 50th Annual Meeting of the Association for arXiv:https://direct.mit.edu/nol/article-pdf/2/2/176/18
Computational Linguistics (2012) 136–145. [24] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT:
[15] F. Hill, A. Korhonen, Learning abstract concept em- Pre-training of deep bidirectional transformers for</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>McRae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Matsuki</surname>
          </string-name>
          ,
          <article-title>People use their knowledge of common events to understand language, and do so as quickly as possible</article-title>
          .,
          <source>Lang Linguist Compass</source>
          .
          <volume>3</volume>
          (
          <issue>6</issue>
          ) (
          <year>2009</year>
          )
          <fpage>1417</fpage>
          -
          <lpage>1429</lpage>
          . doi:
          <volume>10</volume>
          .1111/j.1749-818
          <string-name>
            <surname>X.</surname>
          </string-name>
          <year>2009</year>
          .
          <volume>00174</volume>
          .x.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Baltrušaitis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ahuja</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.-P. Morency,</surname>
          </string-name>
          <article-title>Multimodal machine learning: A survey and taxonomy</article-title>
          ,
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>41</volume>
          (
          <year>2018</year>
          )
          <fpage>423</fpage>
          -
          <lpage>443</lpage>
          . language understanding,
          <source>in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <source>Association for Computational Linguistics</source>
          , Minneapolis, Minnesota,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          -1423.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          ,
          <year>2019</year>
          . Cite arxiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <surname>Faster</surname>
          </string-name>
          r-cnn:
          <article-title>Towards real-time object detection with region proposal networks</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>28</volume>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bansal</surname>
          </string-name>
          ,
          <article-title>Lxmert: Learning crossmodality encoder representations from transformers</article-title>
          ., in: K. Inui,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          Wan (Eds.),
          <source>EMNLP/IJCNLP (1)</source>
          ,
          <source>Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>5099</fpage>
          -
          <lpage>5110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Son</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kim</surname>
          </string-name>
          ,
          <article-title>Vilt: Vision-and-language transformer without convolution or region supervision</article-title>
          ,
          <source>in: International Conference on Machine Learning</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Debut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chaumond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delangue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cistac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Louf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Funtowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Davison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shleifer</surname>
          </string-name>
          , P. von Platen, C. Ma,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jernite</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Plu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Le</given-names>
            <surname>Scao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gugger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Drame</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Lhoest</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rush</surname>
          </string-name>
          , Transformers:
          <article-title>Stateof-the-art natural language processing</article-title>
          ,
          <source>in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>38</fpage>
          -
          <lpage>45</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .emnlp-demos.
          <volume>6</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>J.</given-names>
            <surname>Salazar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kirchhof</surname>
          </string-name>
          ,
          <article-title>Masked language model scoring, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>2699</fpage>
          -
          <lpage>2712</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>240</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>240</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>G.</given-names>
            <surname>Vigliocco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Meteyard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Andrews</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Kousta</surname>
          </string-name>
          ,
          <article-title>Toward a theory of semantic representation</article-title>
          ,
          <source>Language and Cognition</source>
          <volume>1</volume>
          (
          <year>2009</year>
          )
          <fpage>219</fpage>
          -
          <lpage>247</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>G.</given-names>
            <surname>Löhr</surname>
          </string-name>
          ,
          <article-title>Does the mind care about whether a word is abstract or concrete? why concreteness is probably not a natural kind, Mind &amp; Language n/a (</article-title>
          <year>2023</year>
          ). doi:https://doi.org/10.1111/mila.12473.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>C.</given-names>
            <surname>Fields</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kennington</surname>
          </string-name>
          , Vision language transformers: A survey,
          <year>2023</year>
          . arXiv:
          <volume>2307</volume>
          .
          <fpage>03254</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>D.</given-names>
            <surname>Driess</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S. M.</given-names>
            <surname>Sajjadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lynch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chowdhery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ichter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wahid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tompson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Vuong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chebotar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sermanet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Duckworth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vanhoucke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hausman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Toussaint</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gref</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Mordatch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Florence</surname>
          </string-name>
          , Palm-e:
          <article-title>An embodied multimodal language model</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>03378</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>