<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CONCRETEXT @ EVALITA2020: The Concreteness in Context Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lorenzo Gregori</string-name>
          <email>lorenzo.gregori@unifi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Montefinese Daniele P. Radicioni</string-name>
          <email>daniele.radicioni@unito.it</email>
          <email>maria.montefinese@unipd.it</email>
          <email>maria.montefinese@unipd.it daniele.radicioni@unito.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Amelio Ravelli</string-name>
          <email>andreaamelio.ravelli@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rossella Varvara</string-name>
          <email>rossella.varvara@unifi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Istituto di Linguistica Computazionale, “Antonio Zampolli” (ILC-CNR) - ItaliaNLP Lab</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Florence</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Padua University of Turin</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>415</fpage>
      <lpage>428</lpage>
      <abstract>
        <p>Focus of the CONCRETEXT task is conceptual concreteness: systems were solicited to compute a value expressing to what extent target concepts are concrete (i.e., more or less perceptually salient) within a given context of occurrence. To these ends, we have developed a new dataset which was annotated with concreteness ratings and used as gold standard in the evaluation of systems. Four teams participated in this first edition of the task, with a total of 15 runs submitted. Interestingly, these works extend information on conceptual concreteness available in existing (non contextual) norms derived from human judgments with new knowledge from recently developed neural architectures, in much the same multidisciplinary spirit whereby the CONCRETEXT task was organized.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Concept concreteness – that is, how directly a
concept is related to sensorial experience
        <xref ref-type="bibr" rid="ref7 ref8">(Brysbaert
et al., 2014a)</xref>
        – is a fundamental dimension of
conceptual semantic representation that has attracted
more and more interest and attention in
psycholinguistics in the last decade. This dimension is
usually assessed by participants ratings on a Likert
scale: concrete concepts lie herein on one side of
the scale and refer to something that exists in
reality and can be experienced immediately through
      </p>
      <p>
        Copyright c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
the senses; abstract concepts lie on the opposite
side of the scale and are grounded in the
internal sensory experience and linguistic information.
While concrete concepts have direct sensory
referents
        <xref ref-type="bibr" rid="ref12">(Crutch and Warrington, 2005)</xref>
        and greater
availability of contextual information
        <xref ref-type="bibr" rid="ref11 ref19 ref24">(Connell et
al., 2018; Kousta et al., 2011; Montefinese et al.,
2020)</xref>
        , abstract concepts tend to be more
emotionally valenced
        <xref ref-type="bibr" rid="ref19">(Kousta et al., 2011)</xref>
        and less
imageable
        <xref ref-type="bibr" rid="ref15 ref24">(Montefinese et al., 2020; Garbarini et al.,
2020)</xref>
        .
      </p>
      <p>
        The CONCRETEXT task challenges
participants to build NLP systems to automatically
assign a concreteness value to words in context. It is
aimed at investigating how the concreteness
information affects sense selection: different from past
research
        <xref ref-type="bibr" rid="ref23 ref7 ref8">(Brysbaert et al., 2014b; Montefinese et
al., 2014)</xref>
        , we are interested in assessing the
concreteness of concepts within the context of real
sentences rather than in isolation. Additionally,
the concreteness score is assumed to be a property
of meanings rather than a property of word forms;
thus, scoring the concreteness of a concept in
context implicitly requires to individuate its
underlying sense, by handling lexical phenomena such as
polysemy and homonymy.
      </p>
      <p>
        Ordinary experience suggests that concepts’
concrete/abstract status can affect their semantic
representation, and lexical access and processing:
concrete meanings are acknowledged to be more
quickly and easily delivered in human
communication than abstract meanings
        <xref ref-type="bibr" rid="ref2">(Bambini et al.,
2014)</xref>
        . Historically, it has been observed that
concrete concepts are responded to more quickly than
abstract concepts in lexical decision tasks
        <xref ref-type="bibr" rid="ref20 ref4">(Bleasdale, 1987; Kroll and Merves, 1986)</xref>
        , although
more recent experiments have shown that abstract
concepts might have an advantage when other
variables have been accounted for
        <xref ref-type="bibr" rid="ref19">(Kousta et al.,
2011)</xref>
        . Concrete concepts are also easier to encode
and retrieve than abstract concepts
        <xref ref-type="bibr" rid="ref22 ref25 ref30">(Romani et al.,
2008; Miller and Roodenrys, 2009)</xref>
        , are easier to
make associations with
        <xref ref-type="bibr" rid="ref13">(de Groot, 1989)</xref>
        , and are
more thoroughly described in definition tasks
        <xref ref-type="bibr" rid="ref27">(Sadoski et al., 1997)</xref>
        . Moreover, it takes generally
less time to comprehend a concrete sentence than
an abstract one
        <xref ref-type="bibr" rid="ref16 ref28">(Haberlandt and Graesser, 1985;
Schwanenflugel and Shoben, 1983)</xref>
        . Thus, it has
been proposed that different organizational
principles govern semantic representations of concrete
and abstract concepts: concrete concepts are
predominantly organized by featural similarity
measures, and abstract concepts by associative
relations, co-occurrence patterns and syntactic
information
        <xref ref-type="bibr" rid="ref30">(Vigliocco et al., 2009)</xref>
        .
      </p>
      <p>
        All surveyed features make aspects ingrained in
the distinction between concreteness/abstractness
a stimulating and challenging field also for
computational linguistics. Among the earliest attempts
at grasping concreteness, we find works that
investigated on concreteness/abstractness
information in its interplay with metaphor identification
and figurative language more in general
        <xref ref-type="bibr" rid="ref29">(Turney et al., 2011)</xref>
        (and, more recently
        <xref ref-type="bibr" rid="ref10 ref21">(Mensa
et al., 2018b)</xref>
        ). Although concreteness
information is acknowledged to be central to, e.g.,
word-sense induction and compositionality
modeling
        <xref ref-type="bibr" rid="ref18">(Hill et al., 2013)</xref>
        , the contribution of
concreteness/abstractness to semantic representations
is not fully grasped and exploited in existing
approaches and resources, with the notable
exception of works aimed i) at learning multimodal
embeddings, and how abstract and concrete
representations can be acquired by multi-modal
models
        <xref ref-type="bibr" rid="ref17 ref2 ref23 ref8">(Hill and Korhonen, 2014)</xref>
        ; and ii) at exploring
in how far concreteness information is represented
in the distributional patterns in corpora
        <xref ref-type="bibr" rid="ref18">(Hill et
al., 2013)</xref>
        . Moreover, some approaches exist that
attempted to create lexical resources by also
employing common-sense information
        <xref ref-type="bibr" rid="ref10 ref10 ref21">(Mensa et al.,
2018a; Colla et al., 2018)</xref>
        .
      </p>
      <p>Characterizing tokens within sentences with
their concreteness requires integrating both
wordspecific and contextual information. In our view,
the CONCRETEXT Task entails dealing with a
relaxed form of word sense disambiguation; such
aspects were faced by our participants by devising
methods relying on both traditional
knowledgebased approaches, and more recent language
models and sequence-to-sequence models. Finally,
like in many real-world cases, the provided trial
data is rather scarce, in the order of hundred
sentences for the Italian language, and as many for
English. This aspect forced our participants to
face something similar to a ‘cold start’ problem.
We hope that this edition of the CONCRETEXT
task will be the first appointment in a series for
those who are interested in the issues posed by the
contextual conceptual concreteness to research on
natural language semantics.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Task Definition</title>
      <p>The task CONCRETEXT (so dubbed after
CONcreteness in conTEXT) focuses on automatic
concreteness (and conversely, abstractness)
recognition. Given a sentence along with a target word,
we asked participants to propose a system able
to assess the concreteness of a concept expressed
by a given word within a sentence, on a 7-point
Likert-like scale where 1 stands for completely
abstract (e.g., ‘freedom’) and 7 for completely
concrete (e.g., ‘car’). For example, in the sentence
“In summer, wheat fields are coloured in yellow”
the noun field refers to an entity that can smell, be
touched, and pointed to. In this case, in a scale
ranging from 1 to 7 its concreteness may be
evaluated as 7, because it refers to an extremely
concrete concept. In contrast, the same noun field
in the sentence “Physics is Alice’s research field”
refers to a scientific subject, i.e., something that
cannot be perceived through the five senses, but
that can be explained through a linguistic
description. In this sentence, the noun field may be
evaluated 1 because it refers to an extremely abstract
concept. Moreover, the task targets can be halfway
between completely abstract and completely
concrete, as in the case of “Magnetic field attracts
iron”, where the noun field refers to something
more abstract compared to “wheat fields” but more
concrete compared to “research field”. As
anticipated, the concreteness score being assigned to
the word should be evaluated in context: the word
should not be considered in isolation, but as part
of a given sentence.</p>
      <p>
        Participants were invited to exploit all possible
strategies to solve the task, including (but not
limited to) knowledge bases, external training data,
word embeddings, etc.
The dataset used for this task has been taken from
the English-Italian parallel section of The Human
Instruction Dataset
        <xref ref-type="bibr" rid="ref11 ref21 ref9">(Chocron and Pareti, 2018)</xref>
        ,
derived from WikiHow instructions.1 All such
documents had been anonymized beforehand, so
that downloaded data present no privacy nor data
sensitivity issues.
      </p>
      <p>The dataset is composed of overall 1; 096
sentences, arranged as follows: 562 Italian sentences
plus 534 English sentences. Each sentence
contains a target term (either verb or noun) with its
associated concreteness score (1–7 scale). Such
score is derived from the average of at least 30
human judgments from native Italian and English
speakers about the concreteness of a target word in
a given sentence (see Table 1 for the dataset
numbers).</p>
      <p>The reliability of the collected data within
each language (Italian, English) for the trial and
test phases was evaluated separately by
applying the split-half correlations corrected with the
Spearman-Brown formula after randomly
dividing the participants into two subgroups of equal
size. All the reliability indexes were calculated
on 10; 000 different randomizations of the
participants. The mean correlations between the two
groups are very high for both the trial and test
phases, ranging from a minimum of r = 0:87
for English (at the test phase) to a maximum of
r = 0:98 for Italian (at the trial phase), showing
that the resulting ratings are highly reliable and
1The whole Human Instruction
dataset is freely available on
https://www.kaggle.com/paolop/
human-instructions-multilingual-wikihow
Dataset
Kaggle,
(a) English dataset.</p>
      <p>(b) Italian dataset.
can be used across the entire Italian – and English
– speaking populations.</p>
      <p>The dataset has been split into trial and test data,
with a 20–80 ratio. Trial data has been released
with the concreteness scores, while the test data
has been provided at the beginning of the
evaluation window without any score.2
4</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Measures and Baselines</title>
      <p>We chose the Spearman correlation indices as our
main evaluation measure; for the sake of
completeness, we also report Pearson indices
(substantially in accord with the previous metrics). We
chose the former measure because the collected
ratings are not normally distributed, which makes
the Spearman correlation more suited to the data.
In fact, by running the Shapiro–Wilk test we
obtained a p-value &lt; 0:001. The non normal
distribution of data is also confirmed by the plot of the
gold standard ratings, as illustrated in Figure 1.</p>
      <p>
        Two baselines have been designed for this task.
Baseline One. The first baseline for the Italian
language is derived as follows. The fastText word
embeddings have been acquired beforehand by
training the model on the Italian dump of the
WikiHow instructions. We chose fastText for its
support to the handling of OOV terms
        <xref ref-type="bibr" rid="ref5">(Bojanowski et
al., 2017)</xref>
        , which is a crucial feature in the present
setting. The cited norms by Montefinese et al.
(2014) (referred to as ‘the norms’ hereafter) have
been used herein. The average score of terms in
each input sentence S = ft1; t2; : : : tK g has been
2The dataset employed in the CONCRETEXT task is
available at the URL https://lablita.github.io/
CONcreTEXT/.
computed by scrolling through the content words
of the sentence. Each term t is searched in the
norms: if the term is found, the associated
concreteness score c(t) is returned; otherwise, if the
term is not present in the norms, the ranking of
the l (l = 20; 000) elements most similar to t is
generated through fastText. In this case, we scan
the whole norms list and employ the concreteness
score of the element in the norms closest to those
in the fastText ranking. In either case we obtain
a score for each and every term in the input
sentence, so that the concreteness score of the target
token t^ is computed as the averaged score of the
terms in the input sentence:
c(t^) = 1 XK c(ti):
      </p>
      <p>K i=1</p>
      <p>The first baseline for the English language is
analogous to the Italian one, except for the fact that
the English tokens from the norms are accessed in
this case. The same strategy governs the handling
of the fastText resource, that in this case has been
trained on the English dump of the Human
Instruction Dataset.</p>
      <p>Baseline Two. The second baseline for the
Italian language implements a simple lookup
function. More specifically, input sentences have been
translated into English through the Google
Translate ajax API implementation, and then the
concreteness scores associated to the terms in the
norms by Brysbaert et al. (2014b) are retrieved
(in the unlikely case the term is not found, it is
dropped, thus not contributing to the final score).
The concreteness score of the target term is thus
assigned to the average concreteness of terms in
the given input sentence. The baseline two for the
English language employs the concreteness score
—by also employing the norms by Brysbaert et
al. (2014b)— associated to all terms in the input
sentence, finally assigning to the target token the
average concreteness score for the whole sentence.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Systems Descriptions</title>
      <p>
        In this Section we briefly describe the systems that
participated in the competition. As a first edition,
the CONCRETEXT task recorded a good
feedback from the community, with 4 teams, overall
7 participants and 15 submitted system runs. In
the next Section we report the results obtained by
all such systems, while anonymizing a withdrawn
participant.
The ANDI team
        <xref ref-type="bibr" rid="ref26">(Rotaru, 2020)</xref>
        proposed a system
based on multiple classes of concreteness score
predictors. The first class of predictors has been
derived from large datasets of behavioral norms,
collected for a wide variety of psycholinguistic
factors. Beside well known concreteness norms,
ANDI takes into account also semantic diversity,
age of acquisition, emotional and sensori-motor
dimensions, as well as frequency and contextual
diversity counts. The vocabulary resulting from
the merging of these words collections comprises
more than 70K words, and it is the base
vocabulary used to extract all the predictors. The second
class of predictors has been derived from
contextindependent distributional models, namely
Skipgram, GloVe, and NumberBatch embeddings, as
well as from the concatenation of the three. The
third class of predictors has been derived from
features obtained through recent transformers
models, i.e. context-dependent representations. The
models exploited are: BERT, GPT-2, Bart, and
ALBERT. The final rating has been computed
through a ridge regression over the three classes.
5.2
      </p>
      <sec id="sec-4-1">
        <title>CAPISCO</title>
        <p>
          The CAPISCO Team
          <xref ref-type="bibr" rid="ref6">(Bondielli et al., 2020)</xref>
          submitted 3 systems for both Italian and English.
NON-CAPISCO. The first system computes a
variation of the Baseline Two; that is, the target
concreteness is obtained by combining the
concreteness value of the target term (taken in
isolation), and the average concreteness of the whole
sentence. Improvement from baseline comes from
considering differently the weight of the
concreteness of the target term and of the context.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>CAPISCO-CENTROIDS. This system is based</title>
        <p>on the assumption that close semantic spaces are
featured by similar concreteness scores. In this
case the authors first build two centroids, one for
concrete and one for abstract concepts based on
the norms by Brysbaert et al. (2014b) and Della
Rosa et al. (2010), by employing fastText
pretrained embeddings. The concreteness score of a
term is then computed by averaging the distance of
the first 50 lexical substitutes of the target
(identified through BERT) from the two polarized
centroids. Introducing a list of target substitutes in a
given context is thus the gist of this approach.</p>
      </sec>
      <sec id="sec-4-3">
        <title>System run</title>
        <p>ANDI
NON-CAPISCO
KONKRETIKA 3
KONKRETIKA 1
Baseline 2
KONKRETIKA 4
CAPISCO CENTR
KONKRETIKA 2
CAPISCO TRANS
Baseline 1
withdrawn run3
withdrawn run1
withdrawn run2</p>
      </sec>
      <sec id="sec-4-4">
        <title>System run</title>
        <p>
          ANDI
CAPISCO TRANS
CAPISCO CENTR
NON-CAPISCO
Baseline 2
Baseline 1
CAPISCO-TRANSFORMERS. In this variant,
the CAPISCO team fine-tuned a pre-trained BERT
model on the concreteness rating task, by
complementing the CONCRETEXT training data with
newly generated training data. The new data
generation is twofold: for each original sentence, new
sentences are generated by replacing the target
term with the first lexical substitutes derived with
BERT target masking approach. Then, more
sentences are borrowed from Italian and English
reference corpora.
The KONKRETIKA team
          <xref ref-type="bibr" rid="ref1">(Badryzlova, 2020)</xref>
          presented a system that first assigns a concreteness
and an abstractness score to the target lemma, and
then it adjusts these values based on the
surrounding context. In the first step, the system computes
semantic similarity between the target vectors and
a “seed list” consisting of abstract and concrete
words (extracted from the MRC Psycholinguistic
Database). In the second step, the values where
adjusted to the sentential context considering the
mean concreteness index of the entire sentence.
The team submitted 4 runs based on a heuristically
selected coefficient.
6
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>Four teams participated in the CONCRETEXT
competition: ANDI, CAPISCO, KONKRETIKA,
and a withdrawn team. ANDI and CAPISCO
developed a system for both languages (English and
Italian), while KONKRETIKA participated in the
English track only, and the same did the
withdrawn participant. Each team was allowed to
submit the output of up to 4 system runs; the final
ranking has been compiled based on the results of
the best run.</p>
      <p>In Tables 2 and 3 we present the score of each
run for the English and Italian language,
respectively. Although, as mentioned, the Spearman
indices were adopted as our main evaluation metrics,
we also report Pearson correlation indices and
Euclidean distance, that may be useful to complete
the assessment of the results. The final ranking is
provided in Tables 4 and 5.</p>
      <p>We can observe a substantial agreement
between Spearman and Pearson indices: the
averaged delta between such figures amounts to 0:012
and to 0:008 on the English and Italian dataset,
respectively. Also the Euclidean distance seems to
substantially confirm the results: for the results on
English (Table 2) it is minimal for the output of
the ANDI system, and it increases while Spearman
correlation values decrease. The same trend is also
confirmed on Italian results (Table 3).</p>
      <p>Tables 6 and 7 report disaggregated Spearman
correlations for verbs and nouns. This allows
to highlight if and to what extent the
participating systems obtained better results on either POS.
ANDI obtained the best results on both verbs and
nouns in both languages. This system (and
NONCAPISCO as well) obtained analogous results on
verbs and nouns. On the whole, the rest of the
systems obtained results clearly better on English
verbs and slightly better on Italian nouns. In
particular, KONKRETIKA (English only) is strongly
biased on verbs: its performances on verbs are
higher in all 4 runs. CAPISCO systems exhibit the
most varied behavior.
7</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>The obtained results confirm transformers as a
good device to compute concreteness score for
words in context. The virtues of
transformers in grasping contextual information are largely
known, but in the present setting we observe that
their output can be further improved by
integrating behavioral information (this seems to be one
major difference between the systems ANDI and
CAPISCO-TRANSFORMERS).</p>
      <p>The most important output of this challenge is
definitely the great performance of the ANDI
system, that proves to be robust and reliable for the
considered task: the system obtains the best
ranking in both languages, a low deviation from the
gold standard and a substantial stability in
processing both verbs and nouns. Moreover, the proposed
system is ready to be applied in a multi-language
environment, given that non-English sentences are
automatically translated into English. The ANDI
system exploits different kinds of available
resources and works with local and contextual
information. This shows that deriving the
concreteness score of a word in context is a complex task,
involving different semantic, cognitive and
experiential levels.</p>
      <p>The high correlation obtained by the
NONCAPISCO in the English task is somehow
surprising, since this system makes use only of the mean
concreteness of the sentence (computed from
existing norms) as contextual information. This
result is thus related to the availability of existing
norms, but it shows that there is a link between
the concreteness score of a target word in context
and the concreteness scores of the words it
occurs with. Further analysis are needed, but it
suggests that concrete interpretations of a target word
are associated with concrete context words. Of
course, systems based exclusively on behavioral
norms are strongly dependent on the coverage of
the considered vocabulary. In fact, the
NONCAPISCO Italian performances (obtained
exploiting a 1:2K vocabulary) are lower than all the
other systems, while on the English track it ranks
second (using a 70K vocabulary).
8</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>
        We presented the results of the CONCRETEXT
task at EVALITA 2020
        <xref ref-type="bibr" rid="ref3">(Basile et al., 2020)</xref>
        .
The task challenges participants to build NLP
systems to automatically assign a concreteness
score to words in context, evaluating to what
extent target concepts are concrete (i.e., more or
less perceptually salient) within a given context
of occurrence. A novel dataset was developed
for this task as a multilingual comparable
corpus composed of 550 Italian sentences and 534
English sentences, annotated with the
concreteness/abstractness rating of target nouns and verbs.
Three teams completed their participation to the
task, obtaining the following ranking: ANDI
        <xref ref-type="bibr" rid="ref26">(Rotaru, 2020)</xref>
        , CAPISCO
        <xref ref-type="bibr" rid="ref6">(Bondielli et al., 2020)</xref>
        , and
KONKRETIKA
        <xref ref-type="bibr" rid="ref1">(Badryzlova, 2020)</xref>
        .
      </p>
      <p>Future work will address the following steps.
First of all, we will improve our dataset by
including further languages, also from different language
families and under-resourced languages. Also the
set of considered targets should be expanded, to
ensure a broader coverage to the dataset, and more
significant results (thanks to the larger
experimental base) to its future users as well.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Yulia</given-names>
            <surname>Badryzlova</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>KONKRETIKA @ CONCRETEXT: Computing concreteness indexes with sigmoid transformation and adjustment for context</article-title>
          .
          <source>In Valerio Basile</source>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valentina</given-names>
            <surname>Bambini</surname>
          </string-name>
          , Donatella Resta, and
          <string-name>
            <given-names>Mirko</given-names>
            <surname>Grimaldi</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>A dataset of metaphors from the italian literature: Exploring psycholinguistic variables and the role of context</article-title>
          .
          <source>PloS one</source>
          ,
          <volume>9</volume>
          (
          <issue>9</issue>
          ):
          <fpage>e105634</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Evalita 2020: Overview of the 7th evaluation campaign of natural language processing and speech tools for italian</article-title>
          .
          <source>In Valerio Basile</source>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Fraser A</given-names>
            <surname>Bleasdale</surname>
          </string-name>
          .
          <year>1987</year>
          .
          <article-title>Concreteness-dependent associative priming: Separate lexical organization for concrete and abstract words</article-title>
          .
          <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>
          ,
          <volume>13</volume>
          (
          <issue>4</issue>
          ):
          <fpage>582</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Bondielli</surname>
          </string-name>
          , Gianluca E. Lebani,
          <string-name>
            <surname>Lucia C. Passaro</surname>
            , and
            <given-names>Alessandro</given-names>
          </string-name>
          <string-name>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>CAPISCO @ CONCRETEXT: (Un)supervised Systems to Contextualize Concreteness with Norming Data</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Marc</given-names>
            <surname>Brysbaert</surname>
          </string-name>
          , Michae¨l Stevens, Simon De Deyne, Wouter Voorspoels, and
          <string-name>
            <given-names>Gert</given-names>
            <surname>Storms</surname>
          </string-name>
          . 2014a.
          <article-title>Norms of age of acquisition and concreteness for 30,000 dutch words</article-title>
          .
          <source>Acta psychologica</source>
          ,
          <volume>150</volume>
          :
          <fpage>80</fpage>
          -
          <lpage>84</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Marc</given-names>
            <surname>Brysbaert</surname>
          </string-name>
          , Amy Beth Warriner, and
          <string-name>
            <given-names>Victor</given-names>
            <surname>Kuperman</surname>
          </string-name>
          . 2014b.
          <article-title>Concreteness ratings for 40 thousand generally known english word lemmas</article-title>
          .
          <source>Behavior research methods</source>
          ,
          <volume>46</volume>
          (
          <issue>3</issue>
          ):
          <fpage>904</fpage>
          -
          <lpage>911</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Paula</given-names>
            <surname>Chocron</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Pareti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Vocabulary alignment for collaborative agents: a study with real-world multilingual how-to instructions</article-title>
          .
          <source>In IJCAI</source>
          , pages
          <fpage>159</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Colla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Porporato</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Conceptual Abstractness: From Nouns to Verbs</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), volume
          <volume>2253</volume>
          . CEUR.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Louise</given-names>
            <surname>Connell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Dermot</given-names>
            <surname>Lynott</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Briony</given-names>
            <surname>Banks</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Interoception: the forgotten modality in perceptual grounding of abstract and concrete concepts</article-title>
          .
          <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source>
          ,
          <volume>373</volume>
          (
          <issue>1752</issue>
          ):
          <fpage>20170143</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Sebastian J Crutch and Elizabeth K Warrington</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Abstract and concrete concepts have structurally different representational frameworks</article-title>
          .
          <source>Brain</source>
          ,
          <volume>128</volume>
          (
          <issue>3</issue>
          ):
          <fpage>615</fpage>
          -
          <lpage>627</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Annette M de Groot</surname>
          </string-name>
          .
          <year>1989</year>
          .
          <article-title>Representational aspects of word imageability and word frequency as assessed through word association</article-title>
          .
          <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>
          ,
          <volume>15</volume>
          (
          <issue>5</issue>
          ):
          <fpage>824</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Pasquale A Della Rosa</surname>
            , Eleonora Catricala`,
            <given-names>Gabriella</given-names>
          </string-name>
          <string-name>
            <surname>Vigliocco</surname>
          </string-name>
          , and Stefano F Cappa.
          <year>2010</year>
          .
          <article-title>Beyond the abstract-concrete dichotomy: Mode of acquisition, concreteness, imageability, familiarity, age of acquisition, context availability, and abstractness norms for a set of 417 italian words</article-title>
          .
          <source>Behavior research methods</source>
          ,
          <volume>42</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1042</fpage>
          -
          <lpage>1048</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Francesca</given-names>
            <surname>Garbarini</surname>
          </string-name>
          , Fabrizio Calzavarini, Matteo Diano, Monica Biggio, Carola Barbero, Daniele P Radicioni, Giuliano Geminiani, Katiuscia Sacco, and Diego Marconi.
          <year>2020</year>
          .
          <article-title>Imageability effect on the functional brain activity during a naming to definition task</article-title>
          .
          <source>Neuropsychologia</source>
          ,
          <volume>137</volume>
          :
          <fpage>107275</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Karl F Haberlandt and Arthur C Graesser</surname>
          </string-name>
          .
          <year>1985</year>
          .
          <article-title>Component processes in text comprehension and some of their interactions</article-title>
          .
          <source>Journal of Experimental Psychology: General</source>
          ,
          <volume>114</volume>
          (
          <issue>3</issue>
          ):
          <fpage>357</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Felix</given-names>
            <surname>Hill</surname>
          </string-name>
          and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Korhonen</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Learning abstract concept embeddings from multi-modal data: Since you probably can't see what i mean</article-title>
          .
          <source>In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>255</fpage>
          -
          <lpage>265</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Felix</given-names>
            <surname>Hill</surname>
          </string-name>
          , Douwe Kiela, and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Korhonen</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Concreteness and corpora: A theoretical and practical study</article-title>
          .
          <source>In Proceedings of the Fourth Annual Workshop on Cognitive Modeling and Computational Linguistics (CMCL)</source>
          , pages
          <fpage>75</fpage>
          -
          <lpage>83</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Stavroula-Thaleia</surname>
            <given-names>Kousta</given-names>
          </string-name>
          , Gabriella Vigliocco, David P Vinson,
          <source>Mark Andrews, and Elena Del Campo</source>
          .
          <year>2011</year>
          .
          <article-title>The representation of abstract words: why emotion matters</article-title>
          .
          <source>Journal of Experimental Psychology: General</source>
          ,
          <volume>140</volume>
          (
          <issue>1</issue>
          ):
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Judith F Kroll and Jill S Merves</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Lexical access for concrete and abstract words</article-title>
          .
          <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):
          <fpage>92</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Mensa</surname>
          </string-name>
          , Aureliano Porporato, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          . 2018a.
          <article-title>Annotating concept abstractness by common-sense knowledge</article-title>
          .
          <source>In Chiara Ghidini</source>
          , Bernardo Magnini, Andrea Passerini, and Paolo Enrico Mensa, Aureliano Porporato, and
          <string-name>
            <given-names>Daniele P.</given-names>
            <surname>Radicioni</surname>
          </string-name>
          . 2018b.
          <article-title>Grasping metaphors: Lexical semantics in metaphor analysis</article-title>
          .
          <source>In Aldo Gangemi</source>
          , Anna Lisa Gentile, Andrea Giovanni Nuzzolese,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Rudolph</surname>
          </string-name>
          , Maria Maleshkova, Heiko Paulheim, Jeff Z Pan, and Mehwish Alam, editors,
          <source>The Semantic Web: ESWC 2018 Satellite Events</source>
          , pages
          <fpage>192</fpage>
          -
          <lpage>195</lpage>
          , Cham. Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Leonie M Miller</surname>
            and
            <given-names>Steven</given-names>
          </string-name>
          <string-name>
            <surname>Roodenrys</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The interaction of word frequency and concreteness in immediate serial recall</article-title>
          .
          <source>Memory &amp; Cognition</source>
          ,
          <volume>37</volume>
          (
          <issue>6</issue>
          ):
          <fpage>850</fpage>
          -
          <lpage>865</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Montefinese</surname>
          </string-name>
          , Ettore Ambrosini, Beth Fairfield, and
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Mammarella</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The adaptation of the affective norms for english words (anew) for italian</article-title>
          .
          <source>Behavior research methods</source>
          ,
          <volume>46</volume>
          (
          <issue>3</issue>
          ):
          <fpage>887</fpage>
          -
          <lpage>903</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Montefinese</surname>
          </string-name>
          , Ettore Ambrosini, Antonino Visalli, and
          <string-name>
            <given-names>David</given-names>
            <surname>Vinson</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Catching the intangible: a role for emotion? Behavioral and</article-title>
          Brain Sciences,
          <volume>43</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Romani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sheila</given-names>
            <surname>Mcalpine</surname>
          </string-name>
          , and
          <string-name>
            <surname>Randi C Martin</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Concreteness effects in different tasks: Implications for models of short-term memory</article-title>
          .
          <source>Quarterly Journal of Experimental Psychology</source>
          ,
          <volume>61</volume>
          (
          <issue>2</issue>
          ):
          <fpage>292</fpage>
          -
          <lpage>323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Armand</given-names>
            <surname>Rotaru</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>ANDI @ CONCRETEXT: Predicting concreteness in context for English and Italian using distributional models and behavioural norms</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Mark</given-names>
            <surname>Sadoski</surname>
          </string-name>
          , William A Kealy,
          <string-name>
            <surname>Ernest T Goetz</surname>
            , and
            <given-names>Allan</given-names>
          </string-name>
          <string-name>
            <surname>Paivio</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Concreteness and imagery effects in the written composition of definitions</article-title>
          .
          <source>Journal of Educational Psychology</source>
          ,
          <volume>89</volume>
          (
          <issue>3</issue>
          ):
          <fpage>518</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Paula J Schwanenflugel and Edward J Shoben</surname>
          </string-name>
          .
          <year>1983</year>
          .
          <article-title>Differential context effects in the comprehension of abstract and concrete verbal materials</article-title>
          .
          <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>
          ,
          <volume>9</volume>
          (
          <issue>1</issue>
          ):
          <fpage>82</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Peter</given-names>
            <surname>Turney</surname>
          </string-name>
          , Yair Neuman, Dan Assaf, and Yohai Cohen.
          <year>2011</year>
          .
          <article-title>Literal and metaphorical sense identification through concrete and abstract context</article-title>
          .
          <source>In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>680</fpage>
          -
          <lpage>690</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Vigliocco</surname>
          </string-name>
          , Lotte Meteyard,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Andrews</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Stavroula</given-names>
            <surname>Kousta</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Toward a theory of semantic representation</article-title>
          .
          <source>Language and Cognition</source>
          ,
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <fpage>219</fpage>
          -
          <lpage>247</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>