<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>KonKretiKa @ CONcreTEXT: Computing Concreteness Indexes with Sigmoid Transformation and Adjustment for Context</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yulia Badryzlova</string-name>
          <email>yuliya.badryzlova@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HSE University Moscow</institution>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The present paper is a technical report of KonKretiKa, a system for computation of concreteness indexes of words in context, submitted to the English track of the CONcreTEXT shared task. We treat concreteness as a bimodal problem and compute the concreteness indexes using paradigms of concrete and abstract seed words and distributional semantic similarity. We also conduct sigmoid transformation to achieve greater similarity to the psycholinguistically attested data, and apply dynamic adjustment of static indexes for sentential context. One of the modifications of the presented system ranked third in the task, with rs = .6634 and r = .6685 against the gold standard.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This paper is a description of the system with the
working title KonKretiKa, which was submitted
to the English track of CONcreTEXT, the shared
task on evaluation of concreteness in context
        <xref ref-type="bibr" rid="ref10">(Gregori et al., 2020)</xref>
        offered at EVALITA 2020,
the 7th evaluation campaign of Natural Language
Processing and speech tools for the Italian
language
        <xref ref-type="bibr" rid="ref5">(Basile et al., 2020)</xref>
        .
      </p>
      <p>KonKretiKa stems from our previous work on
computation of such indexes for the purposes of
metaphor identification.</p>
      <p>
        Computationally obtained indexes of
concreteness are extensively explored in experiments for
automated metaphor identification. Application of
concreteness indexes to metaphor identification
relies on the assumptions made by the theories of
embodied and grounded cognition
        <xref ref-type="bibr" rid="ref4">(Barsalou,
2008)</xref>
        , and primary and conceptual metaphor
        <xref ref-type="bibr" rid="ref13">(Lakoff and Johnson, 1980)</xref>
        . These theories claim
that human thinking is intrinsically metaphoric,
since the conceptual representations underlying
knowledge are grounded in sensory and motor
systems, and conceptual metaphor is the primary
mechanism for transferring conventional mental
imagery from sensorimotor domains to the
domains of subjective experience.
      </p>
      <p>An established method to compute the
concreteness index of a word is to collect two sets of
lexemes (‘seed lists’, or ‘paradigms’) consisting
of abstract and concrete words – and to measure
the lexical similarity between each word in the
lexicon and each of the paradigm words.</p>
      <p>
        <xref ref-type="bibr" rid="ref18">Turney et al. (2011)</xref>
        use concreteness indexes
to identify linguistic metaphor in the TroFi dataset
        <xref ref-type="bibr" rid="ref6">(Birke and Sarkar, 2006)</xref>
        . They compute the
concreteness index of a word by comparing its
distributional semantic embedding to the vector
representations of 20 abstract and 20 concrete words.
      </p>
      <p>
        The paradigm words are automatically selected
from the MRC Psycholinguistic Database
Machine Usable Dictionary
        <xref ref-type="bibr" rid="ref8">(Coltheart, 1981)</xref>
        , a
collection of 4,295 English words rated with degrees
of abstractness by human subjects in
psycholinguistic experiments.
      </p>
      <p>
        <xref ref-type="bibr" rid="ref17">Tsvetkov et al. (2013)</xref>
        also compute the
concreteness indexes of English words by using a
distributional semantic model and the MRC
database. They train a logistic regression classifier on
1,225 most abstract and 1,225 most concrete
words from MRC; the degree of concreteness of a
word is the posterior probability produced by the
classifier. The Tsvetkov et al. system for
metaphor identification with concreteness indexes is
based on cross-lingual model transfer, when the
model is trained on English data, and then the
classification features are translated into other
languages by means of an electronic dictionary.
r
c
n
o
C
albatross, balloon, bench, bridge, catfish, cauliflower, chicken, clown, corkscrew, crab,
daisy, deer, eagle, egg, frog, garlic, goat, harpsichord, lion, mattress, mussel, nightgown,
nightingale, owl, ox, pants, peach, piano, pig, potato, quilt, rabbit, saxophone, sheep,
shrimp, skyscraper, sofa, stoat, tulip, turtle
affirmation, animosity, demeanour, derivation, determination, detestation, devotion,
enunciatc tion, etiquette, fallacy, forethought, gratitude, harm, hatred, ignorance, illiteracy, impatience,
a
tr independence, indolence, inefficiency, insufficiency, integrity, intellect, interposition,
justifis
b cation, malice, mediocrity, obedience, oblivion, optimism, prestige, pretence, reputation,
reA
sentment, tendency, unanimity, uneasiness, unhappiness, unreality, value
∀   , ∀  ∃   = {
(  ,  1), 
(  ,  2), … , 
(  ,   ), … , 
(  ,   )} ,
where  is the set of words in the vocabulary,
 is the set of words in the seed list, k is the number of elements in S

= {  1′ ,   ′2, … ,   1′0} ,
 = 
{
      </p>
      <p>}
where   ′ is a linearly ordered set of   (in ascending order)</p>
    </sec>
    <sec id="sec-2">
      <title>Equations 1-3. Computation of indexes with paradigm lists.</title>
      <p>(1)
(2)
(3)</p>
      <p>
        <xref ref-type="bibr" rid="ref1">Badryzlova (2020)</xref>
        explores concreteness and
abstractness indexes for linguistic metaphor
identification in Russian and English. The paradigm
words are selected in a semi-automatic fashion:
the Russian paradigm is derived from the Open
Semantics of the Russian Language, the
semantically annotated dataset of the KartaSlov database
        <xref ref-type="bibr" rid="ref11">(Kulagin, 2019)</xref>
        ; the English paradigm is selected
from the MRC database
        <xref ref-type="bibr" rid="ref8">(Coltheart, 1981)</xref>
        . The
indexes of concreteness and abstractness are
computed for large sets of Russian and English words
(about 18,000 and 17,000 lexemes, respectively).
The metaphor identification in Russian is
conducted on the RusMet corpus
        <xref ref-type="bibr" rid="ref2 ref3">(Badryzlova, 2019;
Badryzlova and Panicheva, 2018)</xref>
        , and the
English on the TroFi dataset. The author shows that
the distributions of concreteness and abstractness
indexes in the two languages follow the same
pattern: in the lexicon, there is a distinct group of
highly concrete words, which have very high
concreteness and very low abstractness indexes;
similarly, there is a group of distinctly abstract
vocabulary, with low concreteness and high
abstractness scores. Moreover, there is a general trend for
abstractness indexes to increase as the
corresponding concreteness indexes decrease. The
author also observes statistical correlation between
two Russian abstractness ratings, which may
indicate that the category of abstractness is more
semantically homogeneous than the category of
concreteness.
      </p>
      <p>
        The present work develops and extends the
method of
        <xref ref-type="bibr" rid="ref1">Badryzlova (2020)</xref>
        in two directions:
(a) we apply sigmoid transformation to fit the
curve comprised of the computed concreteness
and abstractness indexes to the distribution of
indexes in psycholinguistic data; and (b) we suggest
a method for dynamic adjustment of the obtained
indexes for sentential context, according to the
requirements of the CONcreTEXT shared task
        <xref ref-type="bibr" rid="ref10">(Gregori et al., 2020)</xref>
        . The working title of the
proposed system is KonKretiKa.
2
      </p>
      <sec id="sec-2-1">
        <title>Description of the system</title>
        <p>
          We demonstrate a method for evaluating
concreteness on English data; however, it can be
transferred to any other language provided that the
following
types
of resources
are
available:
(1) a lexicon with semantic
          <xref ref-type="bibr" rid="ref11 ref9">(e.g. Fellbaum, 1998;
Kulagin, 2019)</xref>
          or psycholinguistic
          <xref ref-type="bibr" rid="ref7 ref8">(e.g. Brysbaert
et al., 2014; Coltheart, 1981)</xref>
          annotation to select
the paradigm words from; (2) a pre-trained
distributional semantic
        </p>
        <p>model; and (3) a relatively
large wordlist containing lexemes with different
frequencies of occurrence (ipm) in order to ensure
the maximum possible variation in concreteness
across the lexicon.</p>
        <p>When analyzing the distribution of
psycholinguistic
concreteness
ratings,</p>
        <p>
          <xref ref-type="bibr" rid="ref7">Brysbaert et al.
(2014)</xref>
          observe that “concreteness and
abstractness may be not the two extremes of a quantitative
continuum […], but two qualitatively different
characteristics.” of a word. Following this
observation and the previous work in
          <xref ref-type="bibr" rid="ref1">(Badryzlova,
2020)</xref>
          , we treat concreteness as a bimodal
property investing the word with two characteristics:
the rate of concreteness and the rate of
abstractness. Thus, we start by computing the standalone
indexes of concreteness and of abstractness; then,
the single aggregate index is computed as a
function of these two indexes.
2.1
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Computation of raw indexes with paradigm words and distributional semantic similarity</title>
        <p>
          Computation of the standalone concreteness and
abstractness indexes is based on paradigm lists of
concrete and abstract words; we use the English
concrete and abstract paradigms from
          <xref ref-type="bibr" rid="ref1">Badryzlova
(2020)</xref>
          . These paradigms were compiled from the
MRC Psycholinguistic Database: nouns from the
top and from the end of the MRC concreteness
rating were drawn to populate the concrete and the
abstract paradigms, respectively. The paradigm
lists are presented in Table 1.
        </p>
        <p>
          The indexes of concreteness and abstractness
were computed using a Continuous Skip-Gram
model
          <xref ref-type="bibr" rid="ref12">(Kutuzov et al., 2017)</xref>
          which had been
pre-trained on the lemmatized Gigaword 5th
Edition corpus
          <xref ref-type="bibr" rid="ref15">(Parker et al., 2011)</xref>
          .
        </p>
        <p>As shown in Equations 1-3, to compute a
concreteness or an abstractness index ( ) of a word,
we measured semantic similarity (cosine distance)
Sim between the vectors of this word and each
word in the paradigm (concrete or abstract,
respectively), and took the mean of the ten nearest
semantic neighbors (NN).</p>
        <p>
          In total, we computed concreteness and
abstractness indexes for approximately 23,000
English words (nouns, verbs, adjectives, and adverbs);
this lexicon was taken from the
          <xref ref-type="bibr" rid="ref7">Brysbaert et al.
(2014)</xref>
          ranking, which allowed us to analyze the
correlation between the computational and the
large-scale psycholinguistic data at the
subsequent stages of the present study (see Section 3).
        </p>
        <p>The obtained computational sets of
concreteness and abstractness indexes were normalized to
the range [1, 7]1 in order to comply with the scale
set by the CONcreTEXT shared task. In order to
obtain an aggregate single-value index of a word,
which would be representative of both its
concreteness and abstractness, we subtracted the
abstractness indexes from the concreteness indexes.
m
e
t
s
y</p>
        <p>S</p>
        <sec id="sec-2-2-1">
          <title>Leader-1</title>
        </sec>
        <sec id="sec-2-2-2">
          <title>Leader-2</title>
        </sec>
        <sec id="sec-2-2-3">
          <title>KonKretiKa-3</title>
        </sec>
        <sec id="sec-2-2-4">
          <title>KonKretiKa-1</title>
        </sec>
        <sec id="sec-2-2-5">
          <title>Baseline-2</title>
        </sec>
        <sec id="sec-2-2-6">
          <title>KonKretiKa-4 KonKretiKa-2</title>
          <p>on )
ita lau (tcn
rm ep tx emResult (rs) Result (r)
fso ty ten ts
an oC jud
rT a
2
1
2
1
0.5
0.5
0.8
0.8
where  defines the slope of the function and 
defines the inflection point. Consequently, we can
transform the sigmoid by changing the  and 
coefficients.</p>
          <p>In the submissions to the CONcreTEXT shared
task, we experimented with two transformations
of the raw KonKretiKa curve (Figure 2). In the
first transformation, we applied a heuristically
chosen combination of  and  which was
intended to increase the slope and the curvature
while preserving the S-shape of the sigmoid. The
second transformation was intended to attain
maximum resemblance of its shape to the
Brysbaert et al. curve. We used grid search with different
combinations of coefficients  and  to maximize
the correlation between the two curves. During
2 The KonKretiKa ranking is available at:
https://github.com/yubadryzlova/CONcreTEXT-2020
this fitting, only the values of the indexes are
adjusted, while their initial ranks remain intact –
thus, there is no data leakage from the
psycholinguistic ranking. 2
2.3</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Contextual adjustment</title>
        <p>Since the CONcreTEXT shared task requires that
the concreteness indexes of target words be
dynamically adjusted to their sentential context, the
following heuristic was applied in the submitted
KonKretiKa models. We computed the mean
concreteness of all content words in the sentence
(with the target word excluded) and adjusted the
concreteness value of the target word accordingly.
The adjusted index  was computed as follows:
  =   − ( ∗  )
where  is the target word,  is the raw index
from the KonKretiKa ranking,  is the mean
concreteness of the sentence, and  is the adjustment
coefficient. In the models submitted to the
CONcreTEXT shared task, we applied two
heuristically defined  coefficients:  = 0.5 and  = 0.8.</p>
        <p>Thus, the four modifications of KonKretiKa
submitted to the shared task were differentiated by
the two parameters: the type of transformation and
the contextual adjustment coefficient.
3</p>
      </sec>
      <sec id="sec-2-4">
        <title>Results and discussion</title>
        <p>The parameters of the four modifications and their
results are presented in Table 2 (along with the
Baselines and the Leaders). The results indicate
that systems with the lower coefficient of
sentential adjustment (0.5) perform better than systems
with the higher adjustment coefficient (0.8)
irrespective of the type of sigmoid transformation;
yet, the system with Type 2 (fitted to the
psycholinguistic data) transformation somewhat
outperforms the system with Type 1 (S-shaped)
transformation.</p>
        <p>The best of our modifications, KonKretiKa-3,
demonstrated Spearman correlation with the gold
standard rs = .6634 and Pearson correlation
r = .6685, ranking our system third in the track,
yet by a substantial margin behind the two
winning system (with rs = .83313 and r = .83406 and
rs = .78541 and r = .78682, respectively).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Dataset</title>
    </sec>
    <sec id="sec-4">
      <title>KKK (static)</title>
    </sec>
    <sec id="sec-5">
      <title>Gold</title>
      <p>(dynamic)</p>
      <p>BRY
(static)
rs =.743
r =.751
We carried out a post hoc analysis of the contextal
adjustment coefficient (c) by using grid search to
maximize the correlation between KonKretiKa
(Type 2 transformation) and the gold standard.
Moreover, we altered the scope of the context
words for which the mean sentential concreteness
(M) was computed – by taking 2-3 nearest
semantic neighbors (either of any part of speech, or only
nouns, or only verbs); this was done in order to
reduce the possible noise from the words that are
not semantically related to the target in the
sentence. The change of the contextual scope did not
lead to a substantial difference in the result. As for
the contextual adjustment coefficient, the grid
search showed that c = 0.32 – which is lower than
the most efficient coefficient from our earlier
submissions (c = 0.5 in KonKretiKa-3) – results in a
slight increase of correlations: rs = .678 and
r = .688.</p>
      <p>A closer analysis of the test sentences suggests
that contribution of contextual adjustment
presumably may be increased by considering a
broader context of a sentence – for instance,
spanning over 1-3 adjacent sentences from the left and
the right contexts; this option constitutes a
possible direction for future work.
3.1</p>
      <sec id="sec-5-1">
        <title>Comparison of computational and psycholinguistic data</title>
        <p>Pairwise correlations between the computational
(KonKretiKa, KKK) and the psycholinguistic
rankings (Brysbaert et al., BRY and the gold
standard) are shown in Table 3. It can be seen that
KKK better correlates with the BRY data than
with the gold standard (rs = .743, r = .751 vs.
rs = .663, r = .669, respectively). Presumably,
such difference in the two correlations is due to
the much larger size of the BRY lexicon. The
correlation between the two psycholinguistic datasets
word
handmaiden (N)
tire (V)
bedrock (N)
alarm (N)
text (N)
nonreactive (ADJ)
temptingly (ADV)
hail (V)
stance (N)
nudge (N)
chasm (N)</p>
        <p>7
6.18
(BRY vs. Gold) is rs = .755, r = .761, which is
close to the correlation between KKK and BRY.</p>
        <p>We undertook closer pairwise comparative
analysis between two pairs of rankings:
1.</p>
        <p>Static KonKretiKa indexes (the indexes
after Type 2 sigmoid transformation, without
contextual adjustment) vs. the
Brysbaert et al. ranking (which is also static):
approximately 23,000 words – nouns, verbs,
adjectives, and adverbs (the two wordlists
are identical).</p>
        <p>Indexes of the target words from the
CONcreTEXT test data as presented in the
dynamic version of KonKretiKa (the
sigmoid-transformed Type 2 indexes with
contextual adjustment coefficient c = 0.32)
vs. the Gold standard (where the target
words are also ranked dynamically in
context): 436 words – verbs and nouns.</p>
        <p>The top residuals between the KonKretiKa and
the Brysbaert et al. indexes are presented in
Table 4. Analysis of these discrepancies suggests
that most of them stem from polysemy and the
differences between its representation in
distributional semantic models and in psycholinguistic
reality. Thus, distributional semantic models do not
discriminate between various meanings of words;
if occurrences of one of the meanings
substantially outnumber the other meanings in discourse
and, as a consequence, in the training corpus, the
resulting vector reflects the more frequent
meaning.
e
c
n
e
t
n
e
S
t
e d
g r
r o
a
T w
d
l
o
G</p>
        <p>K
K
K
f
f
i
D
399 vision (N)</p>
        <p>Check your &lt; vision &gt; to see if you are seeing blurry or double.
353 vision (N)</p>
        <p>With retinal migraine, you may experience loss of &lt; vision &gt; in one
eye and a headache that starts behind your eyes.
61
385
81
155 spirit (N)
324 pain (N)
6</p>
        <p>Gin is an alcoholic &lt; spirit &gt; made from distilled grain or malt.</p>
        <p>See your doctor if you are experiencing &lt; pain &gt; or discomfort.
answer (N)
war (N)</p>
        <p>Be sure to write your final &lt; answer &gt; without the negative sign.</p>
        <p>They have escaped from civil &lt; war &gt; in Liberia or Zimbabwe.
answer (N)</p>
        <p>Final &lt; answers &gt; for equations are considered wrong unless you
have broken them down to their simplest form.
237 heart (N)
163 pain (N)</p>
        <p>The &lt; heart &gt; pumps blood due to an internal electrical system.
Take your medications to ease your physical &lt; pain &gt;.
176 agreement (N) 5.16 1.85 3.31
After signing the indemnification &lt; agreement &gt;, you can sign the
legally binding bond agreement.</p>
        <p>For example, the nearest semantic neighbors of
the noun handmaiden in the distributional
semantic model3 are: embodiment, personification,
epitome, and paragon – associating this word with its
abstract, metaphoric meaning ‘something that
supports something else that is more important’4,
whereas for speakers of English the other,
concrete meaning ‘a woman who is someone’s
servant’ apparently stands out as being more salient.</p>
        <p>Similarly, among the nearest semantic neighbors
of the noun chasm in the distributional semantic
model are: disparity, schism, rich-poor divide,
mistrust, (the) haves, divergence, antagonism, and
inequality – indicating that the distributional
vector of chasm is biased towards the abstract
meaning of this word (‘a very big difference that
separates one person or group from another’) rather
than the concrete one (‘a very deep crack in rock
or ice’), while human subjects see the latter
meaning as more salient or prevalent.</p>
        <p>As for nonreactive and temptingly, which are
more concrete in the computational data, this
could be explained by their perceived vagueness
to human subjects, since these words do not have
meanings that would be markedly juxtaposed to
each other in terms of concreteness-abstractness –
thus ranking them rather low in the
psycholinguistic data. Meanwhile, the nearest semantic
neighbors of temptingly in the distributional semantic
model are: strappy sandal, capelet, knee-length
skirt, enticingly, floral-print, high-heeled sandal,
lace-trimmed, harem pants, and puffed sleeve – all
rather concrete objects (or the properties of such
objects).</p>
        <p>The top residuals between KonKretiKa and the
gold standard are shown in Table 5. The
discrepancy between the abstract meaning of vision (‘the
ability to think about and plan for the future, using
intelligence and imagination, especially in politics
and business’) and its concrete meaning (‘the
ability to see’) can also be attributed to the differences
between representation of meanings in
distributional semantic models and in psycholinguistic
reality – the reason already discussed above. Thus,
the nearest distributional semantic neighbors of
vision are: worldview, ideal, visionary, thinking,
perspective, idea, dream, and blueprint – rather
than terms related to eyesight.</p>
        <p>
          The noun spirit in Table 5 (Sentence 155) is
used in the sense of ‘strong alcoholic drink’.
However, its nearest neighbors in the distributional
semantic model are ethos, ideal, idealism, tradition,
essence, enthusiasm, passion, faith, chivalric,
zeal, credo, and compassion – indicating that the
meaning ‘your attitude to life or to other people’
3 Continuous Skip-Gram model
          <xref ref-type="bibr" rid="ref12">(Kutuzov et al.,
2017)</xref>
          , pre-trained on Gigaword 5th Edition corpus
4 Definitions are cited according to Macmillan
Dictionary (n.d.)
is dominant in the model, and the contextual
adjustment we apply is not sufficient for overcoming
the abstractness of the dominant meaning.
        </p>
        <p>As for the noun war, its nearest neighbors in the
distributional semantic model are conflict,
warfare, invasion, 1991-95 Serbo-Croatian,
IsraelHezbollah, genocide, Bosnia war, Jehad,
civilwar, Croatia war, Cold War, Iran-Iraq, wartime,
Vietnam-like, etc. – that is, rather abstract
concepts. The only more concrete words referring to
physical combat action that occur in the
distributional semantic neighborhood of war are
battlefield and bloodshed, but this is not enough to
outweigh the abstract terms. Thus, the distributional
semantic model models warfare in terms of
abstract rather than concrete (such as names of
weapons, military equipment, military personnel,
etc.) concepts. As a result, military action is not
sufficiently juxtaposed to the metaphoric meaning
of war as ‘a situation in which two people or
groups of people fight, argue, or are extremely
unpleasant to each other’.</p>
        <p>In the case of answer and agreement, their
nearest distributional semantic neighbors in the model
are fairly abstract concepts: explanation, answer,
reply, solution, unanswerable, query, TV-talkback
answer, question, and yes (for answer), and
accord, pact, deal, treaty, initial, negotiation,
memorandum, compromise, and negotiate (for
agreement). Meanwhile, human subjects rank answer
and agreement rather high in concreteness;
presumably, this is a consequence of conflating the
mental representations of the action of
answering / reaching an agreement with their two modes
– the spoken and the written, i.e. with the physical
actions of speaking and writing. This conflation is
not reflected in discourse – it largely exists in the
mental representations of answer and agreement
and, therefore, is not very distinguishable on the
level of linguistic representation.</p>
        <p>Of interest are the cases of heart and pain,
which have much lower concreteness in
KonKretiKa than in the gold standard sentences
where these words are used in their physical,
concrete meanings. The nearest distributional
semantic neighbors of heart are heart-related, coronary
artery, kidney, liver, lung, arrhythmia, cardiac,
angina, and aneurism. The nearest neighbors of
pain are discomfort, ache, agony, tingling
sensation, numbness, soreness, menstrual cramp,
lightheadedness, stiffness, nausea, and arthritis. It
would be quite expected for such semantic
neighborhood to entitle heart and pain to higher
concreteness values than what they receive in
KonKretiKa. A more in-depth analysis into this
contradiction revealed that it stems from the
vulnerability in the semantic composition of the
concrete paradigm which was used to compute the
raw indexes (see Table 1). The words of this
paradigm belong to the two major semantic classes –
living organisms (animals and plants) and
manmade artifacts. The class of words denoting
human beings was intentionally excluded when the
paradigm was compiled on the grounds that such
nouns tend to indicate abstract social roles rather
than physical humans. As a consequence, physical
organic objects such as body parts and organs, or
physical sensations and physiological conditions
received non-uniform indexes in KonKretiKa:
those that refer to humans as well as to animals
(e.g. in veterinary or gastronomic discourse)
ranked rather high in concreteness: e.g. liver (6.6),
pancreas (6.4), foot (6.3), encephalitis (6.25),
kidney (6.25), entrails (6.05), tummy (5.92),
womb (5.6) – whereas those that tend to be
primarily associated with humans received lower
indexes, e.g. heart (2.63), heartburn (2.57),
scar (2.53), nausea (2.5), headache (1.61),
distress (1.5), pain (1.21), queasiness (1.12), etc.</p>
        <p>
          Thus, comparison of the KonKretiKa
computational indexes with the psycholinguistic data of
CONcreTEXT allowed us to detect a potential
shortcoming in our approach to the design of the
concrete paradigm. As was noted in previous
study
          <xref ref-type="bibr" rid="ref1">(Badryzlova, 2020)</xref>
          , the class of concrete
words seems to be more semantically
heterogeneous than of abstract words; therefore, it may
reasonable in future experiments to diversify the
concrete paradigm and expand it in size by including
words that denote human beings.
4
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Conclusions</title>
        <p>We presented KonKretiKa system for computing
concreteness indexes of English words in context;
the system was submitted to the English track of
the CONcreTEXT shared task. The best
modification of KonKretiKa ranked third in the task, with
rs = .6634 and r = .6685 against the gold standard.
We treat concreteness as a bimodal problem and
use paradigm lists of concrete and abstract words
to compute two indexes for each word, that of
concreteness and of abstractness. The single
aggregate index indicative of both the word’s
concreteness and abstractness is computed as the
function of the two respective indexes. The set of
raw aggregate indexes is transformed using
sigmoid transformation to increase the variance and
to attain greater similarity to the psycholinguistic
data. To dynamically adjust the concreteness
indexes to the context, we apply an adjustment
coefficient. Post hoc analysis of the adjustment
coefficient indicates that lower coefficients lead to
better performance. We hypothesize that the
contribution of the adjustment coefficient could be
increased by expanding the scope of the context, for
example, by considering one or more sentences
from the left and the right contexts of the target
sentence. According to our analysis, the main
source of divergence between the computational
and the psycholinguistic indexes lies in the
different representation, or salience, of word meanings
in distributional semantic models and in
psycholinguistic reality. Besides, analysis of divergences
between the computational and the
psycholinguistic rankings prompted us a potential direction for
reducing the bias in composition of the
concreteness paradigm, which can be overcome by
diversifying the paradigm.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Badryzlova</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <year>2020</year>
          .
          <article-title>Exploring Semantic Concreteness and Abstractness for Metaphor Identification and Beyond</article-title>
          .
          <source>Computational Linguistics and Intellectual Technologies</source>
          <volume>33</volume>
          - 47.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Badryzlova</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Automated metaphor identification in Russian texts</article-title>
          . National Research University Higher School of Economics, Moscow.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Badryzlova</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Panicheva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2018</year>
          .
          <article-title>A Multi-feature Classifier for Verbal Metaphor Identification in Russian Texts</article-title>
          ,
          <source>in: Conference on Artificial Intelligence and Natural Language</source>
          . Springer, pp.
          <fpage>23</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Barsalou</surname>
            ,
            <given-names>L.W.</given-names>
          </string-name>
          ,
          <year>2008</year>
          .
          <article-title>Grounded cognition</article-title>
          .
          <source>Annu. Rev. Psychol</source>
          .
          <volume>59</volume>
          ,
          <fpage>617</fpage>
          -
          <lpage>645</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di</surname>
            <given-names>Maro</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Passaro</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.C.</surname>
          </string-name>
          ,
          <year>2020</year>
          . EVALITA 2020:
          <article-title>Overview of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          , in: Basile,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Croce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Maro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Passaro</surname>
          </string-name>
          , L.C. (Eds.),
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ). CEUR.org, Online.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Birke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarkar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2006</year>
          .
          <article-title>A Clustering Approach for Nearly Unsupervised Recognition of Nonliteral Language</article-title>
          ., in: EACL.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Brysbaert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warriner</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuperman</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <year>2014</year>
          .
          <article-title>Concreteness ratings for 40 thousand generally known English word lemmas</article-title>
          .
          <source>Behavior research methods 46</source>
          ,
          <fpage>904</fpage>
          -
          <lpage>911</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Coltheart</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <year>1981</year>
          .
          <article-title>The MRC psycholinguistic database</article-title>
          .
          <source>The Quarterly Journal of Experimental Psychology Section A</source>
          <volume>33</volume>
          ,
          <fpage>497</fpage>
          -
          <lpage>505</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <year>1998</year>
          .
          <article-title>WordNet: An electronic database</article-title>
          . MIT Press, Cambridge, MA.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Gregori</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montefinese</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radicioni</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravelli</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varvara</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <year>2020</year>
          .
          <article-title>CONcreTEXT @ Evalita2020: the Concreteness in Context Task</article-title>
          , in: Basile,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Croce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Maro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Passaro</surname>
          </string-name>
          , L.C. (Eds.),
          <source>Proceedings of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA</source>
          <year>2020</year>
          ). CEUR.org, Online.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Kulagin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <year>2019</year>
          . Opy`
          <article-title>t sozdaniya mashinnoproveryaemoj semanticheskoj razmetki russkix sushhestvitel`ny`x [Developing computationally verifiable semantic annotation of Russian nouns]</article-title>
          . Presented at the Annual International Conference “Dialogue,” Moscow.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Kutuzov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fares</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oepen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velldal</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <year>2017</year>
          .
          <article-title>Word vectors, reuse, and replicability: Towards a community repository of large-text resources</article-title>
          ,
          <source>in: Proceedings of the 58th Conference on Simulation and Modelling</source>
          . Linköping University Electronic Press, pp.
          <fpage>271</fpage>
          -
          <lpage>276</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Lakoff</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Johnson,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ,
          <year>1980</year>
          .
          <article-title>Metaphors we Live by</article-title>
          , 2nd ed. The University of Chicago Press, Chicago-London.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Macmillan</given-names>
            <surname>Dictionary</surname>
          </string-name>
          ,
          <article-title>Free English Dictionary and Thesaurus [WWW Document]</article-title>
          , n.d. URL https://www.macmillandictionary.
          <source>com/ (accessed 11.7</source>
          .20).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Parker</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graff</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kong</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maeda</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
          <string-name>
            <given-names>English</given-names>
            <surname>Gigaword Fifth Edition LDC2011T07 (Tech. Rep</surname>
          </string-name>
          .).
          <source>Technical Report. Linguistic Data Consortium</source>
          , Philadelphia.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
          <article-title>Scikit-learn: Machine Learning in</article-title>
          <source>Python Journal of Machine Learning Research.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Tsvetkov</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mukomel</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gershman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2013</year>
          .
          <article-title>Cross-lingual metaphor detection using common semantic features</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Metaphor in NLP</source>
          . pp.
          <fpage>45</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Turney</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neuman</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Assaf</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
          <article-title>Literal and metaphorical sense identification through concrete and abstract context</article-title>
          ,
          <source>in: Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics</source>
          , pp.
          <fpage>680</fpage>
          -
          <lpage>690</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>